Speech output control device, speech output method, and program
Patent Information
- Application Number
- PCT/JP2026/000051
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-09-11
- Filing Date
- 2026-01-05
- Publication Date
- 2026-10-01
Smart Images

Figure JP2026000051_01102026_PF_FP_ABST
Abstract
Description
Audio output control device, audio output method and program
[0001] The present invention relates to an audio output control device, an audio output method and a program.
[0002] There is a technology for extracting characters from an image captured by a camera and outputting the characters as audio (see, for example, Patent Document 1). Patent Document 1 discloses a glasses-type device that identifies and reads out character groups from an image captured around a user. In Patent Document 1, in order to address slow processing caused by information overload, audio is output for character groups with high priority based on factors such as character size.
[0003] Japanese Unexamined Patent Application Publication No. 2022-092837
[0004] Devices with a camera like those described in Patent Document 1 have been increasingly reduced in size and weight, such as in the glasses-type and earphone-type forms, and it is expected that such devices will be worn on a daily basis. In such a situation, it is required that for information that is difficult for the user to visually identify, the user can check the information while facing the information without blocking the user's visual field.
[0005] The present disclosure has been made in view of the above, and an object of the present disclosure is to enable a user to check information that is difficult for the user to visually identify while the user is facing the information without blocking the user's visual field.
[0006] In order to solve the above-mentioned problem and achieve the object, an audio output control device according to the present disclosure comprises: a detection unit that detects specific information, which is a candidate to be read aloud, from video captured by a camera that captures video in front of a user; a determination unit that determines, with respect to the specific information detected by the detection unit, whether or not the specific information is to be conveyed to the user; an information acquisition unit that acquires read-aloud information for conveying the specific information, which has been determined by the determination unit to be information to be conveyed to the user, to the user; a speech synthesis processing unit that generates audio information indicating the read-aloud information acquired by the information acquisition unit; and an output control unit that controls the audio information generated by the speech synthesis processing unit to be output as audio from an audio output unit to the user.
[0007] The audio output method relating to this disclosure involves an audio output control device performing the following actions: detecting specific information that is a candidate for reading aloud from video footage captured by a camera that captures video footage in front of the user; determining whether the detected specific information is to be conveyed to the user; acquiring reading information to convey to the user the specific information that has been determined to be to be conveyed to the user; generating audio information indicating the acquired reading information; and controlling the generated audio information to be output as audio from the audio output unit to the user.
[0008] The program relating to this disclosure causes an audio output control device to perform the following processes: detecting specific information that is a candidate for reading aloud from video footage captured by a camera that captures video footage in front of the user; determining whether the detected specific information is to be conveyed to the user; acquiring reading aloud information to convey to the user the specific information that has been determined to be to be conveyed to the user; generating audio information indicating the acquired reading aloud information; and controlling the generated audio information to be output as audio to the user from the audio output unit.
[0009] According to this disclosure, the system has the effect of allowing users to check information that is difficult for them to identify visually, without obstructing their field of view, while they are facing that information.
[0010] Figure 1 is a schematic diagram showing an example configuration of the audio output system according to the first embodiment. Figure 2 is a schematic front view showing the earphones according to the first embodiment. Figure 3 is a block diagram showing an example configuration of the earphones according to the first embodiment. Figure 4 is a block diagram showing an example configuration of the information terminal device according to the first embodiment. Figure 5 is a diagram showing an example of video captured by a camera. Figure 6 is a flowchart showing the processing flow in the audio output system according to the first embodiment. Figure 7 is a flowchart showing the processing flow in the audio output system according to the first embodiment. Figure 8 is a block diagram showing an example configuration of the information terminal device according to the second embodiment. Figure 9 is a flowchart showing the processing flow in the audio output system according to the first embodiment.
[0011] Embodiments of the audio output system 1 according to this disclosure will be described in detail below with reference to the attached drawings. However, the present invention is not limited to the following embodiments.
[0012] [First Embodiment] <Audio Output System> Figure 1 is a schematic diagram showing an example of the configuration of an audio output system according to the first embodiment. The audio output system 1 includes an earphone 10 and an information terminal device (audio output control device) 40. The earphone 10 and the information terminal device 40 are connected by wire or wireless communication. In the following description, the earphone 10 and the information terminal device 40 will be described as being connected by wireless communication. If specific information detected from the image captured by the camera 16 placed in the earphone 10 is difficult to identify by visual inspection, the audio output system 1 outputs spoken information as audio to convey the specific information to the user. After outputting the audio, the audio output system 1 acquires additional information based on the captured image and outputs it as audio.
[0013] The specific information is detected from the video captured by the camera 16. The specific information is a candidate for reading aloud via the audio output unit 19 of the earphone 10. The detection process for this specific information will be described later.
[0014] <Earphones> The earphones 10 have multiple independent housings 11 that are worn on the user's ears. The earphones 10 can be, for example, wireless earphones, wired earphones, etc. The form of the earphones 10 can be any form, for example, open-ear type, canal type, etc. Figures 1 and 2 show an example of open-ear type earphones with TWS (True Wireless Stereo). The earphones 10 have a left earphone 10L for the left ear that is worn on the left ear, and a right earphone 10R for the right ear that is worn on the right ear. The left earphone 10L outputs left channel data of the audio signal. The right earphone 10R outputs right channel data of the audio signal. In the following description, when it is not necessary to distinguish between the left earphone 10L and the right earphone 10R, they will be described as earphones 10.
[0015] The earphone 10 will be explained using Figures 1 to 3. Figure 1 is a schematic side view of the earphone 10 according to the embodiment, showing how it looks when the user is wearing the earphone 10 and is viewed from the left or right direction of the user. In Figure 1, the left earphone 10L is shown as viewed from the left side of the user when it is worn on the user's left ear, and the right earphone 10R is shown as viewed from the right side of the user when it is worn on the user's right ear. Figure 2 is a schematic front view of the earphone according to the first embodiment, showing how it looks when the user is wearing the earphone 10 and is viewed from the front of the user. In Figure 2, the right earphone 10R is shown, and the left earphone 10L has a similar configuration. Figure 3 is a block diagram showing an example of the configuration of the earphone according to the first embodiment. The earphone 10 acquires data, including music data, from the information terminal device 40. The earphone 10 includes an operation unit 15, a camera 16, a microphone 17, an amplification unit 18, an audio output unit 19, a communication unit 27, a battery 28, a power supply unit 29, and a control unit 30.
[0016] The housing 11 of the earphone 10 is worn on the user's ear. The housing 11 defines the external shape of the earphone 10. The housing 11 comprises a main body 11a, a first ear hook 11b, and a second ear hook 11c. The main body 11a is the part that is positioned inside the user's auricle when the user wears the earphone 10. The first ear hook 11b is positioned between the helix of the auricle and the head when the user wears the earphone 10. The second ear hook 11c is the part that is positioned behind the auricle when the user wears the earphone 10. The main body 11a, the first ear hook 11b, and the second ear hook 11c are integrally formed. The housing 11 houses an operation unit 15, a camera 16, a microphone 17, an amplification unit 18, an audio output unit 19, a communication unit 27, a battery 28, a power supply unit 29, and a control unit 30. The arrangement of the operation unit 15, camera 16, microphone 17, amplification unit 18, audio output unit 19, communication unit 27, battery 28, power supply unit 29, and control unit 30 in the housing 11 is just one example.
[0017] The earphone 10 has a left housing 11L that is worn on the user's left ear and a right housing 11R that is worn on the right ear. When it is not necessary to distinguish between the left housing 11L and the right housing 11R, they will be described as housing 11.
[0018] The operation unit 15 is a touch sensor located on the main unit 11a. The operation unit 15 can receive various operations, such as playing and stopping music data on the earphone 10. The operation unit 15 can receive various operations, such as voice recognition and video recognition on the earphone 10. The operation unit 15 outputs an operation signal indicating the received operation to the operation control unit 32.
[0019] Camera 16 captures images of the area around the earphone 10. Camera 16's shooting is controlled by the shooting control unit 33. Camera 16 outputs the captured image to the shooting control unit 33. Camera 16 comprises a left camera 16L located in the left housing 11L of the left earphone 10L, and a right camera 16R located in the right housing 11R of the right earphone 10R. In the description of this embodiment, when it is not necessary to distinguish between the left camera 16L and the right earphone 10R, they will be described as camera 16.
[0020] The left camera 16L, located on the left earphone 10L, is positioned on the main body 11La of the left housing 11L. The left camera 16L is positioned on the main body 11La so that it faces forward when the user wears the left earphone 10L on their left ear; in other words, the shooting direction is in front of the user.
[0021] The right camera 16R, located on the right earphone 10R, is positioned on the main body 11Ra of the right housing 11R. The right camera 16R is positioned on the main body 11Ra such that it faces forward when the user wears the right earphone 10R in their right ear; in other words, the shooting direction is in front of the user.
[0022] The microphone 17 picks up sounds from the vicinity of the earphone 10. The microphone 17 can pick up various sounds, including, for example, voice commands to the earphone 10. The microphone 17 outputs the picked-up sounds to the voice input control unit 34. The microphone 17 is located in at least one of the left housing 11L and the right housing 11R. The microphone 17 is located in the main body 11a.
[0023] The amplification unit 18 amplifies the audio output from the audio output unit 19. The amplification unit 18, for example, performs D / A conversion on audio channel data, including music data acquired from the information terminal device 40, and amplifies it. The audio output unit 19 is located in the main unit 11a.
[0024] The audio output unit 19 outputs sound for the user to hear. The audio output unit 19 outputs sound based on a music signal, for example. The audio output control unit 35 controls the output of sound corresponding to text data of read-aloud information generated in the information terminal device 40, which will be described later. The audio output unit 19 outputs an audio signal amplified by the amplification unit 18. The audio output unit 19 is located in at least one of the left housing 11L and the right housing 11R. The audio output unit 19 is located in the main body 11a.
[0025] The communication unit 27 is a communication unit. The communication unit 27 is capable of wireless communication including, for example, Bluetooth®, Wi-Fi®, or NFMI (Near Field Magnetic Induction). In this embodiment, the communication unit 27 is connected to the information terminal device 40 using the Bluetooth method. The communication unit 27 is located, for example, in the second ear hook portion 11c.
[0026] The communication unit 27 is connected to the information terminal device 40 so as to be able to send and receive data. The communication unit 27 receives data, including music data, from the information terminal device 40, for example. In this embodiment, the communication unit 27 transmits to the information terminal device 40 video footage captured by the camera 16 acquired by the shooting control unit 33, and audio data based on audio signals picked up by the microphone 17 acquired by the audio input control unit 34. In this embodiment, the communication unit 27 receives audio data of generated read-aloud information from the information terminal device 40.
[0027] The battery 28 supplies power for the earphone 10 to operate. The battery 28 is a rechargeable battery that can be repeatedly charged and discharged. The battery 28 is, for example, a nickel-metal hydride rechargeable battery, a lithium-ion rechargeable battery, or a lithium polymer rechargeable battery. The battery 28 is built into the earphone 10. The charging and discharging of the battery 28 is controlled by the power control unit 39. The battery 28 is located, for example, in the second ear hook portion 11c.
[0028] The power supply unit 29 can be connected to the power supply of an earphone case (not shown) and supplies power to the battery 28. The power supply unit 29 is connected to the power control unit 39. The power supply unit 29 may also be capable of supplying power from a DC power source such as a mobile battery or from a contactless charging device.
[0029] <Earphone Control Unit> The control unit 30 is a processing unit composed of, for example, a CPU (Central Processing Unit). The control unit 30 loads a program stored in a storage unit (not shown) into memory and executes the instructions contained in the program. The control unit 30 includes an internal memory (not shown) which is used for temporary data storage, etc. The control unit 30 has an operation control unit 32, a shooting control unit 33, an audio input control unit 34, an audio output control unit 35, a communication control unit 37, and a power supply control unit 39.
[0030] The operation control unit 32 acquires operation signals from the operation unit 15 in response to operations performed on the operation unit 15. The operation control unit 32 outputs control signals corresponding to the acquired operation signals to each part of the earphone 10 and the information terminal device 40.
[0031] The shooting control unit 33 controls the shooting by the camera 16. The shooting control unit 33 acquires the video captured by the camera 16. In this embodiment, the shooting control unit 33 uses the communication unit 27 via the communication control unit 37 to transmit the captured video to the information terminal device 40.
[0032] The audio input control unit 34 acquires the audio signal picked up by the microphone 17. The audio input control unit 34 performs A / D conversion on the audio signal picked up by the microphone 17 and acquires it as audio data.
[0033] The audio output control unit 35 controls the output of audio from the audio output unit 19. The audio output control unit 35 performs processing to output audio to the audio output unit 19. The audio output control unit 35 controls the output of audio data, such as music data, from the audio output unit 19. The audio output control unit 35 controls the output of audio corresponding to the text data of the read-aloud information generated in the information terminal device 40 (described later) to the user as audio from the audio output unit 19. The audio output control unit 35 controls the output of audio information indicating additional information as audio. The audio output control unit 35 outputs the audio data acquired from the information terminal device 40 to the amplification unit 18.
[0034] The communication control unit 37 communicates wirelessly with the information terminal device 40 by controlling the communication unit 27. For example, the communication control unit 37 controls the information terminal device 40 to receive data including music data. In this embodiment, the communication control unit 37 transmits to the information terminal device 40 video footage captured by the camera 16 and audio data based on audio signals picked up by the microphone 17 acquired by the audio input control unit 34. In this embodiment, the communication control unit 37 controls the information terminal device 40 to receive audio information corresponding to the read-aloud information generated in the information terminal device 40, which will be described later.
[0035] The power control unit 39 is a charging circuit that charges the battery 28. The power control unit 39 is connected to the battery 28 and the power supply unit 29. The power control unit 39 supplies power from the power supply unit 29 to the battery 28.
[0036] <Information Terminal Device> The information terminal device 40 will be described using Figure 4. Figure 4 is a block diagram showing an example configuration of the information terminal device according to the first embodiment. The information terminal device 40 is a portable electronic device used by the user of the earphone 10, such as a smartphone or a tablet device. When specific information detected from the image captured by the camera 16 placed on the earphone 10 is difficult to identify by visual inspection, the information terminal device 40 outputs spoken information as audio to convey the specific information to the user. After outputting the audio, the information terminal device 40 acquires additional information based on the captured image and outputs it as audio. The information terminal device 40 comprises a communication unit 41, a visual acuity information storage unit 42, and a control unit 50.
[0037] The communication unit 41 is a communication unit. The communication unit 41 is connected to the earphone 10 so as to be able to send and receive data. For example, the communication unit 41 transmits data including music data to the earphone 10. In this embodiment, the communication unit 41 receives video footage captured by the camera 16 and audio data based on audio signals picked up by the microphone 17 acquired by the audio input control unit 34 from the earphone 10. In this embodiment, the communication unit 41 transmits the video footage captured by the camera 16 from the earphone 10, along with instructions such as prompts to perform the acquisition of specific information. In this embodiment, the communication unit 41 receives text data indicating spoken information of specific information, which is the result of object recognition, from the internet search engine, map information, or AI server. In this embodiment, the communication unit 41 converts text data indicating audio information corresponding to the spoken information generated by the information terminal device 40 (described later) into audio data and transmits it to the earphone 10. The communication unit 41 is capable of short-range wireless communication, including Bluetooth, for example. In this embodiment, the communication unit 41 is connected to the earphone 10 using the Bluetooth method.
[0038] The visual acuity information storage unit 42 is a storage device that stores the user's visual acuity information.
[0039] Visual acuity information includes, for example, uncorrected visual acuity and myopia information, such as the degree of myopia expressed as 1.00D. Visual acuity information also includes, for example, hyperopia information, such as the add power or age (accommodative power). Visual acuity information may also include the degree of vision correction, such as whether the user currently wears glasses.
[0040] Visual acuity information is input in advance by the user via an input unit (not shown) of the information terminal device 40. Visual acuity information can also be input by the user via voice input through the microphone 17 of the earphone 10.
[0041] The video acquired by the information terminal device 40 from the earphone 10 can be used for various purposes. For example, the video may be recorded. For example, specific information that is a candidate for reading aloud may be detected from the video, and audio data indicating the reading aloud information related to that specific information may be fed back to the earphone 10. The process of acquiring the reading aloud information is performed by the information acquisition unit 55, which will be described later. The process of acquiring the reading aloud information may use an internet search engine, map information, or an AI server. In this case, for example, instructions to perform the process of acquiring specific information are sent to the internet search engine, map information, or AI server along with the video acquired from the earphone 10. The internet search engine, map information, or AI server performs the process of acquiring specific information and sends text data, audio data, etc., indicating information about the specific information that is a candidate for reading aloud to the information terminal device 40. The information terminal device 40 converts the specific information that is a candidate for reading aloud received from the internet search engine, map information, or AI server into audio data and sends it to the earphone 10.
[0042] The user's spoken voice, which is the audio acquired by the information terminal device 40 from the earphone 10, can be used for various purposes. For example, the voice may be recorded.
[0043] <Control Unit of Information Terminal Device> The control unit 50 is an arithmetic processing unit composed of, for example, a CPU. The control unit 50 loads a program stored in a storage unit (not shown) into memory and executes the instructions contained in the program. The control unit 50 includes an internal memory (not shown) which is used for temporary data storage, etc. In this embodiment, the control unit 50 implements the function of a control unit of the voice output system 1. If specific information detected from the video captured by the camera 16 placed on the earphone 10 is difficult to identify by visual inspection, the control unit 50 outputs spoken information as voice to convey the specific information to the user. After outputting the voice, the control unit 50 acquires additional information based on the captured video and outputs it as voice. The control unit 50 includes a communication control unit (output control unit) 51, a voice recognition processing unit 52, a detection unit 53, a determination unit 54, an information acquisition unit 55, and a voice synthesis processing unit 59.
[0044] The communication control unit 51 performs wireless communication with the earphone 10 by controlling the communication unit 41. For example, the communication control unit 51 controls transmission of data including music data to the earphone 10. In the present embodiment, the communication control unit 51 performs control to receive, from the earphone 10, video captured by the camera 16 and audio data based on an audio signal collected by the microphone 17 acquired by the audio input control unit 34. In the present embodiment, the communication control unit 51 transmits, to an Internet search engine, map information, or an AI server, an instruction for causing the same to execute acquisition processing of specific information along with the video captured by the camera 16 from the earphone 10. In the present embodiment, the communication control unit 51 receives text data indicating an object recognition result from the Internet search engine, map information, or the AI server. In the present embodiment, the communication control unit 51 performs control to transmit, to the earphone 10, audio data indicating specific information that is a reading target candidate converted by the speech synthesis processing unit 59. The communication control unit 51 implements a function as an output control unit that controls output of the audio information generated by the speech synthesis processing unit 59 to the user as audio from the audio output unit 19 of the earphone 10.
[0045] The speech recognition processing unit 52 performs speech recognition processing on audio data based on an audio signal collected by the microphone 17 disposed in the earphone 10. More specifically, the speech recognition processing unit 52 performs speech recognition processing for recognizing utterance content by natural language processing or the like from audio data based on an audio signal collected by the microphone 17 disposed in the earphone 10 via the communication unit 41.
[0046] The detection unit 53 detects specific information that is a reading target candidate from the video captured by the camera 16.
[0047] For example, when the user turns his or her face, if there is information such as characters, character strings, and symbols present at the center position of the video, the detection unit 53 detects such character strings or the like as the specific information.
[0048] The center position is, for example, a predetermined range including the vertical and horizontal centers of the video.
[0049] A symbol is, for example, a logo, a symbol mark, an image, a pictogram, or a sign.
[0050] The detection unit 53 detects, for example, information such as characters, strings of characters, and symbols as specific information if, for example, the user turns their face towards the camera and information such as characters, strings of characters, and symbols is continuously present in the center of the camera's field of view for a predetermined period of time, such as three seconds or more.
[0051] The detection unit 53 may, for example, detect characters, strings of characters, symbols, or other information located in the center of an image captured at a position where the user's head is temporarily stopped while the user is moving their head to look around, as specific information.
[0052] Figure 5 shows an example of video footage captured by a camera. In Figure 5, the central position of the camera's field of view is the area A1 enclosed by a dashed line. In the example shown in Figure 5, the detection unit 53 detects the string "Kanagawa Hospital" as specific information that exists within area A1, the central position of the camera's field of view. Area A1 can be arbitrarily set by the camera's focal length or user settings.
[0053] The determination unit 54 determines whether the specific information detected by the detection unit 53 is intended to be communicated to the user. In this embodiment, the determination unit 54 determines whether the specific information is intended to be communicated to the user based on whether the specific information is difficult for the user to identify visually. More specifically, the determination unit 54 determines whether the specific information detected by the detection unit 53 is difficult for the user to identify visually.
[0054] The determination unit 54 may determine, based on the size of the specific information in the video and the user's visual acuity information, whether or not it is difficult for the user to visually read the characters or recognize the shape of the specific information. The determination unit 54 determines whether or not the size of the specific information in the video is below a size threshold that makes it difficult for the user to visually identify it.
[0055] The higher the visual acuity, the smaller the size of specific information that can be identified; conversely, the lower the visual acuity, the larger the size of specific information that can be identified. This relationship between visual acuity and the size of specific information that can be identified, in other words, the relationship between visual acuity and the size threshold at which information becomes difficult to identify by visual inspection, is assumed to be stored in a memory unit (not shown) beforehand.
[0056] The size threshold may be determined based on the user's visual acuity information, which is entered in advance by the user. The size threshold may also be changed based on the size of the specific information in the video captured when the instruction to read the specific information is given, or based on the distance from the user to the specific information. The size threshold may also be changed based on the ambient illumination.
[0057] If there are multiple pieces of specific information, the determination unit 54 may determine the order in which to output audio based on the size of the specific information in the video. If there are multiple pieces of specific information, the determination unit 54 may obtain the size of each piece of specific information from the video and determine the order in which to output audio in descending order of size. Alternatively, the determination unit 54 may do the opposite and determine the order in which to output audio in ascending order of character size.
[0058] In the case of Figure 5, the size of the text "Kanagawa Hospital" in the video is determined by the determination unit 54 to be below the size threshold at which it becomes difficult for the user to visually identify it, and is therefore determined to be specific information that is difficult for the user to visually identify.
[0059] The information acquisition unit 55 acquires read-aloud information to be conveyed to the user regarding specific information that the determination unit 54 has determined to be the subject of communication with the user. More specifically, the information acquisition unit 55 acquires read-aloud information to be conveyed to the user regarding specific information that the determination unit 54 has determined to be difficult for the user to identify visually, via the communication unit 27 from internet search engines, map information, or AI servers.
[0060] The information acquisition unit 55, based on the video captured after the audio output unit 19 of the earphone 10 outputs audio corresponding to the read-aloud information, acquires additional information related to the audio information from an internet search engine, map information, or AI server via the communication unit 27.
[0061] Additional information is further information related to specific information. Additional information may be presented in a stepwise manner, deepening the information, or in other words, becoming more detailed. The level of detail of additional information may vary depending on its distance from each specific piece of information. For example, the closer the distance, the more detailed the additional information; the further away, the more general the additional information. The level of detail of additional information may also be changed by user settings.
[0062] The information acquisition unit 55 acquires, for example, changes in the position of specific information in the image captured a predetermined time after the output of the audio corresponding to the spoken information, and acquires additional information if the change in position is smaller than the change threshold. The information acquisition unit 55 also acquires additional information if, for example, after the output of the audio corresponding to the spoken information, the user is intently watching the specific information, that is, the specific information is continuously located in the center of the camera's field of view for a predetermined time or longer. This means that additional information is acquired if the user has not moved very close to the specific information after the output of the audio corresponding to the spoken information, and the information remains identifiable by visual inspection.
[0063] The change threshold is, for example, 10% vertically and 10% horizontally across the entire video capture range.
[0064] The predetermined time is, for example, 3 seconds later.
[0065] The information acquisition unit 55 acquires additional information, for example, when a predetermined time has passed since the output of the audio corresponding to the spoken information, and the specific information has changed from a state where it was difficult for the user to visually identify to a state where it has become identifiable. This means that additional information is acquired when the user approaches the specific information after the output of the audio corresponding to the spoken information, making it identifiable by visual inspection. For example, this occurs when the user has heard the audio corresponding to the spoken information and is approaching the specific information while remaining interested in it.
[0066] Here, we will explain using the video shown in Figure 5 as an example. In the case of Figure 5, if the specific information detected by the detection unit 53 is "Kanagawa Hospital", the information acquisition unit 55 acquires information that will be used to convey to the user, such as "Kanagawa Byoin" through speech synthesis, which will be the reading of the detected string.
[0067] The information acquisition unit 55, for example, after outputting the string of specific information "Kanagawa Hospital" as audio, if the user continues to look at the specific information for a predetermined period of time, such as three seconds or more, it acquires further information related to that specific information and outputs it as additional information as audio. For example, this occurs when the user has heard the audio corresponding to the read-aloud information and continues to look at the specific information with interest.
[0068] The information acquisition unit 55, for example, after the audio output of the string of specific information "Kanagawa Hospital," compares the video taken a predetermined time later, such as 3 seconds later, with the video taken at the time the specific information was detected by the detection unit 53, and acquires the change in the position of the specific information "Kanagawa Hospital" in the video. If the change in position is below a change threshold, the information acquisition unit 55 determines that the user is continuously looking at the specific information "Kanagawa Hospital" and acquires information related to the specific information "Kanagawa Hospital" from an internet search engine, map information, or an AI server. The information acquisition unit 55 acquires, for example, "Kanagawa Hospital, reception starts at 8am" as additional information.
[0069] For example, after the audio output of the string "Kanagawa Hospital" of the specific information, if the user continues to face the specific information and approaches it, and the specific information becomes identifiable, the information acquisition unit 55 acquires further information related to that specific information and acquires it as additional information.
[0070] The information acquisition unit 55, for example, compares the video footage taken after a predetermined time, such as 3 seconds, following the audio output of the string of specific information "Kanagawa Hospital," with the video footage taken at the time the specific information was detected by the detection unit. The unit acquires the change in the font size of the specific information "Kanagawa Hospital" in the video and determines whether the font size is below a size threshold that makes it difficult for the user to visually identify it. If the change in font size is greater than the change threshold, the information acquisition unit 55 determines that the user has approached the specific information "Kanagawa Hospital" and can now visually identify it. In this case, the information acquisition unit 55 acquires additional information about the specific information "Kanagawa Hospital" from an internet search engine, map information, or an AI server.
[0071] The speech synthesis processing unit 59 generates speech information indicating the read-aloud information acquired by the information acquisition unit 55. For example, the speech synthesis processing unit 59 converts the text data of the read-aloud information acquired by the information acquisition unit 55 via the communication control unit 51 from an internet search engine, map information, or an AI server into speech data.
[0072] In the case of Figure 5, the speech synthesis processing unit 59 performs speech synthesis based on text data such as "Kanagawa Hospital" and outputs the voice to the user from the voice output unit 19 of the earphone 10. The speech synthesis processing unit 59 is not limited to outputting only the voice of "Kanagawa Hospital," but may also synthesize additional voice to inform the user that what is "Kanagawa Hospital," such as "The full display is Kanagawa Hospital."
[0073] The speech synthesis processing unit 59 generates speech information indicating additional information.
[0074] <Audio Output Method> Next, an example of information processing in the audio output system 1 will be explained using Figure 6. Figure 6 is a flowchart showing the processing flow in the audio output system according to the first embodiment. When the earphone 10 is activated and the object recognition function is turned ON, the processing shown in the flowchart in Figure 6 is executed. During the execution of the processing shown in the flowchart in Figure 6, the audio acquired by the microphone 17 and the video captured by the camera 16 are transmitted to the information terminal device 40.
[0075] The control unit 50 starts the detection process using the detection unit 53 (step S101). The control unit 50 uses the detection unit 53 to detect, for example, characters, strings of characters, and symbols that are located in the center of the image when the user turns their face towards it, and uses those strings of characters, etc., as specific information. The control unit 50 proceeds to step S102.
[0076] The control unit 50 determines whether or not specific information has been detected in the video by the detection unit 53 (step S102). If the control unit 50 determines that specific information has been detected in the video by the detection unit 53 (step S102; Yes), proceed to step S103. If the control unit 50 does not determine that specific information has been detected in the video by the detection unit 53 (step S102; No), proceed to step S109.
[0077] If the control unit 50 determines that specific information has been detected in the video (step S102; Yes), the control unit 50 uses the determination unit 54 to determine whether or not it is difficult for the user to visually identify the information (step S103). The control unit 50 uses the determination unit 54 to determine whether or not it is difficult for the user to visually identify the specific information detected by the detection unit 53. If the control unit 50 determines that it is difficult for the user to visually identify the information (step S103; Yes), the control unit 50 proceeds to step S104. If the control unit 50 does not determine that it is difficult for the user to visually identify the information (step S103; No), the control unit 50 proceeds to step S109.
[0078] If the control unit 50 determines that it is difficult for the user to identify the information visually (step S103; Yes), the control unit 50 uses the information acquisition unit 55 to acquire information related to the specific information (step S104). The control unit 50 uses the information acquisition unit 55 to acquire read-aloud information to convey to the user about the specific information that the user has determined is difficult to identify visually, from an internet search engine, map information, or an AI server. The control unit 50 then proceeds to step S105.
[0079] The control unit 50 performs speech synthesis processing using the speech synthesis processing unit 59 (step S105). The control unit 50 converts the text data of the read-aloud information acquired by the information acquisition unit 55 into speech data using the speech synthesis processing unit 59. The control unit 50 then proceeds to step S106.
[0080] The control unit 50 outputs the audio data converted by the communication control unit 51 from the audio output unit 19 of the earphone 10 (step S106). The control unit 50 then proceeds to step S107.
[0081] The control unit 50 determines, using the information acquisition unit 55, whether or not there is movement in the image (step S107). The control unit 50 determines, using the information acquisition unit 55, for example, that there is movement in the image if the change in the position of specific information in the image captured a predetermined time after the output of the audio corresponding to the spoken information is greater than the change threshold. Alternatively, the control unit 50 determines, using the information acquisition unit 55, that there is movement in the image if, for example, after the output of the audio corresponding to the spoken information, the specific information has not been continuously present at the center of the camera's field of view for a predetermined time or longer. If the control unit 50 determines that there is movement in the image (step S107; Yes), it proceeds to step S109. If the control unit 50 does not determine that there is movement in the image (step S107; No), it proceeds to step S108.
[0082] If it is not determined that there is movement in the video (step S107; No), the control unit 50 acquires additional information, which is further information, using the information acquisition unit 55 (step S108). The control unit 50 repeats the process of step S105. In this case, steps S105 and S106 perform audio synthesis and output based on the additional information.
[0083] The control unit 50 determines whether or not to terminate the process (step S109). More specifically, the control unit 50 determines to terminate the process, for example, when an operation to turn off the object recognition function is detected, or when the power of the earphone 10 is turned off. If the control unit 50 determines to terminate the process (step S109; Yes), it terminates the process in this flowchart. If the control unit 50 does not determine to terminate the process (step S109; No), it executes the process in step S102 again.
[0084] Next, another example of information processing in the voice output system 1 will be described using Figure 7. Figure 7 is a flowchart showing the processing flow in the voice output system according to the first embodiment. Steps S201 to S206, S208, and S209 in Figure 7 perform the same processing as steps S101 to S106, S108, and S109 in the flowchart of Figure 6.
[0085] The control unit 50 uses the determination unit 54 to determine whether the previously identified information is difficult for the user to identify visually (step S207). The control unit 50 uses the determination unit 54 to determine whether the previously identified information detected in step S202 is difficult for the user to identify visually. If the control unit 50 uses the determination unit 54 to determine whether the previously identified information is difficult for the user to identify visually (step S207; Yes), it proceeds to step S209. If the control unit 50 uses the determination unit 54 to determine whether the previously identified information is difficult for the user to identify visually (step S207; No), it proceeds to step S208.
[0086] <Effects> As described above, in this embodiment, if specific information detected from the video captured by the camera 16 positioned in the earphone 10 is difficult to identify visually, it is possible to output spoken information as audio to inform the user about the specific information. In this embodiment, additional information can be acquired and output as audio based on the video captured after the audio is output. According to this embodiment, even for information that is difficult for the user to identify visually, the user can confirm the information by audio without obstructing their field of view, while facing the information.
[0087] In this embodiment, it is possible to determine whether or not a user has difficulty visually identifying specific information based on the size of the specific information in the image and the user's visual acuity information. According to this embodiment, it is possible to appropriately determine whether or not a user has difficulty visually identifying specific information.
[0088] In this embodiment, it is possible to determine whether or not the specific information is something that should be conveyed to the user, based on the calculated distance from the user to the specific information and the user's visual acuity information. According to this embodiment, it is possible to appropriately determine whether or not the user has difficulty visually identifying the information.
[0089] In this embodiment, if there are multiple pieces of specific information, the order in which the audio is output can be determined based on the size of the specific information in the video.
[0090] [Second Embodiment] The voice output system 1 according to this embodiment will be described with reference to Figure 8. Figure 8 is a block diagram showing an example of the configuration of an information terminal device according to the second embodiment. The basic configuration of the voice output system 1 is the same as that of the voice output system 1 of the first embodiment. In the following description, components that are the same as those of the voice output system 1 are denoted by the same reference numerals or corresponding reference numerals, and their detailed descriptions are omitted. In this embodiment, the control unit 50A of the information terminal device 40A is equipped with a distance calculation unit 56A, and the processing in the visual acuity information storage unit 42A, the determination unit 54A and the information acquisition unit 55A differs from that of the first embodiment.
[0091] For example, when characters of different sizes are mixed in specific information displayed as a single group, such as a character string within a single sign, only characters of a small size are detected as the specific information. In this case, only the small-size characters detected as the specific information are read aloud, and there is a possibility that interrupted information that does not make sense is output as audio. For example, if a character string "1ごうしゃ3ばんどあ" exists in a sign, there is a possibility that only the small-size character "ゃ" is detected as the specific information. Furthermore, if the size of the numeral characters in the sign "1号車3番ドア" (Car 1, Door 3) is large, there is a possibility that the character string excluding the numerals, " 号車番ドア", is detected as the specific information. In the present embodiment, specific information displayed as a single group, such as a character string within a single sign, is detected as one piece of specific information.
[0092] The visual acuity information storage unit 42A may store a plurality of pieces of visual acuity information for each user. For example, the visual acuity information storage unit 42A may store visual acuity information indicating unassisted visual acuity for each user and visual acuity when wearing glasses.
[0093] The visual acuity information is input for each state of the user when wearing the earphone 10. In other words, the visual acuity information can be switched according to the state of the user when wearing the earphone 10.
[0094] A distance calculation unit 56A calculates the distance from the user to the specific information.
[0095] When there are a plurality of pieces of specific information, the distance calculation unit 56A acquires the distance to each piece of specific information. For example, when a plurality of characters are mixed, the distance calculation unit 56A acquires the distance to each character. For example, for specific information displayed as a single group such as a character string within a single sign, the distance calculation unit 56A calculates the same distance for all of the specific information.
[0096] The determination unit 54A determines whether or not the information is to be conveyed to the user, based on the distance from the user to the specific information calculated by the distance calculation unit 56A and the user's visual acuity information. More specifically, the determination unit 54A determines whether or not the specific information detected by the detection unit 53 is to be conveyed to the user, based on the distance from the user to the specific information calculated by the distance calculation unit 56A and the user's visual acuity information. The determination unit 54A determines whether or not the specific information detected by the detection unit 53 can be identified by reading the characters or recognizing the shape of the specific information, based on the distance from the user to the specific information and the user's visual acuity information. As a result, the determination unit 54A similarly determines whether or not specific information displayed as a single unit, such as a string of characters within a single sign, is to be conveyed to the user, for which the same distance has been calculated.
[0097] The determination unit 54A may switch the visual acuity information according to the user's state and determine what to convey to the user. For example, the determination unit 54A may switch the visual acuity information according to the degree of the user's vision correction and determine what to convey to the user.
[0098] User status refers to whether the user is seeing with or without glasses, etc.
[0099] The determination unit 54A may determine whether to detect specific information based on whether the near and far points of the clear vision range, based on the user's visual acuity information, are outside the threshold range, regardless of the size of the characters in the image. The determination unit 54A determines whether the distance from the user to the specific information in the image is outside the threshold range (outside the clear vision range). If the distance from the user to the specific information in the image is outside the threshold range, the determination unit 54A determines that it is something to be conveyed to the user.
[0100] The range of clear vision generally refers to the depth range in which objects can be clearly seen with the naked eye or with corrected vision (including glasses and contact lenses). The farthest point is called the "far point," and the closest point is called the "near point." The area between the far point and the near point is called the range of clear vision. The range of clear vision may be adjusted based on the ambient light level.
[0101] The near and far points can be calculated based on the user's visual acuity, myopia degree, hyperopia add power, or age. Age is related to the derivation of accommodative power, which is the ability of the lens to focus. This relationship between visual acuity and the range of clear vision, in other words, the relationship between visual acuity and the range of areas that are difficult to distinguish by visual inspection, is assumed to be stored in a memory unit (not shown) beforehand.
[0102] The determination unit 54A may calculate the user's near point and far point from multiple visual acuity information to determine the threshold for the clear vision range, and then determine whether or not specific information exists outside the threshold.
[0103] In the case of Figure 5, the string "Kanagawa Hospital" in the video is determined by the determination unit 54A to be outside the threshold distance and is therefore specific information that should be conveyed to the user.
[0104] The determination unit 54A may determine the order in which to output audio based on the distance to the specific information if there are multiple specific pieces of information. The determination unit 54A may also obtain the specific information from the internet or the like, and refer to at least one of the following: the degree of congestion, the distance from the user, the current time and business hours, to determine the order in which to output audio in order of ease of entry.
[0105] The information acquisition unit 55A acquires read-aloud information to be conveyed to the user for specific information that the determination unit 54A has determined to be to be conveyed to the user.
[0106] The level of detail of additional information may vary depending on the distance to each specific piece of information. For example, the level of detail of additional information may be more detailed for closer distances and more general for farther distances. The level of detail of additional information may also be changed by user settings.
[0107] <Effects> As described above, this embodiment can determine whether or not specific information detected from an image is something to be conveyed to the user, based on the distance from the user to the specific information and the user's visual acuity. According to this embodiment, for example, if specific information displayed as a single unit, such as a string of characters within a single sign, contains characters of different sizes, it can determine whether or not it is something to be conveyed to the user as a single unit. According to this embodiment, if the detected specific information is difficult to identify by visual inspection, it can output audio information to convey the specific information to the user.
[0108] In this embodiment, additional information can be acquired and audio output based on video footage captured after audio output.
[0109] In this embodiment, it is possible to determine which visual acuity information to convey to the user by switching it according to the user's state.
[0110] In this embodiment, if there are multiple pieces of specific information, the order in which audio output is performed can be determined based on the distance to each piece of specific information.
[0111] [Third Embodiment] The voice output system 1 according to this embodiment will be described with reference to Figure 9. Figure 9 is a flowchart showing the processing flow in the voice output system according to the first embodiment. In this embodiment, the processing in the voice recognition processing unit 52, detection unit 53, determination unit 54, and information acquisition unit 55 of the control unit 50 of the information terminal device 40 differs from that of the first embodiment.
[0112] The voice recognition processing unit 52 recognizes the search request voice from the voice data based on the voice signal picked up by the microphone 17 located in the earphone 10 via the communication unit 41.
[0113] A search request voice is a designated voice message that requests the search for specific information, such as a store name or facility name, including phrases like "○○ Cafe" or "Where is ○○ Cafe?".
[0114] The detection unit 53 detects specific information that is a candidate for reading aloud from the video captured by the camera 16. The detection unit 53 may detect multiple specific information items that are candidates for reading aloud from a single video.
[0115] The determination unit 54 determines whether or not the information to be conveyed to the user is to be conveyed to the user, based on the user's voice search request. If the determination unit 54 recognizes the user's voice search request, it determines that the information to be conveyed to the user is to be conveyed to the user. If the determination unit 54 recognizes the user's voice search request and further determines that the information to be conveyed to the user is to be conveyed to the user if the store name or facility name included in the voice search request is detected as specific information by the detection unit 53. Specifically, if the user makes a voice search request that includes "○○ Cafe" and "○○ Cafe" is detected as specific information, the information acquisition unit 55 acquires the read-aloud information to be conveyed to the user as follows.
[0116] The information acquisition unit 55, based on the user's voice search request determined by the determination unit 54, acquires, via the communication unit 27, read-aloud information to be conveyed to the user about specific information that it has determined is the target of the user's request. In this case, the read-aloud information to be conveyed to the user might be something like, "There is a XX Cafe XX branch. It is 30m ahead on your right. It is open until 8pm." Furthermore, location information from the user's current location or related information may be added to this read-aloud information to be conveyed to the user. In addition, visual information such as the shape and color of the target information recognized from the video may be added to the read-aloud information to be conveyed to the user.
[0117] If a user makes a voice search request that includes "○○ Cafe," and "○○ Cafe" is not detected as specific information, the system will determine that there is nothing to tell the user and will output a voice message such as, "There are no ○○ Cafes in the vicinity." If the user requests further additional information, the system may retrieve relevant information from an AI server or similar and output the retrieved additional information as voice.
[0118] When a user makes a voice search request that includes "○○ Cafe," and "○○ Cafe" is not detected as specific information, the system should output a voice message indicating that it is not found. Then, the AI server or the like should obtain, via the communication unit 27, read-aloud information to inform the user about the alternative specific information, specifically "coffee shops, cafes," which is a higher category of "○○ Cafe," from an internet search engine, map information, or the AI server.
[0119] Next, an example of information processing in the audio output system 1 will be explained using Figure 9. The processing in steps S302, S304, and S309 is the same as the processing in steps S102, S104, and S109 of the flowchart shown in Figure 6.
[0120] The control unit 50 starts the detection process using the detection unit 53 (step S301). The control unit 50 detects the characters, strings of characters, and symbols contained in the video as specific information using the detection unit 53. The control unit 50 recognizes the search request voice using the voice recognition processing unit 52. The control unit 50 proceeds to step S302.
[0121] If the control unit 50 determines that specific information has been detected in the video (step S302; Yes), the control unit 50 uses the determination unit 54 to determine whether or not there is something to convey (step S303). The control unit 50 uses the determination unit 54 to determine whether or not there is something to convey to the user based on the user's search request voice. The control unit 50 uses the determination unit 54 to determine if there is something to convey to the user based on the recognition result of the voice recognition processing unit 52, and if it recognizes the user's search request voice, it determines that there is something to convey to the user. If the control unit 50 determines that there is something to convey (step S303; Yes), it proceeds to step S304. If the control unit 50 does not determine that there is something to convey (step S303; No), it proceeds to step S309.
[0122] <Effects> As described above, this embodiment can determine whether or not something is the subject to be conveyed to the user based on the user's voice search request. According to this embodiment, it is possible to determine whether or not a candidate for specific information in the video that matches the voice search request uttered by the user is the subject to be conveyed to the user. According to this embodiment, based on the voice search request from the user, it is possible to output as audio read-aloud information to convey to the user about the specific information that has been determined to be the subject to be conveyed to the user.
[0123] The components of the illustrated audio output system are functionally conceptual and do not necessarily have to be physically configured as shown. In other words, the specific form of each device is not limited to that shown, and all or part of them may be functionally or physically distributed or integrated in any unit depending on the processing load and usage conditions of each device.
[0124] The configuration of the audio output system is implemented, for example, as software, such as a program loaded into memory. In the above embodiment, these were described as functional blocks implemented through the cooperation of hardware or software. That is, these functional blocks can be implemented in various forms using hardware alone, software alone, or a combination thereof.
[0125] The components described above include those that are easily conceivable by those skilled in the art, and those that are substantially identical. Furthermore, the components described above can be combined as appropriate. In addition, various omissions, substitutions, or modifications of the components are possible without departing from the spirit of the present invention.
[0126] In the above embodiment, the operation unit 15, microphone 17, operation control unit 32, and voice recognition processing unit 52 are not essential components.
[0127] In the above description, the earphone 10 was described as having a left earphone 10L and a right earphone 10R that can communicate with the information terminal device 40, but it is not limited to this. The earphone 10 may also be a master unit in which the left earphone 10L or the right earphone 10R can communicate with the portable electronic device 100, and a slave unit in which the right earphone 10R or the left earphone 10L can communicate with the left earphone 10L or the right earphone 10R.
[0128] In the above example, the control unit 50 of the information terminal device 40 implements a function to acquire read-aloud information about specific information, but the invention is not limited to this. For example, the control unit 30 of the earphone 10 may implement a function to acquire read-aloud information about specific information. In this case, the left earphone 10L or the right earphone 10R, which has a control unit 30 that implements the function to acquire read-aloud information about specific information, receives video footage captured from the other right earphone 10R or left earphone 10L.
[0129] In the third embodiment described above, the determination of whether or not something is to be communicated to the user is made based on the voice search request from the user, but the embodiment is not limited to this. For example, when a predetermined operation, including text input, is received by the operation unit 15 of the earphone 10, the determination of whether or not something is to be communicated to the user may be made.
[0130] The audio output control device, audio output method, and program of this disclosure can be used, for example, in an audio output control device which is an information terminal device connected to a camera-equipped device such as an earphone with a camera.
[0131] 1. Audio output system 10. Earphone 11. Housing 15. Operation unit 16. Camera 17. Microphone 18. Amplification unit 19. Audio output unit 27. Communication unit 28. Battery 29. Power supply unit 30. Control unit 32. Operation control unit 33. Shooting control unit 34. Audio input control unit 35. Audio output control unit 37. Communication control unit 39. Power supply control unit 40. Information terminal device (audio output control device) 41. Communication unit 42. Visual acuity information storage unit 50. Control unit 51. Communication control unit (output control unit) 52. Voice recognition processing unit 53. Detection unit 54. Judgment unit 55. Information acquisition unit 59. Voice synthesis processing unit
Claims
1. A voice output control device comprising: a detection unit that detects specific information that is a candidate for reading aloud from video footage captured by a camera that captures video footage in front of the user; a determination unit that determines whether or not the specific information detected by the detection unit is to be conveyed to the user; an information acquisition unit that acquires reading aloud information to be conveyed to the user for the specific information that the determination unit has determined to be to be conveyed to the user; a voice synthesis processing unit that generates voice information indicating the reading aloud information acquired by the information acquisition unit; and an output control unit that controls the voice information generated by the voice synthesis processing unit to be output as voice to the user from a voice output unit.
2. The voice output control device according to claim 1, wherein the determination unit determines whether the specific information is to be conveyed to the user based on whether or not the specific information is difficult for the user to identify by visual inspection.
3. The audio output control device according to claim 2, wherein the determination unit determines whether or not the user has difficulty visually identifying the specific information based on the size of the specific information in the video and the user's visual acuity information.
4. The voice output control device according to claim 3, further comprising: a distance calculation unit for calculating the distance from the user to the specific information, wherein the determination unit determines whether or not the specific information is something to be conveyed to the user, based on the distance from the user to the specific information calculated by the distance calculation unit and the user's visual acuity information.
5. The audio output control device according to claim 2, wherein, if there are multiple pieces of specific information, the determination unit determines the order in which to output audio based on the size of the specific information in the video.
6. The voice output control device according to claim 1, further comprising: a distance calculation unit for calculating the distance from the user to the specific information, wherein the determination unit determines whether or not the information is to be conveyed to the user based on the distance from the user to the specific information calculated by the distance calculation unit and the user's visual acuity information.
7. The audio output control device according to claim 2 or 6, wherein the information acquisition unit acquires additional information relating to the audio information based on video footage captured after the output of the audio information, the audio synthesis processing unit generates audio information indicating the additional information, and the output control unit controls the output of the audio information indicating the additional information as audio.
8. The voice output control device according to claim 6, further comprising: a vision information storage unit that stores multiple vision information for each user, wherein the determination unit determines which of the vision information to switch and convey to the user according to the user's state.
9. The audio output control device according to claim 6, wherein the distance calculation unit obtains the distance to each of the specified pieces of information if there are multiple specified pieces of information, and the determination unit determines the order in which to output audio based on the obtained distances to the specified pieces of information.
10. The voice output control device according to claim 1, wherein the determination unit determines whether or not the search request voice from the user is the subject to be conveyed to the user.
11. An audio output method in which an audio output control device performs the following actions: detecting specific information that is a candidate for reading aloud from video footage captured by a camera that captures video footage in front of the user; determining whether the detected specific information is to be conveyed to the user; acquiring reading aloud information to convey to the user the specific information that has been determined to be to be conveyed to the user; generating audio information indicating the acquired reading aloud information; and controlling the generated audio information to be output as audio from the audio output unit to the user.
12. A program that causes an audio output control device to execute a process including: detecting specific information that is a candidate for reading aloud from video footage captured by a camera that captures video footage in front of the user; determining whether the detected specific information is to be conveyed to the user; acquiring reading aloud information to convey to the user the specific information that has been determined to be to be conveyed to the user; generating audio information indicating the acquired reading aloud information; and controlling the generated audio information to be output as audio to the user from the audio output unit.