Voice guidance device, voice guidance method, and voice guidance program

The voice guidance system addresses the challenge of differentiating voice guidance for visually impaired users by using a detection unit, assignment unit, and output control unit to generate and output unique voice guidance for each user, ensuring effective and confusion-free navigation.

JP7699471B2Active Publication Date: 2025-06-27ALSOK INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021088998
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-05-27
Publication Date
2025-06-27
Estimated Expiration
2041-05-27

AI Technical Summary

Technical Problem

Conventional voice guidance systems for visually impaired users struggle to differentiate between voice guidance intended for one user and that intended for others, leading to confusion and potential misdirection.

Method used

A voice guidance system that includes a detection unit to identify visually impaired users and their positions, an assignment unit to generate guidance voices with different voice quality data for each user, and an output control unit to manage the voice output via corresponding speaker devices.

Benefits of technology

The system effectively allows visually impaired users to recognize that the output voice guidance is intended for themselves, preventing confusion and ensuring that voice guidance functions accurately for each user.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007699471000001
    Figure 0007699471000001
  • Figure 0007699471000002
    Figure 0007699471000002
  • Figure 0007699471000003
    Figure 0007699471000003
Patent Text Reader

Abstract

To make a user who has visual impairment recognize that the guidance voice that is being output is the voice guidance to himself / herself and make the voice guidance effectively function for each user without confusion.SOLUTION: A voice guidance device comprises: a detection unit which detects a user who has visual impairment by analyzing a photographed image captured by a camera device and detects at least a current position of the user who has visual impairment; an allocation unit which generates the guidance voice on the basis of the guidance voice data with the different vocal quality allocated to each user when the detection unit detects the plurality of users who have visual impairment; and an output control unit which controls output of the guidance voice generated on the basis of the guidance voice data with the different vocal quality allocated to each user via a voice output device corresponding to at least the current position of each user who has visual impairment detected by the detection unit.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a voice guidance device, a voice guidance method, and a voice guidance program.

Background Art

[0002] Today, for example, inside or outside facilities such as stations, voice guidance such as "There is a ticket gate 5m ahead" or "There is a crosswalk 3m ahead" is provided to users with visual impairments and the like.

[0003] Also, Patent Document 1 (Japanese Unexamined Patent Application Publication No. 2020-125907) discloses a voice guidance system for visually impaired persons that provides voice guidance according to the moving direction of a user (visually impaired person) passing through the inside of a station. This can prevent unnecessary voice guidance from being given to users passing through the inside of the station, and can reduce the output of unnecessary voice guidance.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, in conventional voice guidance systems including the voice guidance system for visually impaired persons in Patent Document 1, when a plurality of users with visual impairments are present in a nearby position, it is difficult to recognize whether the heard voice guidance is for oneself or for other users with visual impairments.

[0006] For example, assume that a plurality of visually impaired users are located in the same place, and one user and another user are walking in different directions. In this situation, assume that voice guidance such as "There is a ticket gate at the place 5m straight ahead" is given to one user. If the voice guidance given to this one user is misrecognized by another user who is walking in the direction opposite to the one user as voice guidance for oneself, the other user will encounter the inconvenience of not being able to reach the ticket gate even after walking 5m straight ahead.

[0007] The present invention has been made in view of the above-described problems, and an object thereof is to provide a voice guidance device, a voice guidance method, and a voice guidance program that enable a visually impaired user to recognize that the output voice guidance is for oneself, and enable the voice guidance to function effectively for each user without confusion.

Means for Solving the Problems

[0008] In order to solve the above-described problems and achieve the object, the present invention includes a detection unit that detects a visually impaired user and at least the current position of the visually impaired user by analyzing a captured image captured by a camera device; and when a plurality of visually impaired users are detected by the detection unit, an assignment unit that generates a guidance voice based on different voice quality guidance voice data assigned to each user; and an output control unit that controls the output of the guidance voice generated based on the different voice quality guidance voice data assigned to each user via a voice output device corresponding to at least the current position of each visually impaired user detected by the detection unit.

Effects of the Invention

[0009] According to the present invention, it is possible to make a visually impaired user recognize that the output guidance voice is for oneself. Therefore, the voice guidance can function effectively for each user without confusion.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, with reference to the drawings, the voice guidance system according to the embodiment for providing the present invention will be described.

[0012] (System Configuration) FIG. 1 is a diagram showing the system configuration of the voice guidance system according to the embodiment. As shown in FIG. 1, the voice guidance system is configured by connecting a plurality of terminal devices 60 and an analysis device 3, which is an administrator terminal device provided in, for example, a management room, to each other via a wide area network such as the Internet or a private network such as a LAN (Local Area Network).

[0013] The terminal devices 60 are provided at geographically different positions, for example, at predetermined intervals along the passage through which the user passes. Each terminal device 60 includes a camera device 1 and a speaker device 2. The camera device 1 is, for example, a fixed-point camera device, and images a user passing through a passage or the like within a fixed imaging area. Note that the camera device 1 may be a camera device capable of changing the imaging area. The speaker device 2 is an example of a voice output device and outputs guidance voice.

[0014] The analysis device 3 analyzes the captured images of the camera devices 1 of the respective terminal devices 60 and analyzes the characteristics of visually impaired users. Then, the analysis device 3 assigns different voices to each visually impaired user and outputs guidance voice from the speaker device 2 provided at the position where the user moves. Thereby, even when there are a plurality of visually impaired users in adjacent positions, voice guidance can be provided for each user without confusion.

[0015] (Hardware Configuration of Analysis Device) FIG. 2 is a block diagram showing the hardware configuration of the analysis device 3. As shown in FIG. 2, the analysis device 3 includes a CPU (Central Processing Unit) 11, a ROM (Read Only Memory) 12, a RAM (Random Access Memory) 13, and a communication unit 14. The analysis device 3 also includes an HDD (Hard Disk Drive) 15, an input / output interface (input / output I / F) 16, and a communication interface (communication I / F) 17.

[0016] The communication unit 14 performs wired communication via a network such as the Internet or a LAN, and also performs wireless communication such as Bluetooth (registered trademark) or Wi-Fi (registered trademark). The HDD 15 stores a voice guidance program for providing voice guidance to users with visual impairments, map data 50, guidance voice data 51, and a user information table 52.

[0017] As shown in FIG. 3, the map data 50 includes facility information and location information including facility names or names such as facilities, tenants, ticket gates, elevator devices, etc. located within the geographical range for which voice guidance is provided (hereinafter, the expression "service area" is also used in the same meaning), and output condition information indicating conditions for providing voice guidance to the facility, which are stored in association with each other.

[0018] The guidance voice data 51 is data for providing voice guidance based on the map data 50, and a plurality of guidance voice data 51 that are different aurally are stored. More specifically, the guidance voice data 51 is stored as segmented voice data that is voice data for each of various words such as, for example, "I", "Taro", "Hanako", "5", "meters", "ahead", "to", "ticket", "gate", "is", "there", "black", "white", "color", "of", "cardigan", "coat", "baseball", "cap", "long", "short", "hair", etc.

[0019] In one example, as shown in FIG. 4, a plurality of guidance voice data 51 that are aurally different are each assigned a guidance voice ID that uniquely represents them, and are stored separately as guidance voice data 51 with a male speaker and guidance voice data 51 with a female speaker. Also, even when the speakers have the same gender, guidance voice data 51 that are easily distinguishable aurally are stored. Further, for each guidance voice data 51 with a male speaker, information indicating the voice frequency, sound pressure, pitch, and speaking speed are each stored. The same applies to the guidance voice data 51 with a female speaker, and guidance voice data 51 that are easily distinguishable aurally are stored. Also, for each guidance voice data 51 with a female speaker, information indicating the voice frequency, sound pressure, pitch, and speaking speed are each stored.

[0020] In the case of the voice guidance system of the embodiment, when there are a plurality of users with visual impairments, using factors such as the gender of the speaker, voice frequency, sound pressure, pitch, and speaking speed as shown in FIG. 4, guidance voice data 51 that are easily distinguishable by each user with a visual impairment are assigned for voice guidance.

[0021] FIG. 5 shows a schematic diagram of the user information table 52. As shown in this FIG. 5, the user information table 52 is a table that stores the correspondence between users with visual impairments and the guidance voice data 51 assigned to each user. Although it will be described in detail later, it stores the correspondence of the user ID, information on the characteristics (personal characteristics) of the user, and the ID (guidance voice ID) of the guidance voice data 51 assigned to the user.

[0022] To the input / output I / F 16, the display unit 18 and the operation unit 19 are connected when necessary. The communication I / F 17 is connected to the network 5 via a network cable when necessary.

[0023] (Functional Configuration of the Analysis Device) FIG. 6 is a functional block diagram of each function realized software-wise by the CPU 11 executing the voice guidance program stored in the HDD 15. As shown in this FIG. 6, by executing the voice guidance program, the CPU 11 functions as a video acquisition unit 21, a map data acquisition unit 22, an image analysis unit 23, an output voice assignment unit 24, a communication control unit 25, a speaker switching unit 26, and an emergency processing unit 27.

[0024] The video acquisition unit 21 acquires a captured image of a user traveling within the geographical range imaged by each camera device 1. The map data acquisition unit 22 acquires map data 50 corresponding to the longitude and latitude of the geographical range imaged by each camera device 1 from the HDD 15. The image analysis unit 23, which is an example of a detection unit, analyzes the user's personal characteristics based on the captured image captured by each camera device 1, and also determines the presence or absence of visual impairment, etc.

[0025] When it is determined that the user has a visual impairment, the image analysis unit 23 determines whether information on characteristics (personal characteristics) that match the analyzed characteristics (personal characteristics) of the user is stored in the user information table 52. If it is not stored in the user information table 52, the image analysis unit 23 issues a user ID that uniquely represents the analyzed user, associates the issued user ID with the information on the characteristics (personal characteristics) of the analyzed user, and stores them in the user information table 52.

[0026] In addition, the image analysis unit 23 detects the current position of the user based on the longitude and latitude corresponding to each coordinate of the captured image captured by the camera device 1. Furthermore, the image analysis unit 23 detects the moving direction of the user from, for example, the difference in the current positions of the same user shown in a series of captured images of several frames. Also, the image analysis unit 23 detects the presence or absence of a "help-seeking movement" such as a visually impaired user raising a white cane about 50 cm above the head or swinging the white cane left and right in front of the user's face.

[0027] The output voice allocation unit 24 is an example of an allocation unit, and allocates the guidance voice data 51 stored in the HDD 15 to a visually impaired user detected by the image analysis unit 23. Specifically, when a new user is registered by the image analysis unit 23, the output voice allocation unit 24 allocates another guidance voice data 51 with a voice quality different from the already allocated guidance voice among the guidance voice data 51 stored in the HDD 15. Then, the output voice allocation unit 24 stores the ID (guidance voice ID) of the guidance voice data 51 in the user information table 52 as the guidance voice for the new user. Furthermore, the output voice allocation unit 24 refers to the user information table 52 and generates a guidance voice for each user based on the guidance voice data 51 with different voice qualities allocated to each user.

[0028] The speaker switching unit 26 is an example of an output control unit, and switches and controls the speaker device 2 that outputs the guidance voice according to the current position of a visually impaired user, and outputs the guidance voice generated from the guidance voice data 51 with the voice quality allocated to that user. Thereby, voice guidance is performed following the movement of each user with the guidance voice of the voice quality allocated to each user. The communication control unit 25 communicates with each terminal device 60, and performs acquisition of the captured image captured by the camera device 1 and transmission of the guidance voice to the speaker device 2, etc. The emergency processing unit 27, when an operation in which a visually impaired user or the like asks for help is analyzed in the image analysis unit 23, makes an emergency notification to an administrator or the like based on this analysis result. In addition, the emergency processing unit 27 performs output control of a message indicating that a staff member is going to provide emergency assistance, etc. via the speaker device 2 corresponding to the position of the user asking for help.

[0029] In this example, the video acquisition unit 21 to the emergency processing unit 27 are realized by software according to the voice guidance program. However, all or part of these may be realized by hardware such as an IC (Integrated Circuit).

[0030] In addition, the voice guidance program may be provided by being recorded on a recording medium readable by a computer device such as a CD-ROM or a flexible disk (FD) in the form of file information in an installable or executable form. Further, the voice guidance program may be provided by being recorded on a recording medium readable by a computer device such as a CD-R, a DVD (Digital Versatile Disc), a Blu-ray (registered trademark) disc, or a semiconductor memory. Further, the voice guidance program may be provided in a form that is installed via a network such as the Internet. Further, the voice guidance program may be provided by being pre-embedded in a ROM or the like in the device.

[0031] (Voice guidance operation) FIGS. 7 and 8 are flowcharts showing the flow of the voice guidance operation in the voice guidance system of the embodiment. Among these, FIG. 7 is a flowchart showing the first half of the flow of the voice guidance operation. Further, FIG. 8 is a flowchart showing the second half of the flow of the voice guidance operation.

[0032] (Step S1) First, in the flowchart of FIG. 7, in step S1, the video acquisition unit 21 acquires the captured image of the user (pedestrian) captured by each camera device 1.

[0033] (Step S2) After step S1, in step S2, the image analysis unit 23 analyzes the characteristics (person characteristics), current position, and moving direction of the user from the captured image acquired in step S1.

[0034] As the characteristics of the user, the image analysis unit 23 detects the age and gender of the user using a predetermined algorithm. Further, the image analysis unit 23 analyzes the captured image to detect characteristics such as the user's clothing, the color of the clothing, the items carried such as a handbag or a backpack, and the color of the items carried. These characteristics are called person characteristics.

[0035] In addition, the image analysis unit 23 detects the current position of the user based on the longitude and latitude corresponding to each coordinate of the captured image captured by the camera device 1. Further, the image analysis unit 23 detects the moving direction of the user from the difference in the current positions of the same user shown in a series of captured images of, for example, several frames.

[0036] (Step S3) After step S2, in step S3, the image analysis unit 23 determines whether the user shown in the captured image is a healthy person or a user with a visual impairment.

[0037] Although it is an example, a user with a visual impairment owns a white cane for the blind, a white cane. On the other hand, an elderly person who has no visual impairment but has difficulty walking uses a cane such as brown or black. For this reason, the image analysis unit 23 determines whether the user shown in the captured image is a user with a visual impairment based on whether the color of the cane owned by the user is white.

[0038] In addition, a user with a visual impairment has a unique movement of lightly tapping the ground or the like with a white cane while walking in order to check for obstacles or the like. The image analysis unit 23 also uses the presence or absence of such a unique movement as a factor for determining whether the user is a user with a visual impairment.

[0039] In addition, a user with a visual impairment may be accompanied by a guide dog. An ordinary dog has a single string-like lead attached to its collar or harness. In contrast, in the case of a guide dog, a harness with a unique shape called a "U-shaped harness" or a "bar handle type harness" is worn. The image analysis unit 23 also uses the shape of such a harness as a factor for determining whether the user is a user with a visual impairment.

[0040] Also, the "U-shaped harness" or "bar handle type harness" is often white. Therefore, the image analysis unit 23 also uses the color of the harness worn by the dog as a factor for determining whether the user leading the dog is a visually impaired user.

[0041] In addition, for a place where entry with a dog is restricted, if a person enters with a dog, it is likely that the dog is a guide dog and the user is a visually impaired user. Therefore, for a place where entry with a dog is restricted, the image analysis unit 23 determines the user leading the dog as a visually impaired user.

[0042] Also, guide dogs often walk on the left side with respect to the direction of the user's movement. Therefore, the image analysis unit 23 also uses whether the dog being led is walking on the left side with respect to the direction of the user's movement as a factor for determining whether the user is a visually impaired user.

[0043] Furthermore, the image analysis unit 23 also uses factors such as whether the user is wearing sunglasses and whether there is a disabled person mark to determine whether the user is a visually impaired user.

[0044] (Step S3: No → Step S1) Next, when the users shown in the captured image are only healthy people (Step S3: No), the process returns to Step S1, and the processes of Steps S1 to S3 are repeatedly performed by the image analysis unit 23. On the other hand, when it is determined that the users shown in the captured image are visually impaired users (Step S3: Yes), the process proceeds to Step S4.

[0045] (Step S4) In step S4, the image analysis unit 23 refers to the user information table 52 based on the characteristics (person characteristics) of the user determined to have a visual impairment obtained by analyzing the captured image in step S2, and determines whether the user is a registered user based on whether information with the same characteristics (person characteristics) is stored in association with the user ID and the guidance voice ID.

[0046] (Step S4: No → Go to step S14) If the user is not a known user registered in the user information table 52 (step S4: No), the process proceeds to step S14.

[0047] (Step S14: Newly assign a guidance voice ID) In step S14, for the user determined not to be a known user in step S4, the image analysis unit 23 newly issues a user ID. The image analysis unit 23 stores in the user information table 52 the issued user ID in association with the information on the characteristics (person characteristics) of the analyzed user. At the same time, the output voice assignment unit 24 assigns guidance voice data 51 with a voice quality that is not currently assigned to other users. Then, the output voice assignment unit 24 stores in the user information table 52 the ID (guidance voice ID) of the guidance voice data 51, and the user ID and characteristics (person characteristics) of the user in association. That is, the output voice assignment unit 24 assigns guidance voice data 51 with an audible difference as the guidance voice data 51 for each user.

[0048] (Detailed description of step S14) Auditory differences are caused by varying one or more of gender, voice frequency (e.g., formant frequency), sound pressure, pitch, speech rate, etc. Specifically, for example, when there are two visually impaired users, the output voice assignment unit 24 assigns the guidance voice data 51 of a male voice to one user and the voice data of a female voice to the other user. Or, the output voice assignment unit 24 assigns, to one user, the guidance voice data 51 with a high voice frequency and a fast speech rate among the guidance voice data 51 of a male voice, and assigns, to the other user, the guidance voice data 51 with a low voice frequency and a slow speech rate even if it is the same male voice.

[0049] By thus assigning guidance voice data 51 with different voice qualities to each user based on gender, voice frequency, sound pressure, pitch, and speech rate, it is possible to make it easier for each user to recognize in advance the voice quality of the guidance voice for themselves.

[0050] (Step S14 → Step S15; Generation of guidance voice (for pre-recognition)) After step S14, in step S15, the output voice assignment unit 24 generates a guidance voice with the content of the "voice guidance for pre-recognition" described later using the guidance voice data 51 corresponding to the guidance voice ID assigned in step S13. At this time, the output voice assignment unit 24 generates a guidance voice including a proper noun (e.g., "Taro" or "Hanako", etc.) indicating the speaker of the guidance voice data 51, which is associated with each guidance voice data 51. By outputting the thus generated guidance voice to the user, it is possible to make the user recognize in advance that voice guidance will be performed for themselves by the guidance voice of a specific speaker.

[0051] (Detailed description of step S15) The output voice assignment unit 24 generates guidance voice data 51 corresponding to the assigned guidance voice ID, which is a guidance voice including a proper noun indicating a speaker who conducts voice guidance, such as "I'm Hanako guiding you", including proper nouns like "Taro" or "Hanako". As a result, the user can be made to recognize in advance that voice guidance will be provided to themselves with a guidance voice in the voice quality of, for example, "Hanako".

[0052] Note that this example was an example of adding a proper noun of a "name" such as "Taro" or "Hanako". In addition to this, a proper noun of "Mr." or "name" may be added, or other proper nouns such as place names, country names, and building names may be added.

[0053] After the processing of step S15 like this, the processing proceeds to step S6.

[0054] (Step S4: Yes) On the other hand, in step S4, when it is a known user who has already been registered in the user information table 52 (step S4: Yes), the processing proceeds to step S5.

[0055] (Step S5: Acquisition of the assigned guidance voice ID) In step S5, the output voice assignment unit 24 acquires the guidance voice ID assigned to that user from the user information table 52. As a result, the processing proceeds to step S6.

[0056] In step S6, the output voice assignment unit 24 generates a guidance voice including voice guidance for facilities and the like corresponding to the user's current position and moving direction, using the guidance voice data 51 corresponding to the guidance voice ID acquired in step S5 or assigned in step S13.

[0057] (Detailed explanation of step S6) Specifically, when the current position of the user is near a store, for example, the output voice allocation unit 24 reads out the guidance voice data 51 corresponding to the allocated guidance voice ID for each of various words such as "right hand", "to", "store", "A", "is", "here" from the HDD 15. Further, the output voice allocation unit 24 combines the read guidance voice data 51 to generate a guidance voice having as its content a voice guidance corresponding to the current position and moving direction of the user, such as "There is store A on the right hand side".

[0058] (Steps S7 to S8: Determination of speaker device, output of guidance voice) After step S6, the process proceeds to step S7, and the speaker switching unit 26 determines the speaker device 2 that outputs the guidance voice based on the current position of each user, or the current position and moving direction. Then, the process proceeds to step S8, and the speaker switching unit 26 performs output control of the guidance voice having as its content the voice guidance of facilities and the like corresponding to the current position and moving direction of the user, which was generated by the output voice allocation unit 24 in step S6, via the speaker device 2 determined in step S7. Alternatively, the speaker switching unit 26 performs output control via the speaker device 2 determined in step S7 on the voice guidance for pre-recognition generated by the output voice allocation unit 24 in step S14 and the guidance voice having as its content the voice guidance of facilities and the like corresponding to the current position and moving direction of the user.

[0059] (Specific example of steps S15 to S8) Here, a specific example will be shown and explained for the process from generating the guidance voice including the "pre-recognition guidance voice" to outputting it from the speaker device 2. For example, when the user characteristics analyzed by the image analysis unit 23 are a woman wearing a black coat, the output voice allocation unit 24 reads out the guidance voice data 51 corresponding to the guidance voice ID assigned to that user for each of various words such as "black", "color", "of", "coat", "wearing", "woman", "side", etc. from the HDD 15. Also, the output voice allocation unit 24 combines the read guidance voice data 51 to generate a guidance voice including a voice guidance (pre-recognition guidance voice) for making the user recognize that it is a voice guidance for the analyzed user, such as "the woman wearing the black coat", and a voice guidance based on the current position and moving direction of that user.

[0060] Then, based on the current position and moving direction of the user, the speaker switching unit 26 determines, for example, the speaker device 2 of the terminal device 60 where the camera device 1 that captured the captured image of the analyzed user is provided, as the speaker device for outputting the guidance voice. The speaker switching unit 26 controls the output of the above-mentioned guidance voice via the determined speaker device 2. Thereby, since the above-mentioned pre-recognition guidance voice can be made audible to the user, the user can be made to recognize in advance that the voice guidance to be output next is the voice guidance for oneself and the voice quality of that guidance voice.

[0061] When the voice guidance for one or a plurality of users is started in this way, the speaker switching unit 26 selects the speaker device 2 corresponding to the current position and moving direction of the user and outputs the guidance voice with the voice quality assigned to that user. In this way, the voice guidance for the user is always performed with the guidance voice of the voice quality assigned first. Therefore, even when there are a plurality of visually impaired users in the vicinity, since the voice guidance for each user is performed with a different voice quality, each visually impaired user can always recognize that the output voice guidance is the voice guidance for oneself and there will be no confusion. Thus, the voice guidance can function effectively for each user.

[0062] (Steps S9 and S10) Next, after step S8, the process proceeds to step S9 of the flowchart in FIG. 8, and the output voice assignment unit 24 determines whether the user has moved outside the service area. Specifically, the output voice assignment unit 24 refers to the last access date and time in the user information table 52, and if it is more than a certain time (for example, 1 hour) before the current time, it determines that the user has moved outside the service area. When it is determined that the user has moved outside the service area (step S9: Yes), the output voice assignment unit 24 deletes the information about the user from the user information table 52 (step S10).

[0063] Alternatively, the image analysis unit 23 analyzes the camera video such as at the facility entrance and exit to determine whether the user has moved outside the facility (step S9). When the movement of the user outside the facility is confirmed by the camera video (step S9: Yes), the output voice assignment unit 24 deletes the information about the user from the user information table 52 (step S10).

[0064] (Specific examples of voice guidance for multiple users) Furthermore, specifically described, FIGS. 9 to 12 are diagrams schematically showing voice guidance performed for visually impaired users A and B. First, as shown in FIG. 9, it is assumed that user A is going straight from the left direction and user B is going straight from the right direction on the first passage of the store. A second passage is provided for the first passage so as to form a so-called T-junction. Along this first passage and the second passage, terminal devices 60a to 60h corresponding to the terminal device 60 shown in FIG. 1 are arranged at predetermined intervals. It is assumed that user A and user B are imaged by the camera device 1 of the terminal device 60a and the camera device 1 of the terminal device 60f, respectively, and it is analyzed that user A has the characteristic of "a woman wearing a black coat" and user B has the characteristic of "a man wearing a gray suit".

[0065] User A is walking in a position close to the speaker device 2 of the terminal device 60a on the first passage, and user B is walking in a position close to the speaker device 2 of the terminal device 60f on the first passage. In this case, the speaker switching unit 26 selects the speaker device 2 of the terminal device 60a as the speaker device for outputting voice guidance to user A, and selects the speaker device 2 of the terminal device 60f as the speaker device for outputting voice guidance to user B. Also, the output voice assignment unit 24 assigns the guidance voice data 51 (guidance voice ID: M1) of Mr. Taro, a male speaker, to user A, and assigns the guidance voice data 51 (guidance voice ID: F1) of Ms. Hanako, a female speaker with a voice quality different from that of Mr. Taro, a male speaker, to user B.

[0066] The speaker switching unit 26 outputs, via the speaker device 2 of the terminal device 60a, a guidance voice including a pre-recognition voice guidance such as, for example, "For the lady wearing a black coat, I, Taro, will guide you.", which is generated from the guidance voice data 51 with the guidance voice ID M1 assigned to user A. As a result, user A can recognize that the voice guidance for himself / herself is given in the voice of Mr. Taro, a male. Note that by performing the pre-recognition voice guidance based on the personal characteristics as described above, user A can be made to further recognize that the voice guidance to be given next is the voice guidance for himself / herself.

[0067] Similarly, the speaker switching unit 26 outputs, via the speaker device 2 of the terminal device 60f, a guidance voice including a pre-recognition voice guidance such as, for example, "For the gentleman wearing a gray suit, I, Hanako, will guide you. Ahead, there is store B on the left.", which is generated from the guidance voice data 51 with the guidance voice ID F1 assigned to user B. As a result, user B can recognize that the voice guidance for himself / herself is given in the voice of Ms. Hanako, a female. Note that by performing the pre-recognition voice guidance based on the personal characteristics as described above, user B can be made to further recognize that the voice guidance to be given next is the voice guidance for himself / herself.

[0068] Next, as shown in FIG. 10, assume that user A and user B, who are each moving straight ahead, have advanced to a position near the second passage. In this case, the speaker switching unit 26 selects the speaker device 2 of the terminal device 60b as the speaker device 2 that outputs the guidance voice for user A, and selects the speaker device 2 of the terminal device 60d as the speaker device 2 that outputs the guidance voice for user B.

[0069] Then, the speaker switching unit 26 outputs, via the speaker device 2 of the terminal device 60b, a guidance voice such as "Ahead, there is a T-junction. Turn right to store A and go straight to store B." generated from the guidance voice data 51 with the guidance voice ID M1 assigned to user A. Also, the speaker switching unit 26 outputs, via the speaker device 2 of the terminal device 60d, a guidance voice such as "Ahead, there is a T-junction. Turn left to store A." generated from the guidance voice data 51 with the guidance voice ID F1 assigned to user B.

[0070] Next, as shown in FIG. 11, assume that user A and user B have almost simultaneously approached the T-junction. In this case, the selected speaker device is the speaker device 2 of the same terminal device 60c. Then, the speaker switching unit 26 outputs, via the speaker device 2 of the terminal device 60c, a guidance voice such as "There is a T-junction. Turn right to store A and go straight to store B." generated from the guidance voice data 51 with the guidance voice ID M1 assigned to user A. Also, the speaker switching unit 26 outputs, via the speaker device 2 of the terminal device 60c, a guidance voice such as "There is a T-junction. Turn left to store A." generated from the guidance voice data 51 with the guidance voice ID F1 assigned to user B.

[0071] In the example of FIG. 11, although the positions of users A and B are close to each other, each of users A and B has previously recognized the voice quality of the guidance voice for themselves. Also, the guidance voice ID used for the voice guidance for user A is the guidance voice by the voice of M1, and the guidance voice ID used for the voice guidance for user B is the guidance voice by the voice of F1. Since the voice qualities are different, users A and B can distinguish between the guidance voice for themselves and the guidance voice for the other user without confusion. Thus, even when outputting the voice guidance for each of users A and B almost simultaneously via the same speaker device 2, each of users A and B can distinguish the guidance voices with different voice qualities and act according to the voice guidance for themselves. Therefore, the voice guidance for each of users A and B can function effectively.

[0072] Furthermore, as shown in FIG. 12, when user A moves to a position close to the speaker device 2 of the terminal device 60e by going straight along the first passage, the speaker switching unit 26 outputs, via the speaker device 2 of the terminal device 60f, a guidance voice such as "You will soon arrive at store B. Store B is on the right side." generated from the guidance voice data 51 with the guidance voice ID M1 assigned to user A. Thereby, user A can recognize that he / she has moved to store B.

[0073] Also, when user B who has entered the second passage moves to a position close to the speaker device 2 of the terminal device 60h, the speaker device 2 of the terminal device 60h outputs a guidance voice such as "You will soon arrive at store A. Store A is on the left side." generated from the guidance voice data 51 with the guidance voice ID F1 assigned to user B. Thereby, user B can recognize that he / she has moved near store A.

[0074] By outputting the guidance voices with different voice qualities assigned to each user while switching the speaker device 2 according to the movement of the user in this way, it is possible to provide voice guidance for each user without causing confusion.

[0075] (Emergency processing) Next, in step S11 of the flowchart in FIG. 8, the image analysis unit 23 detects whether a visually impaired user makes a "movement seeking help", such as an action of holding a white cane about 50 cm above the head, or an action of swinging the white cane left and right in front of the user's face. If such a "movement seeking help" is not detected (step S11: No), the process returns to step S1.

[0076] On the other hand, if a "movement seeking help" is detected (step S11: Yes), the emergency processing unit 27 issues an emergency notification indicating that the visually impaired user is seeking help, for example, via the display unit 18 (step S12).

[0077] At the same time, in step S13, the speaker switching unit 26 provides an audio guide, such as "An emergency notification has been sent to the administrator. Help will arrive soon, so please wait a moment.", via the speaker device 2 corresponding to the current position of the user seeking help. That is, the speaker switching unit 26 provides an audio guide indicating that the administrator has been contacted in response to the help request and an audio guide asking for a short wait, in the guide audio of the voice quality assigned to the user. This allows the visually impaired user who has requested help to recognize that the administrator or others are taking action in response to their request for help, and can give a sense of security. Also, when this emergency notification is received, an assistant such as an administrator or a security guard can go directly to the position of the user seeking help to provide assistance.

[0078] (Effects of the embodiment) As is clear from the above description, when multiple visually impaired users are present in a proximate location, the voice guidance system according to the embodiment assigns different voice quality guidance voice data 51 to each user to generate guidance voice. Then, the guidance voice of the assigned voice quality is output via the speaker device 2 corresponding to the moving position of each user. Thereby, even when multiple visually impaired users are present in the same place, each user can easily distinguish the guidance voice for themselves, and the voice guidance can function effectively for each user.

[0079] Also, by analyzing the personal characteristics and performing voice guidance (voice guidance for pre-recognition) that enables the user to recognize that it is the voice for themselves with the assigned voice quality, it is possible to make each user more easily recognize the guidance voice for themselves separately from others.

[0080] Also, by including a proper noun indicating the speaker who performs the voice guidance, such as "Taro" or "Hanako", in the voice guidance and outputting it, it is possible to make each user more conscious of the voice guidance for themselves.

[0081] Also, in order to switch the speaker device 2 that outputs the guidance voice according to the movement of the user, by constantly outputting the guidance voice from the same speaker device 2, it is possible to prevent the inconvenience of the voice guidance becoming noise for normal people, store clerks in neighboring stores, neighboring residents, etc.

[0082] Note that in the example of the above-described embodiment, by using the user information table 52 in which the personal characteristics of each visually impaired user (pedestrian) are registered, each visually impaired user (pedestrian) is uniquely identified. However, it is not limited to this, and it may be as follows.

[0083] For example, a visually impaired user (pedestrian) is made to carry a wireless tag such as a BLE tag that transmits radio waves containing its own identification information. BLE is an abbreviation for "Bluetooth (registered trademark) Low Energy". Further, a receiving device for radio waves containing the identification information transmitted by the wireless tag is provided in a terminal device 60 together with, for example, a camera device 1 and a speaker device 2.

[0084] The receiving device receives radio waves from the wireless tag and transmits the identification information contained in the radio waves to an analysis device 3 via a network 5. The analysis device 3 associates the received identification information with an image of a user detected by analyzing a captured image captured by the camera device 1 provided in the terminal device 60 together with the receiving device that has received the identification information, and registers it in a database. As a result, each visually impaired user (pedestrian) can be uniquely identified in the same manner as described above.

[0085] Finally, the above-described embodiments are presented as examples and are not intended to limit the scope of the present invention. This novel embodiment can be implemented in various other forms, and various omissions, replacements, and changes can be made without departing from the gist of the invention. Further, the embodiments and modifications of the embodiments are included in the scope and gist of the invention, and are included in the invention described in the claims and the equivalent scope thereof.

Explanation of Reference Numerals

[0086] 1 Camera device 2 Speaker device 3 Analysis device 5 Network 11 CPU 12 ROM 13 RAM 14 Communication unit 15 HDD 16 Input / output interface (input / output I / F) 17 Communication interface (communication I / F) 18 Display unit 19 Operation unit 21 Video acquisition unit 22 Map data acquisition unit 23 Image analysis unit 24 Output voice assignment unit 25 Communication control unit 26 Speaker switching unit 27 Emergency processing unit 50 Map data 51 Guidance voice data 52 User information table

Claims

1. A detection unit that detects a visually impaired user and at least the current position of the visually impaired user by analyzing a captured image captured by a camera device; An assignment unit that generates a guidance voice based on different voice quality guidance voice data assigned to each user when a plurality of visually impaired users are detected by the detection unit; An output control unit that controls the output of the guidance voice generated based on the different voice quality guidance voice data assigned to each user via a voice output device corresponding to at least the current position of each visually impaired user detected by the detection unit; A voice guidance device comprising the above.

2. The assignment unit assigns guidance voice data in which at least one or more of gender, voice frequency, sound pressure, pitch, and speaking speed are different to each visually impaired user. The voice guidance device according to claim 1, characterized in that.

3. The detection unit detects the characteristics of each visually impaired user respectively; The output control unit performs a pre-recognition voice guidance indicating the characteristics of each detected user using the guidance voice data of the voice quality assigned to each user. The voice guidance device according to claim 1 or claim 2, characterized in that.

4. When generating the guidance voice data, the assignment unit generates the guidance voice data including a proper noun indicating a speaker who conducts the voice guidance. The voice guidance device according to any one of claims 1 to 3, characterized in that.

5. A detection step in which the detection unit detects a visually impaired user and at least the current position of the visually impaired user by analyzing a captured image captured by a camera device; An assignment step in which when a plurality of visually impaired users are detected in the detection step, the assignment unit generates a guidance voice based on different voice quality guidance voice data assigned to each user; An output control step in which the output control unit controls the output of the guidance voice generated based on the different voice quality guidance voice data assigned to each user via a voice output device corresponding to at least the current position of each visually impaired user detected in the detection step; A voice guidance method comprising the above.

6. A computer, A detection unit that detects a visually impaired user and at least the current position of the visually impaired user by analyzing a captured image captured by a camera device; An assignment unit that generates a guidance voice based on different voice quality guidance voice data assigned to each user when a plurality of visually impaired users are detected by the detection unit; Functioning as an output control unit that controls the output of the guidance voice generated based on different voice quality guidance voice data assigned to each user via a voice output device corresponding to at least the current position of each visually impaired user detected by the detection unit; A voice guidance program characterized by the above.

Citation Information

Patent Citations

  • Leading guide system

    JP2002336293A

  • Voice guidance system for visually impaired persons

    JP2020125907A

  • Social Networking with Assistive Technology Device

    US20180053498A1

  • Spatialized verbalization of visual scenes

    US20190333496A1