Method for providing a voice dialogue in sign language in a voice dialogue system for a vehicle
Patent Information
- Application Number
- DE502020011331
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-03-25
- Filing Date
- 2020-03-11
- Publication Date
- 2025-07-17
- Estimated Expiration
- 2040-03-11
AI Technical Summary
Existing in-vehicle voice dialogue systems are inadequate for individuals with speech difficulties or deaf and hard-of-hearing individuals, limiting their ability to effectively communicate with vehicles through spoken language.
A method and system that utilizes an optical capture of input information, such as gestures and facial expressions, to facilitate a non-acoustic speech dialogue in sign language, enabling natural language interaction with the vehicle using image analysis and machine learning to translate and output responses in sign language.
Enables reliable communication with vehicles for individuals who find spoken language challenging, providing a broad spectrum of information through visual outputs in sign language, enhancing accessibility and usability for diverse user groups.
Description
[0001] The present invention relates to a method for providing a speech dialogue in sign language in a speech dialogue system for a vehicle. The invention further relates to a system and a vehicle.
[0002] In-vehicle voice dialogue systems enable the vehicle to be operated and information exchanged with the vehicle through spoken communication. In times of automated driving functions, this can create an additional driving experience by allowing the driver to communicate with the vehicle for information exchange and / or entertainment. Gesture control of in-vehicle instruments is also known. However, engaging in dialogue with the vehicle is still problematic for people with speech difficulties, or for deaf and hard-of-hearing people.
[0003] US 2005 / 0271279 A1 discloses an interaction with machines via gestures in order to provide instructions to the machine.
[0004] US 2013 / 0261871 A1 also discloses a gesture-based control.
[0005] US 2016 / 0107654 A1 discloses an activation method for user inputs.
[0006] Activation methods for user inputs are also known from the documents DE 11 2016 006 769 T5, JP 4 797 858 B2, CN 107 785 017 A, DE 10 2013 014 877 A1, US 2013 / 261 871 A1, CN 106 789 593 B and DE 11 2015 003 655 T5.
[0007] It is therefore an object of the present invention to at least partially remedy the disadvantages described above. In particular, it is an object of the present invention to provide a solution for a non-acoustic speech dialogue system.
[0008] The above object is achieved by a method having the features of claim 1, a system having the features of claim 3 and by a vehicle having the features of claim 6.
[0009] Further features and details of the invention emerge from the respective subclaims, the description, and the drawings. Features and details described in connection with the method according to the invention naturally also apply in connection with the system according to the invention and the vehicle according to the invention, and vice versa, so that with regard to the disclosure of the individual aspects of the invention, reference is always made to each other.
[0010] The object is achieved, in particular, by a method for providing a speech dialogue in sign language in a speech dialogue system, in particular a non-acoustic speech dialogue system, for a vehicle. In particular, the following steps are performed, preferably consecutively in the specified order or in any desired sequence, whereby individual steps can also be performed repeatedly if necessary: Carrying out an optical capture of input information of a vehicle occupant, in particular by an image capture arrangement, carrying out an evaluation of the captured input information, in particular by a processing arrangement, providing a visual output in sign language depending on the evaluation, in particular by an output arrangement.
[0011] In this way, a non-acoustic speech dialogue can be provided using the speech dialogue system, enabling natural language dialogues with the vehicle. The natural language dialogue is formed, in particular, by the input information (e.g., in the form of a "speech," question, answer, or instruction) and the related output (e.g., in the form of a "counter-speech," answer to the question, or feedback to the instruction). A particular advantage here is that the output is provided in sign language, enabling the dialogue even for people who find it difficult to use spoken language for dialogue. It may therefore be possible for the dialogue to be conducted entirely without spoken language.
[0012] It may also be possible for the input information to include information about the vehicle occupant's sign language. This allows the dialogue provided by the speech dialogue system to be provided entirely in sign language. The input information includes, in particular, information about gestures and / or facial expressions and / or silently spoken words and / or body posture of the vehicle occupant, which can be analyzed by a processing arrangement. In order to capture the input information with this information, an image of the vehicle occupant can be recorded, for example, by an image capture arrangement. The analysis may further include image analysis and / or machine learning methods, for example to classify the input information with regard to predefined gestures and / or characters.In this way, the voice dialogue system can analyze the situation in the vehicle interior and use it in the dialogue.
[0013] Furthermore, it is intended that the evaluation will include the following steps: Carrying out a detection of an interaction of the vehicle occupant with the speech dialogue system in sign language on the basis of the input information, wherein the input information is preferably embodied as an image recording of the vehicle occupant in order to detect the interaction at least as a gesture of the vehicle occupant on the basis of the image recording, Carrying out an analysis, in particular image analysis, of the input information to translate the sign language into content translation information that can be evaluated by the speech dialogue system, preferably text information, in particular only upon successful detection of the interaction, Carrying out a detection of a further gesture of the vehicle occupant on the basis of the input information, which is specific to additional information, Generating a response on the basis of the translation information and on the basis of the additional information in order to output the response in sign language via the visual output.
[0014] This enables reliable evaluation of the input information to provide the dialogue. Generating the response may further include, for example, a content analysis of the translation information to determine, for example, a piece of content information that is related to the input information (e.g., as a question) as the response. The output can particularly advantageously be provided in the form of a graphical avatar and / or an abstract form and / or characters (i.e., character-based), and possibly also as audio output.
[0015] It may be possible for at least one gesture by the vehicle occupant to be detected during the evaluation of the input information. This gesture by the vehicle occupant can be detected as a gesture within the context of sign language and can accordingly be used to translate the sign language into the translation information. The gesture can also be detected as a further gesture that is specific to additional information. A further gesture could, for example, be pointing to an object, which accordingly provides the selected object as the additional information. Based on the further gesture detected, the additional information can be used in addition to the translation during the evaluation to generate the answer. For example, the translation information indicates a specific question concerning the object. In other words, the translation information can indicate the function, and the additional information can indicate a parameter for this function.The answer can then be generated as the output of the function.
[0016] The further gesture and the additional information are determined, for example, by the image capture arrangement, in particular a first image capture unit such as a camera, capturing the input information from which the further gesture can be derived (such as pointing at the object). The additional information (i.e., the specific object) can be determined, for example, by querying (further capture) at least one further image capture arrangement and / or at least one further image capture unit of the vehicle, which captures the surroundings of the vehicle and thus the object. For this purpose, directional information is derived, for example, from the further gesture (in which direction is the pointing?) and compared with the further capture of the object.It is also conceivable that the object is identified on the basis of the direction information by evaluating location information about a current location of the vehicle and, in particular, comparing it with environmental information and the direction information.
[0017] According to the invention, it is further provided that the response comprises at least one of the following information (with the following contents): object information about an object in the surroundings of the vehicle, wherein the additional information provides a selection of the object, preferably as a direction indication to the object, weather information, in particular about the weather at the location of the vehicle, environmental information, preferably information about a current location of the vehicle, news information, in particular about local news at the location of the vehicle, media information, such as videos or literature or the like, wherein, particularly preferably, to generate the response, the information is filtered with regard to the availability of an output of the information in sign language. This enables the provision of a broad spectrum of information for the speech dialogue. The further gesture according to the additional information is designed to select the object, e.g., as a finger gesture or the like, with which the object can be pointed at. According to the filtering, for example, only content suitable for output in sign language can be considered. It may also be possible for information in text form to be translated into sign language by a method and / or system according to the invention, e.g., by means of the processing arrangement. For this purpose, the processing arrangement can optionally also communicate with an external data processing system in order to carry out this translation, e.g., via an external server or via cloud computing.
[0018] The invention also relates to a system for a vehicle for providing a voice dialogue in sign language. The system can be implemented, for example, as vehicle electronics and thus be suitable for permanent integration into the vehicle. The system comprises at least one of the following components: an image capture arrangement, in particular a camera system, for capturing input information from a vehicle occupant, wherein the input information is preferably embodied as image information, in particular image recording, a (in particular electronic) processing arrangement for evaluating the captured input information, wherein the processing arrangement is embodied, for example, as a control unit and / or electronics and / or a microcontroller and / or the like, an output arrangement for visual output in sign language depending on the evaluation.
[0019] The system according to the invention thus provides the same advantages as those described in detail with reference to a method according to the invention. Furthermore, the system can be suitable for implementing a method according to the invention. It is further conceivable for the system to be designed as a voice dialog system of a method according to the invention in order to implement the steps of the method.
[0020] The system can, for example, be implemented as a technical component in the vehicle that monitors the vehicle interior using an image capture system, particularly in the form of a camera system, and is capable of detecting when a vehicle occupant (passenger) interacts with the system, particularly a voice dialogue system (voice assistant), and translates the sign language into text. This allows the voice assistant to behave as if the user were speaking to the voice assistant. The image capture system is installed, for example, in the front of the dashboard or the rearview mirror, thus providing precise user monitoring.
[0021] Advantageously, the invention can provide for the image capture arrangement to be configured as an infrared camera to optically capture at least one gesture of the vehicle occupant. The camera can, for example, be configured for capture in the near and / or far infrared range in order to capture the input information particularly reliably and with minimal interference.
[0022] Optionally, the output arrangement may include at least one display, in particular an LED display, for the vehicle interior to output the visual output as a visualization of a response to the input information in sign language. Preferably, the output arrangement is configured as an instrument cluster for the vehicle. This enables convenient recognition of the sign language output.
[0023] The invention also relates to a vehicle with a system according to the invention. Within the scope of the invention, it can be advantageous for the vehicle to be designed as an autonomous vehicle, in particular as a self-driving motor vehicle. Such a vehicle can drive, steer, and park without the influence of a human driver. In this case, the use of the voice dialogue system is particularly useful, since none of the vehicle occupants is involved in controlling the vehicle.
[0024] It is also advantageous if the vehicle is designed as a motor vehicle, in particular a trackless land vehicle, for example as a hybrid vehicle comprising an internal combustion engine and an electric motor for traction, or as an electric vehicle, preferably with a high-voltage electrical system and / or an electric motor. In particular, the vehicle can be designed as a fuel cell vehicle and / or a passenger vehicle. In embodiments of electric vehicles, no internal combustion engine is preferably provided in the vehicle; it is then powered exclusively by electrical energy.
[0025] According to an advantageous development of the invention, it can be provided that the output arrangement is arranged in the vehicle interior, in particular in the area of a center console, and the image capture arrangement is arranged in the area of an instrument cluster and / or the center console and / or an interior and / or rearview and / or exterior mirror of the vehicle. In particular, the image capture arrangement can be arranged on the vehicle in such a way that the driver of the vehicle can be detected. In this case, it is expedient, for example, to arrange the image capture arrangement on the rearview and / or exterior mirror on the side of the vehicle on which the driver of the vehicle is sitting. Furthermore, it is also conceivable that the detection by the image capture arrangement is extended to include other occupants of the vehicle. Accordingly, alternative arrangements in or on the vehicle and / or the use of further image capture arrangements are also suitable.
[0026] Further advantages, features, and details of the invention will become apparent from the following description, which describes exemplary embodiments of the invention in detail with reference to the drawings. The features mentioned in the claims and in the description may be essential to the invention individually or in any combination. They show: Figure 1 shows a schematic representation of parts of a system according to the invention from the perspective of an interior of the vehicle, Figure 2 shows a schematic representation for visualizing a method according to the invention.
[0027] In the following figures, identical reference numerals are used for the same technical features, even in different embodiments.
[0028] In Figure 1A system 100 for a vehicle 1 for providing a voice dialogue in sign language is shown schematically. The illustration is from the perspective of a vehicle occupant 2 in a vehicle interior 40. An image capture arrangement 110 is shown, which can be used to capture input information from the vehicle occupant 2. For this purpose, the image capture arrangement 110 optically monitors the vehicle interior 40 in order to visually record the vehicle occupant 2. In addition, other vehicle occupants 2 can also be captured by the image capture arrangement 110 if necessary.
[0029] Furthermore, a processing arrangement 120 is provided for evaluating the acquired input information, wherein the processing arrangement 120 is embodied, for example, as vehicle electronics. It may also be possible for the evaluation to be performed partially by external components, such as an external data processing system. For this purpose, the processing arrangement communicates with the external data processing system, for example, via a network such as a mobile network and / or the Internet. In this way, additional resources can be used for the evaluation.
[0030] In addition, an output arrangement 130 for visual output in sign language can be provided depending on the evaluation. Typically, in a speech dialogue system, both the input by the user and the response by the system 100 (and / or vice versa) are provided via acoustic speech. According to the invention, at least the response by the system 100 can be provided in sign language. This thus enables communication with the vehicle 1 even when communication via spoken language is not possible.
[0031] The image capture arrangement 110 can be configured as a camera, in particular an infrared camera, to optically capture at least one gesture of the vehicle occupant 2 using the input information. The gesture is, for example, part of the sign language. Furthermore, the input information can also include at least one piece of information about facial expressions and / or silently spoken words and / or body posture of the vehicle occupant, which can also be part of the sign language. This makes it possible to digitally capture the sign language using this information.
[0032] Furthermore, the gesture detected via the input information can also be another gesture that is specific to additional information. This allows the additional information to be detected in addition to the information via sign language. The additional gesture is, for example, pointing at an object 5 (shown is pointing using a finger of the vehicle occupant 2, which provides the directional information). The additional information is then the information about which object 5 was pointed at. This additional information can, for example, also be determined by evaluating the input information, in particular if it also includes information about the object 5 or the exterior of the vehicle 1, or by evaluating the detection of at least one further image capture device (e.g., the exterior of the vehicle 1).This makes it possible to evaluate a command or question conveyed via sign language and compare it with the additional information. The response can then be determined, for example, as information about the object 5 pointed to.
[0033] The output arrangement 130 can comprise at least one display 132, in particular an LED display, for the vehicle interior 40 in order to output the visual output 131 as a visualization of a response to the input information in the sign language, wherein the output arrangement 130 is preferably designed as an instrument cluster 10 for the vehicle 1.
[0034] In addition, the output arrangement 130 can be arranged in the vehicle interior 40, in particular in the region of a center console 30, and the image capture arrangement 110 can be arranged in the region of an instrument cluster 10 and / or the center console 30 and / or a rearview mirror 21 and / or exterior mirror 20 of the vehicle 1.
[0035] In Figure 2 A method according to the invention for providing a speech dialogue in sign language in a speech dialogue system for a vehicle 1 is visualized. According to a first method step 101, an optical capture of input information from a vehicle occupant 2 can be carried out. According to a second method step 102, an evaluation of the captured input information is carried out. Subsequently, according to a third method step 103, a visual output 131 in sign language can be provided depending on the evaluation.
[0036] The above explanation of the embodiments describes the present invention exclusively within the scope of examples. Of course, individual features of the embodiments can be freely combined with one another, provided they are technically feasible, without departing from the scope of the present invention. List of reference symbols
[0037] 1Vehicle, motor vehicle 2Vehicle occupant 3Environment 5Object 10Instrument cluster 20Exterior mirrors 21Rearview mirrors 30Center console 40Vehicle interior 100 101-System 103Procedure steps 110Image acquisition arrangement 120Processing arrangement 130Output arrangement 131Visual output 132Display
Claims
1. Method for providing a speech dialog in sign language in a speech dialog system for a vehicle (1), wherein the following steps are carried out: - carrying out a visual capture of input information from a vehicle occupant (2), - carrying out an evaluation of the captured input information, - providing a visual output (131) in sign language on the basis of the evaluation, characterized in that carrying out the evaluation comprises the following steps: - carrying out a detection of an interaction of the vehicle occupant (2) with the speech dialog system in sign language based on the input information, wherein the input information is preferably implemented as an image recording of the vehicle occupant (2) in order to detect the interaction as a gesture by the vehicle occupant (2) based on the image recording, - carrying out an analysis of the input information in order to translate the sign language into content-based translation information, preferably textual information, that can be evaluated for the speech dialog system, only if the interaction is successfully detected, - carrying out a detection of a further gesture by the vehicle occupant (2) based on the input information which is specific to a piece of additional information, - generating a response based on the translation information and the additional information in order to output the response in sign language by the visual output (131), wherein the response comprises at least one of the following pieces of information: - object information about an object (5) in an environment (3) of the vehicle (1), wherein, for this purpose, the additional information provides selection of the object (5), - weather information, - environment information, - message information, - media information.
2. Method according to claim 1, characterized in that the response also comprises the following information: - information about a current location of the vehicle (1), wherein, in order to generate the response, the information is filtered with regard to the availability of an output of the information in sign language.
3. System (100) for a vehicle (1) for providing a speech dialog in sign language, comprising: - an image capture arrangement (110) for capturing input information from a vehicle occupant (2), - a processing arrangement (120) for evaluating the captured input information, - an output arrangement (130) for visual output in sign language on the basis of the evaluation, characterized in that the system (100) is designed as a speech dialog system of a method according to either of claims 1 or 2 in order to carry out the steps of the method.
4. System (100) according to claim 3, characterized in that the image capture arrangement (110) is designed as an infrared camera in order to optically capture at least one gesture by the vehicle occupant (2).
5. System (100) according to claim 3 or claim 4, characterized in that the output arrangement (130) comprises at least one display (132), in particular an LED display, for the vehicle interior (40) in order to output the visual output (131) as a visualization of a response to the input information in sign language, wherein the output arrangement (130) is preferably designed as an instrument cluster (10) for the vehicle (1).
6. Vehicle (1) comprising a system (100) according to any of claims 3 to 5.
7. Vehicle (1) according to claim 6, characterized in that the vehicle (1) is designed as an autonomous vehicle (1).
8. Vehicle (1) according to claim 6 or claim 7, characterized in that the output arrangement (130) is arranged in the vehicle interior (40), in particular in the region of a center console (30), and the image capture arrangement (110) is arranged in the region of an instrument cluster (10) and / or the center console (30) and / or a rear view mirror (21) of the vehicle (1).