METHOD FOR VOICE CONTROL, SYSTEM FOR VOICE CONTROL AND VEHICLE WITH A SYSTEM FOR VOICE CONTROL

DE502021010462D1Active Publication Date: 2026-06-03CONTINENTAL AUTOMOTIVE TECHNOLOGIES GMBH

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
CONTINENTAL AUTOMOTIVE TECHNOLOGIES GMBH
Filing Date
2021-02-17
Publication Date
2026-06-03

AI Technical Summary

Technical Problem

Existing voice control systems fail to manage voice commands effectively in environments with multiple users, particularly in vehicles, leading to potential safety risks from conflicting or unauthorized commands, especially for safety-critical functions.

Method used

A method and system for voice control in vehicles that identifies users within a predefined area, assigns access rights based on user identification, processes audio signals to recognize and execute authorized commands, and utilizes a combination of visual and audio cues to manage multiple users, including features like near-infrared cameras and deep learning for user recognition and command analysis.

Benefits of technology

Enhances safety by ensuring only authorized users can execute critical functions, reducing the risk of conflicting commands and enhancing security in vehicles with multiple occupants.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

Field of invention

[0001] The invention relates to a method for voice control, in particular in a vehicle, a system for voice control, in particular in a vehicle, and a vehicle with a system for voice control. background

[0002] Methods and systems for voice control are known from the prior art. For example, entertainment functions in a vehicle or certain functions of a smartphone or computer are controlled by voice input. A special case of voice control arises when multiple users can perform this voice control.

[0003] For example, US patent application US 2019 / 0080692 A1 describes a device that facilitates the simultaneous detection and processing of multiple speeches from multiple users. This requires at least two microphones and beamforming logic.

[0004] Voice control becomes problematic for multiple users when, for example, two or more users make different and even contradictory voice commands. This is particularly significant when dealing with safety-relevant voice commands, such as in a vehicle with autonomous or semi-autonomous driving functions. Conflicting voice commands or those from unauthorized users can lead to dangerous situations. However, this aspect is not addressed in the current state of the art. Summary

[0005] The object of the invention is to overcome the disadvantages of the prior art and, in particular, to manage voice control by multiple users. This object is achieved by the subject matter of the independent claims. Further developments of the invention are described in the dependent claims and the following description. The scope of protection is defined by the appended claims.

[0006] One aspect of the invention relates to a method for voice control. Voice control is understood to mean the control of functions of a device, machine, or apparatus by means of voice input. In particular, the invention relates to a method for voice control in a vehicle. This can be a land, water, and / or air vehicle, especially a passenger car or a bus. The controllable functions can be very diverse. In the case of a vehicle, for example, entertainment functions such as music playback or telephony, but also driving functions such as the speed setting of a cruise control or functions of semi-autonomous or autonomous driving can be controlled.

[0007] Users can be located within a predefined area. This predefined area could be, for example, a single room or multiple rooms. In the case of multiple rooms, these rooms could be adjacent or separated from each other. For voice control in a vehicle, the predefined area would typically be the vehicle interior or several separate passenger compartments.

[0008] The process involves user identification. Some or all users within a defined area can be identified. To do this, user characteristics are recorded and compared with characteristics of registered users stored in a database. A user is then either recognized as a registered user or, if not yet registered in the database, classified as an unknown user. This identification process is repeated for each user within the defined area. User identification is computer-aided, for example, using deep learning, and is performed by a computer unit.

[0009] Each user is then assigned access rights based on their user identification. These access rights include the authorization for users to execute voice commands. Access rights can be granted, for example, for each individual function controllable via voice commands and / or for groups of functions. The access rights specify whether a function or group of functions can be executed by a user, cannot be executed, or can only be executed under predetermined circumstances. Additionally, access rights can include a prioritization of individual users, which can be general for all functions or function-specific. Prioritization can, for example, be a numerical value, with a higher priority given to a user the higher the value. For instance, if two users issue opposing voice commands (e.g.,If a user says "Increase speed" and "Decrease speed," the voice command of the user with the higher priority will be executed. If both users have the same priority, a confirmation question can be asked, or the voice commands can be discarded.

[0010] Furthermore, the method involves capturing audio signals using at least one microphone. A microphone is generally understood as a sound transducer that converts airborne sound into a machine-readable signal, particularly an electrical signal. The audio signals are preferably captured after user identification, so that no additional time is spent on user identification after the audio signals have been captured. For example, user identification can be carried out before or at the start of a journey in a vehicle. However, it is also conceivable that, particularly if the number of users is very large and / or the users within the defined area change frequently, user identification is performed only after the audio signals have been captured.

[0011] The captured audio signals are then processed. In this process, the audio signals are assigned to individual users. This assignment is based on the captured audio signals and can be supported by signals from additional detectors, such as cameras. The audio signals are evaluated, for example, with regard to their frequency spectrum. If more than one microphone is used, the different arrival times of the audio signals can be used to determine the user's position within the specified area. Furthermore, a camera—preferably a near-infrared camera with a near-infrared light source, so that it provides usable images even in low light conditions—can be used to capture, for example, the lip movements of individual users. A correlation between the lip movements and the audio signals then assigns the audio signals to the individual users.The aforementioned steps for assigning the audio signals can also be carried out by the computer unit, individually or in combination, using artificial intelligence, for example using a deep learning method.

[0012] Furthermore, the audio signals are converted into machine-readable form. This conversion can be performed by the computer unit, for example, using Natural Language Processing.

[0013] Furthermore, an analysis of the audio signals for voice commands is performed. This analysis can, for example, be carried out based on the audio signals converted into machine-readable form. However, it is also conceivable that the analysis of the audio signals for voice commands is performed in combination with the conversion of the audio signals into machine-readable form. In the latter case, the voice commands are then extracted from the audio signal. This step is also preferably carried out by the computing unit using artificial intelligence.

[0014] Recognized voice commands are then checked against the access rights of the user to whom the voice command was assigned. In other words, it is checked whether the user's access rights permit the execution of the function addressed in the voice command. If the access rights allow the execution of the voice commands, the commands are executed. For this purpose, a signal to execute is sent, for example, to a corresponding controller.

[0015] Access rights manage voice commands from users within a defined space. Only users with the appropriate authorization can execute specific functions via voice command. If these functions are security-relevant, this also increases security, as only authorized users can perform them.

[0016] In some embodiments, user identification is carried out using images, particularly in the visible spectrum and / or near-infrared. Visible light has wavelengths of approximately 380 nm to 780 nm, and near-infrared wavelengths of approximately 780 nm to 3.0 µm. The advantage of visible light is that color information can also be used for user identification. Near-infrared cameras and a corresponding near-infrared light source, on the other hand, allow images to be taken even in darkness without dazzling and disturbing the users. For identification using images, characteristic features, especially of the face, are captured. Alternatively or additionally to user identification using images, identification can also be carried out using audio signals. In this case, for example, the voice tone and / or pronunciation characteristic of each user are analyzed.A combination of the two identification methods provides a more reliable result.

[0017] In some embodiments, each user is assigned a feature vector. This feature vector can include visual features. Visual features are all possible features that can be extracted from one or more images of the user. Such extraction of visual features from one or more images is preferably performed using artificial intelligence, for example, by means of a convolutional neural network in the computing unit. Alternatively or additionally, the feature vector can also include semantic features. Semantic features are defined as those features that can be recognized by a human in the images, such as approximate age, eye color, or interpupillary distance. The semantic features can, for example, be automatically extracted from the visual features via facial recognition.Alternatively or additionally to visual and semantic features, the feature vector can also include auditory features. Auditory features can include, for example, the user's voice pitch, pronunciation, and / or accent. The auditory features can be analyzed by the processing unit using, for example, a short-time Fourier transform and / or artificial intelligence. It is conceivable that the feature vector could also include further features, such as three-dimensional information about the user, obtained, for example, using a stereo camera or by analyzing multiple images of the user in different poses, and / or information about the user's weight, recorded, for example, by pressure sensors within a defined area.

[0018] User identification using feature vectors is performed by determining a distance in multidimensional space between the feature vector assigned to the user and the feature vectors stored in the database. If this distance is, for example, smaller than a predefined or system-adjustable threshold, then the user is recognized as one of the users known from the database.

[0019] In some embodiments, changes in user behavior are detected. A change in behavior could be, for example, a change of position within the predefined area, leaving the predefined area, or entering the predefined area. Such a change can be detected, for example, by a pressure sensor installed in a seat within the predefined area. A decrease in pressure indicates that the user has left their seat, while an increase in pressure indicates that a user has sat down. Alternatively or additionally, a change can be detected by a seatbelt sensor associated with the seatbelt of a seat within the predefined area. Here, too, the unfastening or fastening of the seatbelt indicates that a user has at least changed their position within the predefined area.Furthermore, such a change in user behavior can be detected using a camera in the visible spectrum and / or near-infrared. The images captured by the cameras are then analyzed, for example, using a Gaussian mixing model. The camera can detect both a change in the user's position and a user leaving or entering the designated area. Changes in user behavior can also be detected by analyzing audio signals. This method is particularly effective in identifying newly arrived users whose speech patterns differ from those of users already present in the designated area. If a new user enters the designated area, the user is identified as described above.

[0020] In some implementations, access rights are assigned to a user recognized as a registered user in the database or in a user database. If, however, a user is not recognized, default access rights are assigned to this unknown user.

[0021] In some implementations, the database or user database entries can be configured. In particular, the access rights of individual registered users can be set. Optionally, default access rights can also be set. Such configuration can be done, for example, via a user interface. Individual users can also be removed from the database or user database, or the entire database or user database can be reset. Using a vehicle as an example, children could be granted rights to select the music to be played, adult passengers could have rights to open windows and doors, and the driver could have rights to control driving functions and / or (semi-)autonomous driving. A fleet manager of a rental car fleet could be granted even further rights.

[0022] Certain security requirements must be met to configure entries in the database or user database. For example, configuration of a vehicle might only be possible when the vehicle is stationary to prevent distractions. Furthermore, configuration might only be possible directly in the vehicle, for instance, to prevent attacks by unauthorized individuals such as hackers. Alternatively, or in addition to these security requirements, authorization is also necessary to configure entries. For example, configuration might only be possible after entering a specific security code or if the configuring user possesses a vehicle key.

[0023] In some implementations, background noise is minimized during the processing of the audio signals. This can be achieved, for example, by removing noise. Another possibility is the subtraction of other audio signals that are output via loudspeakers within the specified range. Music played in a vehicle, for example, can thus be easily removed from the audio signals.

[0024] In some implementations, the processing of audio signals is performed at least partially by the computing unit using deep learning. This results in particularly good and robust processing of the audio signals. Alternatively or additionally, the processing of audio signals can be carried out using the context of a conversation between users in the given space. For example, if a user complains about being cold, the subsequent voice command to increase the temperature will be more easily recognized.

[0025] In some implementations, the audio signals are searched for predefined keywords. Such keywords, for example, "Hello, car!", are intended to precede a voice command. The analysis of the audio signals for voice commands is then only performed after a recognized keyword. This prevents, among other things, voice commands accidentally mentioned in a conversation from being recognized and executed as voice commands.

[0026] In some embodiments, keywords can be assigned individually to separate users, meaning each registered user can be assigned their own keyword. When assigning audio signals to users, the keywords are then preferably taken into account in addition to the other attributes. In particular, if the recognized keyword does not match the recognized user, further analysis of the audio signal can be omitted.

[0027] In some implementations, an error message is displayed if the user lacks the necessary access rights for a function to be executed via voice command. This informs the user that while the voice command was recognized, execution of the function was denied due to insufficient access rights. Alternatively, or in addition to displaying the error message, verification by an authorized user can be enabled if access rights are lacking. For example, a user with the appropriate access rights for the requested function can initiate the execution of the previously rejected voice command via a voice command. Furthermore, if an emergency is detected, access rights can be transferred to another user and / or voice commands can be executed even without the necessary access rights.An emergency could be, for example, a user becoming drowsy, inattentive, or experiencing a medical emergency. If, in such a case, no other user in the designated space has the necessary access rights, such as for driving functions of a vehicle, the access rights are transferred to one or more other users in order to, for example, stop the vehicle and thus prevent an accident.

[0028] In some implementations, a token is created for a recognized user. This token includes access rights, a timestamp of its creation, and / or a token lifetime. This eliminates the need to recheck access rights for the recognized user, at least during the token's lifetime. However, if the recognized user leaves the predefined scope, the token can be prematurely deleted. To prevent misuse, the token is preferably encrypted. When instructing the execution of a command recognized as a voice command, as described above, the token is transmitted. A controller tasked with executing the function then decrypts the token and executes the function, provided the access rights contained in the token permit it.

[0029] Furthermore, it is also possible that the token must first be authenticated via a cloud service, with authentication taking place using a wireless transmission method, for example via LTE. A key for encrypting the token is stored in the cloud, and the token is only encrypted upon successful authentication.

[0030] Another aspect of the invention relates to a voice control system. Voice control is understood to mean the control of functions of a device, machine, or apparatus by means of voice input. In particular, the invention relates to a voice control system in a vehicle. This can be a land, water, and / or air vehicle, especially a passenger car or a bus. The controllable functions can be very diverse. In the case of a vehicle, for example, entertainment functions such as music playback or telephony, but also driving functions such as the speed setting of a cruise control or functions of semi-autonomous or autonomous driving can be controlled. The voice control system is configured to carry out the method according to the preceding description.

[0031] The voice control system includes at least one microphone for capturing audio signals. A microphone, in this context, is generally understood to be a sound transducer that converts sound waves into a machine-readable signal, particularly an electrical signal.

[0032] Furthermore, the voice control system comprises at least one computing unit. This computing unit can include various components, which may be separate and / or integrated. For example, the computing unit can include a convolutional neural network, which is used to extract visual features from captured images, perform facial recognition, or separate audio signals for individual users. Additionally, the computing unit can include, for example, a hardware accelerator for artificial intelligence, which is used to convert audio signals into machine-readable form, detect user changes, or identify the user. The hardware accelerator for artificial intelligence is preferably connected to one or more of the sensors, such as the microphone or a camera.The hardware accelerator for artificial intelligence can be designed in such a way that audio signals and / or images are not stored, and only the identification results are transmitted. Furthermore, the computing unit can include a database controller that operates the database and / or the user database. In addition, the computing unit can include a controller for central program control, which, for example, forwards the recognized voice commands for execution, e.g., to a vehicle network.

[0033] In some embodiments, the voice control system further comprises at least one near-infrared camera and one near-infrared light source. The near-infrared light source serves to illuminate the designated area in darkness. The near-infrared camera captures images of the users, which are used to identify the users, to assign the audio signals to individual users, and / or to detect changes in user behavior.

[0034] Another aspect of the invention relates to a vehicle with a voice control system as described above.

[0035] For further clarification, the invention is described with reference to embodiments illustrated in the figures. These embodiments are to be understood as examples only, and not as limitations. Brief description of the characters

[0036] This shows: Fig. 1a schematic top view of an exemplary embodiment of a vehicle, Fig. 2 a flowchart of part of a voice control procedure, Fig. 3 a flowchart of another part of a procedure for voice control and Fig. 4 a flowchart of yet another part of a procedure for voice control. Detailed description of embodiments

[0037] Figure 1Figure 1 shows a schematic top view of an embodiment of a vehicle 1, which is depicted here as a passenger car. The vehicle 1 includes a voice control system 2. Voice control refers to the control of functions by means of voice input, in this case, functions of the vehicle 1. Functions of the vehicle 1 that can be controlled by voice include, for example, entertainment functions such as music playback or telephony, but also driving functions such as the speed setting of a cruise control or functions of semi-autonomous or autonomous driving. The present voice control system 2 is designed such that it can manage several users 3 who are located in a predefined area 4 of the vehicle 1. In this embodiment, the predefined area 4 is the vehicle interior.

[0038] The voice control system 2 includes a microphone 5 for capturing audio signals. Among other things, the microphone 5 captures voice control commands spoken by the users 3. Furthermore, the voice control system 2 includes two near-infrared cameras 6 with corresponding near-infrared light sources, which are not shown here for clarity. The near-infrared light sources illuminate the users 3 in the dark, ensuring that the users 3 are not dazzled or distracted by the near-infrared light and that the near-infrared cameras 6 have sufficient illumination to capture images of the users 3. In the present embodiment, one near-infrared camera 6 is provided for each of the two passenger rows to ensure good image capture of the users 3. However, embodiments with only one near-infrared camera 6 or with more than two near-infrared cameras 6 are also conceivable.The images captured by the near-infrared cameras 6 serve both to identify the users 3 and to assign the audio signals recorded by the microphone 5 to the individual users 3.

[0039] The voice control system 2 may also include additional sensors, such as extra microphones, stereo cameras for three-dimensional scanning, visible light cameras, pressure sensors in the seats, and / or seatbelt sensors. The signals from these sensors can then also be used to identify users 3 and / or to assign audio signals to individual users 3.

[0040] All sensors, in particular the microphone 5 and the near-infrared cameras 6, are connected to a computer unit 7.

[0041] Figure 2This shows a flowchart of part of a voice control process. After startup, user detection (201) is performed based on data captured by the cameras. This user detection (201) is carried out, for example, using artificial intelligence, such as deep learning.

[0042] Using the data obtained during user detection 201 and based on received audio signals, an assignment 202 of audio signals to the individual users 3 is carried out. To improve this assignment 202, additional measures such as minimizing background noise can be implemented.

[0043] A keyword recognition process (203) is performed on the audio signals assigned to users 3, which identifies keywords contained in the audio signals. This keyword recognition process (203) is based, for example, on Natural Language Processing.

[0044] In decision step 204, it is checked whether a keyword has been recognized. If no keyword is recognized, keyword recognition 203 is continued. If, however, a keyword is recognized, speech command processing 205 is started. As soon as speech command processing 205 is completed, keyword recognition 203 is carried out.

[0045] Furthermore, using the data obtained during user detection 201, a change recording 211 is carried out, in which changes of users 3, for example a change of position, leaving vehicle 1 or entering vehicle 1, are recorded.

[0046] In decision step 212, it is checked whether a change has occurred to user 3. If not, change recording 211 continues. If a change has occurred, user identification 213 is performed, based, among other things, on the image data and the audio signal assigned to user 3. User identification 213 also accesses a database 221 in which the characteristics of registered users are stored.

[0047] Figure 3 Figure 205 shows a flowchart of speech command processing. First, the audio signals are converted into a machine-readable form, for example using Natural Language Processing.

[0048] Then, the audio signals generated by the conversion 301 are assigned in machine-readable form to the individual users 3. For this assignment 302, the characteristics of the registered users stored in the database 221 are used.

[0049] The audio signals, in machine-readable form, are then used to perform voice command recognition (303) and analysis of the associated context, preferably using artificial intelligence. This recognition (303) can be further improved by taking into account a history (310) of the most recently spoken voice commands and / or the conversation between users (3).

[0050] Once the recognition 303 of the voice commands is complete, a check 304 is performed to determine whether user 3, to whom the audio signal has been assigned, has the necessary access rights. This check uses database 221, which in this embodiment also manages the access rights. Alternatively, a separate user database could be used, with the user database containing the access rights and database 221 containing the attributes of user 3.

[0051] If user 3 does not possess the necessary rights to execute the function, the voice command processing (205) is terminated. However, if user 3 possesses the necessary rights to execute the function, the voice command is executed (305), and subsequently the voice command processing (205) is terminated.

[0052] In Figure 4A flowchart of an exemplary implementation of a voice control method is shown, in which functions are activated using a token. The analysis of the audio signals is already completed at the beginning of this flowchart. A check (401) is then performed to determine whether a token exists for user 3, to whom an audio signal with a voice command has been assigned. If so, the token is transmitted (407) to the function corresponding to the voice command. The token contains information about user 3's access rights, enabling the function to verify whether the voice command is executed or not.

[0053] If, however, no token exists for user 3, a feature vector for user 3 is first calculated (402). This feature vector is then transmitted (403) to a controller assigned to database 221.

[0054] For the transmitted feature vector, a determination is now made (404) of the distance between the user's feature vector and the other feature vectors of registered users stored in database 221. The feature vector of registered users with the smallest distance to the user's feature vector 3 is selected.

[0055] A check (405) is then performed to determine whether the distance between the feature vector of user 3 and the feature vector of the registered user with the smallest distance to user 3's feature vector is less than a predefined threshold. If so, a new token is generated (406) for user 3. This token is preferably encrypted and / or has a lifetime. The new token is then transmitted (407) to the function corresponding to the voice command. If, however, the distance is greater than the predefined threshold, a token with standard access rights is provided (408), and this token is subsequently transmitted (409) to the corresponding function. In both cases, the function can then check whether sufficient access rights exist to execute the voice command. Reference symbol list

[0056] 1 Vehicle 2 Voice Control System 3 User 4 Predefined Area 5 Microphone 6 Near-infrared Camera 7 Computing Unit 201 User Detection 202 Mapping 203 Keyword Recognition 204 Decision Step 205 Voice Command Processing 211 Change Detection 212 Decision Step 213 User Identification 221 Database 301 Conversion 302 Mapping 303 Recognition 304 Checking 305 Execution 310 History 401 Checking 402 Calculation 403 Transmission 404 Determination 405 Verification 406 Generation 407 Transmission 408 Provisioning 409 Transmission

Claims

1. A voice control method in a vehicle (1), wherein in a predefined region (4), in particular in a vehicle interior, in which users (3) can be located, an identification (213) of users (3) is carried out by acquiring features of the users (3) and comparing the features with features of registered users (3) stored in a database (221), each user (3) being identified as a registered user (3) or being classified as an unknown user (3); access rights are assigned to each user (3) dependent on the identification (213) of the users (3), wherein the access rights comprise the authorization of users (3) to execute voice commands; audio signals are acquired by means of at least one microphone (5); the audio signals are processed, wherein the audio signals are assigned (302) to the individual users (3), the audio signals are converted (301) into machine-readable form, and an analysis of the audio signals for voice commands is carried out; detected voice commands are checked for the access rights of the user (3); and if the access rights allow the voice commands to be executed, the voice commands are ordered to be executed, characterized in that if an emergency is detected, access rights are transferred to another user (3) and / or voice commands are also executed without corresponding access rights.

2. The method as claimed in claim 1, wherein an emergency is a drowsiness, inattention or a medical emergency of a user, and if, in this case, no other user located in the predetermined space has necessary access rights for driving functions of a vehicle, a transfer of the access rights to one or more further users takes place in order to be able to stop the vehicle and thus avoid an accident.

3. The method as claimed in claim 1 or 2, wherein the identification (213) of the users (3) is carried out by means of images, in particular in the range of the visible spectrum and / or in the near infrared, and / or by means of audio signals.

4. The method as claimed in any one of claims 1 to 3, wherein each user (3) is assigned a feature vector, wherein the feature vector comprises visual features, semantic features and / or auditory features, and the identification of the user (3) is carried out via a distance between the feature vector assigned to the user (3) and feature vectors stored in the database.

5. The method as claimed in any one of claims 1 to 4, wherein a change in the user (3) is acquired, in particular by means of a pressure sensor in a seat of the predetermined region (4), a seat belt sensor, a camera in the visible spectrum and / or in the near infrared and / or the audio signals, and if a new user (3) has joined, an identification (213) of the user (3) is carried out.

6. The method as claimed in any one of claims 1 to 5, wherein a user (3) identified as a registered user (3) is assigned access rights in the database (221) or in a user database and an unknown user (3) is assigned standard access rights.

7. The method as claimed in claim 6, wherein security requirements and / or an authorization must be met in order to configure the entries in the database or in the user database.

8. The method as claimed in any one of claims 1 to 7, wherein noise interference is minimized during the processing of the audio signals, and / or the audio signals are processed at least in part by means of deep learning and / or by using the context.

9. The method as claimed in any one of claims 1 to 8, wherein the audio signals are searched for predefined keywords and the analysis of the audio signals for voice commands is carried out only for a recognized keyword.

10. The method as claimed in claim 9, wherein the keywords are individually allocatable for separate users (3) and the keywords are preferably also taken into account when assigning the audio signals to the users (3).

11. The method as claimed in any one of claims 1 to 10, wherein an error message is output in the event of a lack of access rights and / or verification by an authorized user (3) is made possible.

12. The method as claimed in any one of claims 1 to 11, wherein a token is created (406) for a recognized user (3), said token comprising the access rights, a time stamp of generation and / or a lifetime and preferably being encrypted, and being concomitantly transmitted when an order is made to execute the tokens.

13. A voice control system in a vehicle (1) which is designed to execute the method as claimed in any one of claims 1 to 12, comprising at least one microphone (5) for acquiring the audio signals, and at least one computer unit (7).

14. The system as claimed in claim 13, further comprising at least one near-infrared camera (6) and one near-infrared light source.

15. A vehicle comprising a voice control system (2) as claimed in claim 13 or 14.