Vehicle control methods, systems, electronic devices and storage media
By working together with the in-vehicle terminal and server, and utilizing voice recognition and facial recognition technologies, the problem of vehicles failing to start due to users forgetting their car keys has been solved, ensuring vehicle starting security and user experience.
Patent Information
- Application Number
- CN202411809753.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-12-10
AI Technical Summary
When users forget to bring their car keys, the vehicle cannot be started, resulting in a poor user experience.
The vehicle acquires voice data and driver's seat video footage through the in-vehicle terminal, performs voice recognition and facial recognition, and the server performs matching verification. Once the voice data and facial information are successfully matched with the target user information, the vehicle is started.
This technology enables voice-controlled vehicle start-up while ensuring vehicle safety, preventing situations where users cannot start the vehicle if they forget their car keys, and improving the user experience.
Smart Images

Figure CN119636635B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent vehicle technology, and in particular to a vehicle control method, system, electronic device and storage medium. Background Technology
[0002] With the rapid development of smart car technology, users can unlock or lock their vehicles without using a car key. For example, they can unlock or lock the vehicle by touching the unlocking area inside the door handle.
[0003] However, the above method makes it easy for users to forget to bring their car keys, and the car key is still required to start the vehicle. If the user forgets to bring the car key, the vehicle will not be able to start, resulting in a poor user experience. Summary of the Invention
[0004] This application provides a vehicle control method, system, electronic device, and storage medium. The technical solution is as follows:
[0005] Firstly, a vehicle control method is provided, applied to a vehicle control system, the vehicle control system including an on-board terminal and a server, the method comprising:
[0006] The vehicle-mounted terminal acquires first voice data and a first video image from the driver's seat in the vehicle.
[0007] The vehicle terminal performs voice recognition on the first voice data. If the first voice data matches the first keyword in the target keywords provided by the target user and the first video screen indicates that there is a driver in the driver's seat, the first voice data and the first video screen are sent to the server. The first keyword indicates that the vehicle should be started.
[0008] The server matches the first voice data with the target voice data provided by the target user, and matches the driver's facial information in the first video frame with the target facial information provided by the target user. If the first voice data and the target voice data are successfully matched and the facial information is successfully matched with the target facial information, the server sends a first notification message of successful matching to the vehicle terminal.
[0009] The vehicle terminal controls the vehicle to start based on the first notification message.
[0010] In some embodiments, the server matches the first voice data with target voice data provided by the target user in the voice database, and matches the driver's facial information in the first video frame with the target facial information provided by the target user, including:
[0011] The server extracts features from the first speech data to obtain the timbre features of the first speech data, obtains a first matching degree between the timbre features and the timbre features of the target speech data, and determines that the first speech data and the target speech data are successfully matched if the first matching degree is greater than or equal to a first threshold.
[0012] The server performs facial recognition on the first video frame to obtain the driver's facial information, acquires a second matching degree between the facial information and the target facial information, and determines that the facial information and the target facial information are successfully matched if the second matching degree is greater than or equal to a second threshold.
[0013] In some embodiments, the method further includes:
[0014] The server sends a second notification message of matching failure to the vehicle terminal when the first voice data fails to match the target voice data but the face information successfully matches the target face information, or when the first voice data successfully matches the target voice data but the face information fails to match the target face information.
[0015] Based on the second notification message, the vehicle terminal sends a security verification message to the target user, and controls the vehicle to start when the target user confirms the security verification message.
[0016] In some embodiments, when the target user performs a confirmation operation on the security verification message, the method further includes any one of the following:
[0017] If the first voice data fails to match the target voice data but the face information successfully matches the target face information, the server updates the target voice data based on the matching result between the first voice data and the target voice data.
[0018] If the first voice data successfully matches the target voice data but the face information fails to match the target face information, the server updates the target face information based on the matching result between the face information and the target face information.
[0019] In some embodiments, the method further includes:
[0020] If the server fails to match the first voice data with the target voice data and fails to match the face information with the target face information, it sends a third notification message of matching failure to the vehicle terminal.
[0021] The vehicle terminal triggers an anti-theft alarm and notifies the target user based on the third notification message.
[0022] In some embodiments, after the vehicle is started, the method further includes:
[0023] The vehicle terminal acquires the second voice data and performs voice recognition on the second voice data. If the second voice data successfully matches the second keyword in the target keyword, the second voice data is sent to the server, and the second keyword indicates that the vehicle should be turned off.
[0024] The server matches the second voice data with the target voice data in the voice database. If the second voice data and the target voice data are successfully matched, the server sends a fourth notification message of successful matching to the vehicle terminal.
[0025] The vehicle terminal controls the vehicle to shut down based on the fourth notification message.
[0026] In some embodiments, after the vehicle is started, the method further includes:
[0027] The vehicle terminal acquires a second video frame of the vehicle, the second video frame including a video frame inside the vehicle and a video frame outside the vehicle;
[0028] When the vehicle terminal indicates that the vehicle is in the target scene in the second video screen, it acquires third voice data and performs voice recognition on the third voice data. If the third voice data successfully matches the second keyword in the target keywords, it controls the vehicle to shut down, and the second keyword indicates that the vehicle should be shut down.
[0029] Secondly, a vehicle control system is provided, which includes an on-board terminal and a server.
[0030] The vehicle terminal is configured to: acquire first voice data and a first video image of the driver's seat in the vehicle; perform voice recognition on the first voice data; and, if the first voice data successfully matches the first keyword in the target keywords provided by the target user and the first video image indicates that there is a driver in the driver's seat, send the first voice data and the first video image to the server, wherein the first keyword indicates that the vehicle is started.
[0031] The server is configured to match the first voice data with the target voice data provided by the target user, and match the driver's facial information in the first video frame with the target facial information provided by the target user. When the first voice data and the target voice data are successfully matched and the facial information is successfully matched with the target facial information, the server sends a first notification message of successful matching to the vehicle terminal.
[0032] The vehicle-mounted terminal is also used to control the vehicle to start based on the first notification message.
[0033] Thirdly, a vehicle control device is provided, applied to an in-vehicle terminal in a vehicle control system, the vehicle control system further including a server, the device comprising:
[0034] The acquisition module is used for the first voice data and the first video image of the driver's seat in the vehicle;
[0035] The sending module is used to perform speech recognition on the first voice data. If the first voice data matches a first keyword provided by the target user and the first video frame indicates that there is a driver in the driver's seat, the first voice data and the first video frame are sent to the server. The first keyword indicates that the vehicle should be started. The server is used to match the first voice data with the target voice data provided by the target user, and to match the driver's facial information in the first video frame with the target facial information provided by the target user. If the first voice data matches the target voice data and the facial information matches the target facial information, a first notification message indicating successful matching is sent to the vehicle terminal.
[0036] The control module is used to control the vehicle to start based on the first notification message.
[0037] Fourthly, a vehicle control device is provided, which is applied to a server in a vehicle control system. The vehicle control system further includes an on-board terminal. The device includes:
[0038] The receiving module is used to receive the first voice data sent by the vehicle terminal and the first video image from the driver's seat in the vehicle;
[0039] The matching module is used to match the first voice data with the target voice data provided by the target user, and to match the driver's facial information in the first video frame with the target facial information provided by the target user.
[0040] The sending module is used to send a first notification message of successful matching to the vehicle terminal when the first voice data and the target voice data are successfully matched and the face information is successfully matched with the target face information.
[0041] The vehicle terminal is used to perform voice recognition on the first voice data. When the first voice data successfully matches the first keyword in the target keywords provided by the target user and the first video screen indicates that there is a driver in the driver's seat, the first voice data and the first video screen are sent to the server, and the first keyword indicates that the vehicle is started; the vehicle is started based on the first notification message.
[0042] Fifthly, an electronic device is provided, comprising a memory and a processor, wherein the memory stores at least one computer program, the at least one computer program being loaded and executed by the processor to implement the steps performed by the vehicle terminal in the vehicle control method provided in the first aspect above; or to implement the steps performed by the server in the vehicle control method provided in the first aspect above.
[0043] In a sixth aspect, a computer-readable storage medium is provided, wherein at least one computer program is stored therein, the at least one computer program being loaded and executed by a processor to implement the steps performed by the vehicle terminal in the vehicle control method provided in the first aspect above; or to implement the steps performed by the server in the vehicle control method provided in the first aspect above.
[0044] In summary, in the vehicle control method provided in this application embodiment, when the vehicle terminal recognizes the keyword for starting the vehicle based on voice data, and if there is a driver in the current driver's seat, it sends the voice data and a video image of the driver's seat to the server for further security verification. When the server determines that the voice data and the driver both match the relevant information provided by the target user, it notifies the vehicle terminal to control the vehicle to start. This ensures the security of vehicle starting while achieving voice control of vehicle starting, avoids the situation where the user cannot start the vehicle if they forget their car key, and improves the user experience. Attached Figure Description
[0045] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0046] Figure 1This is a schematic diagram of an implementation environment provided in an embodiment of this application;
[0047] Figure 2 This is a flowchart of a vehicle control method provided in an embodiment of this application;
[0048] Figure 3 This is a schematic diagram of the structure of a vehicle control device provided in an embodiment of this application;
[0049] Figure 4 This is a schematic diagram of another vehicle control device provided in an embodiment of this application;
[0050] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0051] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0052] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with relevant laws, regulations, and standards. For example, the voice data, video footage, and facial information involved in this application were all obtained with full authorization.
[0053] Figure 1 This is a schematic diagram of an implementation environment provided in an embodiment of this application. For example... Figure 1 As shown, the implementation environment includes a vehicle control system, which includes an on-board terminal 100 and a server 200. The on-board terminal 100 and the server 200 are connected to each other and transmit data and interact through a network.
[0054] The vehicle terminal 100 is deployed in a physical vehicle and runs a client application to provide various functions, such as vehicle control and status monitoring, navigation and positioning, multimedia entertainment, communication and social networking, personalization settings and user management, etc., which are not limited in this application. In some embodiments, the vehicle terminal 100 includes the following components to achieve various functions.
[0055] The power management component 101 is used to manage the vehicle's power supply. The power management component 101 is connected to the processor 102, microphone 103, and network card 104.
[0056] Processor 102 is used to process various data and instructions. Processor 102 is connected to other components through multiple interfaces. Schematic, processor 102 includes various types of processors, such as microcontroller units (MCUs) and system-on-chips (SoCs), with the MCU and SoC working together to implement the functions of processor 102.
[0057] Microphone 103 is used to collect voice signals.
[0058] Network interface card 104 is used to provide network connectivity. The vehicle terminal 100 transmits various vehicle data to the server 200 through the network interface card 104.
[0059] External interface 105 is used to provide a connection interface with external devices.
[0060] Communication interface 106 includes, for example, a Universal Serial Bus (USB) interface for connecting USB devices, such as storage devices or other USB peripherals; a Controller Area Network (CAN) interface for enabling in-vehicle network communication; and a Local Interconnect Network (LIN) interface, which is an auxiliary bus network that works with the CAN bus to enable communication between different Electronic Control Units (ECUs) in the vehicle.
[0061] Memory 107 is used to store data and programs. Schematic, memory 107 includes, but is not limited to, non-volatile memory (such as FLASH memory), static random access memory, etc.
[0062] Server 200 provides backend services to vehicle terminal 100. In some embodiments, server 200 includes database 201, database operation module 202, network transmission module 203, and control module 204. Database 201 stores voice data, such as voice data obtained after signal enhancement processing of the original voice signal, or the original voice signal itself; this application does not limit this. Database 201 also stores facial information, such as images or videos containing faces, or facial features obtained after feature extraction from images; this application does not limit this. Database operation module 202 reads data from database 201 and stores data in database 201. Network transmission module 203 provides network connectivity. Server 200 receives various data sent by vehicle terminal 100 and sends various data to vehicle terminal 100 through network transmission module 203. Control module 204 controls the other modules to ensure coordinated operation of each module.
[0063] Schematic illustration: Server 200 may be, for example, a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The number of servers 200 may be more or less, and this embodiment of the application does not limit this. Of course, server 200 may also include other functional servers to provide more comprehensive and diversified services.
[0064] In this application, the vehicle-mounted terminal 100 and the server 200 together constitute a vehicle control system for implementing the vehicle control method provided in this application. The vehicle-mounted terminal 100 and the server 200 implement the vehicle control method through interaction.
[0065] In some embodiments, the network uses standard communication technologies and / or protocols. The network is typically the Internet, but can be any network, including but not limited to any combination of Local Area Networks (LANs), Metropolitan Area Networks (MANs), Wide Area Networks (WANs), mobile, wired or wireless networks, private networks, or Virtual Private Networks. In some embodiments, technologies and / or formats including Hyper Text Markup Language (HTML), Extensible Markup Language (XML), etc., are used to represent data exchanged over the network. Furthermore, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Networks (VPNs), and Internet Protocol Security (IPsec) can be used to encrypt all or some links. In other embodiments, custom and / or dedicated data communication technologies can be used to replace or supplement the aforementioned data communication technologies.
[0066] The following is for reference. Figure 2 Taking the interaction between the on-board terminal and the server in a vehicle control system as an example, this paper introduces the vehicle control method provided in this application.
[0067] Figure 2 This is a flowchart of a vehicle control method provided in an embodiment of this application. Figure 2 As shown, the method is applied to a vehicle control system, which includes an on-board terminal and a server. Taking the interaction between the on-board terminal and the server as an example, the method includes the following steps 201 to 208.
[0068] 201. The vehicle terminal acquires the first voice data and the first video image from the driver's seat in the vehicle.
[0069] In this embodiment, the vehicle can be a gasoline-powered vehicle, an electric vehicle, or a hybrid vehicle; this application does not limit the type. Indicatively, the vehicle is equipped with a microphone and a camera. The microphone can be located in multiple positions within the vehicle, such as the center of the roof or in front of the driver's seat. The camera is located in front of or to the side of the driver's seat to ensure that a video image including the driver's seat can be captured. The in-vehicle terminal collects in-vehicle voice signals through the microphone, converts the collected voice signals into digital signals, and obtains first voice data for subsequent processing in the digital domain. Furthermore, the in-vehicle terminal communicates with the camera, for example, via a USB interface or a wireless network; this application does not limit the type of communication. The camera captures an image of the driver's seat and transmits the captured first video image of the driver's seat to the in-vehicle terminal.
[0070] 202. The vehicle terminal performs voice recognition on the first voice data. If the first keyword in the first voice data is successfully matched with the target keyword provided by the target user and the first video screen indicates that there is a driver in the driver's seat, the first voice data and the first video screen are sent to the server.
[0071] In this embodiment, the target user has the authority to control vehicle startup, such as the vehicle owner. The in-vehicle terminal stores target keywords provided by the target user. These target keywords instruct the vehicle to perform various operations, such as starting the vehicle, stopping the vehicle, unlocking the vehicle, locking the vehicle, opening the door, closing the door, opening the window, heating the seat, etc. This application does not limit the number of target keywords. The target keywords include a first keyword, which instructs the vehicle to start. For example, the first keyword could be "start vehicle," "vehicle ignition," or "start vehicle," etc. In this application, starting the vehicle instructs the vehicle engine to change from a stationary state to a running state, that is, entering the driving preparation stage. Taking a gasoline vehicle as an example, starting the vehicle also means controlling the ignition system to perform an ignition operation to ensure the engine starts smoothly. Taking an electric vehicle as an example, starting the vehicle also means driving the motor to start running through the motor control unit, transmitting power to the wheels, and enabling the vehicle to start smoothly.
[0072] Schematic illustration: The in-vehicle terminal performs speech recognition on the first voice data to convert it into text, matches the recognized text with target keywords, and determines whether the recognized text contains the first keyword provided by the target user. If it does (e.g., the recognized text includes the first keyword), the match is considered successful. Additionally, the in-vehicle terminal performs image recognition on the first video frame to determine whether a driver is present in the driver's seat.
[0073] In this step, if the first voice data successfully matches the first keyword and the first video image indicates that there is a driver in the driver's seat, it means that the driver is in place and the vehicle needs to be started. Based on this, the vehicle terminal sends the first voice data and the first video image to the server to perform a safety verification of whether the driver meets the conditions for starting the vehicle, ensuring the safety of starting the vehicle under voice control. In some embodiments, if there is a driver in the driver's seat, the vehicle terminal can also perform posture analysis on the driver to determine whether the driver is in a normal sitting posture, whether there is any obvious obstruction to the body, etc., thereby further confirming whether the driver's condition is safe for driving.
[0074] In addition, during the process of sending the first voice data and the first video image to the server, the vehicle terminal also sends a Vehicle Identification Number (VIN) to the server to identify the vehicle. In some embodiments, the vehicle terminal may also send the acquisition time of the first voice data and the first video image, so that the server can correctly parse and process the data; this application does not limit this.
[0075] 203. The server matches the first voice data with the target voice data provided by the target user, and matches the driver's facial information in the first video frame with the target facial information provided by the target user.
[0076] In this embodiment, the server's database stores target voice data and target facial information provided by the target user. The number of target voice data items can be one or more, such as the voice data of the vehicle owner or voice data of other users authorized by the vehicle owner. Similarly, the number of target facial information items can be one or more, such as the facial information of the vehicle owner or facial information of other users authorized by the vehicle owner. Indicatively, the database uses a relational database (such as MySQL) or a non-relational database (such as MongoDB) to store the target voice data and target facial information. Each record corresponds to the identity information of a target user and their pre-registered target voice data and target facial information.
[0077] In some embodiments, the server extracts features from the first speech data to obtain timbre features, and acquires a first matching degree between the timbre features and the timbre features of the target speech data. If the first matching degree is greater than or equal to a first threshold, it is determined that the first speech data and the target speech data are successfully matched. The first threshold is a preset threshold that can be set according to the feature extraction algorithm used by the server and business requirements; for example, the first threshold is 90%. Illustratively, the server compares the timbre features of the first speech data with those of the target speech data one by one. During the comparison process, a distance metric algorithm (such as Euclidean distance or cosine similarity) is used to calculate the first matching degree between different timbre features. Furthermore, this application does not limit the method by which the server determines whether the first speech data and the target speech data are successfully matched. For example, the server is configured with a speech feature extraction model based on deep learning (such as a convolutional neural network or a recurrent neural network). The server inputs the first speech data and the target speech data into the speech feature extraction model respectively to obtain the timbre features of the first speech data and the timbre features of the target speech data. The timbre features extracted by the speech feature extraction model can accurately characterize the uniqueness of the user from whom the speech data originates, such as the shape of the vocal tract and pronunciation habits.
[0078] In some embodiments, the server performs facial recognition on the first video frame to obtain the driver's facial information, acquires a second matching degree between the facial information and the target facial information, and determines that the facial information and the target facial information are successfully matched if the second matching degree is greater than or equal to a second threshold. The second threshold is a preset threshold that can be set according to the facial recognition algorithm used by the server and business requirements; for example, the second threshold is 95%. Illustratively, the server compares the driver's facial information with the target facial information one by one. During the comparison process, a distance metric algorithm (such as Euclidean distance or cosine similarity) is used to calculate the second matching degree between different facial information. Furthermore, this application does not limit the facial recognition algorithm used by the server. For example, the server may be configured with a deep learning-based facial feature extraction model (such as a convolutional neural network or a recurrent neural network). The server inputs the first video frame into the facial feature extraction model to extract key facial features, such as the position, shape, and texture of facial features, and uses these key features as the driver's facial information to match the target facial information pre-provided by the target user.
[0079] In some embodiments, the server records detailed matching information during the matching process, including the matching time, the similarity value of the match, the vehicle and owner information involved, etc., and stores this information in the operation log table of the database for subsequent statistical analysis and troubleshooting.
[0080] 204. If the server successfully matches the first voice data with the target voice data and the face information successfully matches the target face information, the server sends a first notification message of successful matching to the vehicle terminal.
[0081] In this embodiment, when the first voice data successfully matches the target voice data and the face information successfully matches the target face information, the server generates a first notification message indicating successful matching. Using an encrypted transmission mechanism, the first notification message is sent to the vehicle terminal via the network. Indicatively, this first notification message includes an identification field (such as a status code indicating successful matching) and additional information (such as an identifier indicating successful driver authentication, authorization information, etc.) to facilitate subsequent operations by the vehicle terminal.
[0082] 205. The vehicle terminal controls the vehicle to start based on the first notification message.
[0083] In this embodiment, after receiving the first notification message, the vehicle terminal verifies its legality and integrity. For example, it checks whether the message's source is legitimate and whether it has been tampered with. The reliability of the message is ensured by verifying its digital signature or checksum. After confirming the validity of the first notification message, the vehicle terminal communicates with the vehicle's Electronic Control Unit (ECU) and sends a start command to the relevant ECUs responsible for starting the vehicle (such as the engine control unit of a gasoline vehicle or the motor control unit of an electric vehicle). Upon receiving the start command, these ECUs operate according to the vehicle's predetermined start-up procedure.
[0084] After steps 201 to 205, when the vehicle terminal recognizes the keyword for starting the vehicle based on the voice data, and if there is a driver in the current driver's seat, it sends the voice data and a video image of the driver's seat to the server for further security verification. When the server confirms that the voice data and the driver both match the relevant information provided by the target user, it notifies the vehicle terminal to control the vehicle to start. Since the data stored on the server is not easily tampered with and has high security, security verification through the server can ensure the security of vehicle starting while realizing voice control of vehicle starting, avoiding situations where users cannot start the vehicle if they forget their car keys or the car keys are faulty, thus improving the user experience.
[0085] In some embodiments, if the server fails to match the first voice data with the target voice data but successfully matches the facial information with the target facial information, or if the first voice data matches the target voice data but fails to match the facial information with the target facial information, the server sends a second notification message indicating a matching failure to the vehicle terminal. The second notification message includes the reason for the matching failure. Based on the second notification message, the vehicle terminal sends a security verification message to the target user. If the target user confirms the security verification message, the vehicle starts. Schematic, if either the first voice data or the driver's facial information matches successfully, it indicates that the current driver may meet the conditions for voice-controlled vehicle start. Based on this, the vehicle terminal sends a security verification message to the target user to obtain the target user's authorization. For example, the vehicle terminal sends the security verification message to the target user's mobile terminal via SMS or a phone call. Correspondingly, the target user confirms the security verification message by replying to a confirmation SMS or during a call; this application does not limit this. This approach fully considers various scenarios for voice-controlled vehicle start-up. For example, if the driver's voice is hoarse due to a cold, voice data matching may fail; similarly, if the driver has a facial injury or is wearing makeup, facial information matching may fail. Through further security verification, this method can reduce the false rejection rate and improve the user experience.
[0086] In some embodiments, when the target user performs a confirmation operation on the security verification message, the vehicle terminal sends the confirmation operation back to the server. Then, if the first voice data fails to match the target voice data but the face information matches the target face information successfully, the server updates the target voice data based on the matching result between the first voice data and the target voice data. Similarly, if the first voice data matches the target voice data successfully but the face information fails to match the target face information, the server updates the target face information based on the matching result between the face information and the target face information. This process means that after either the first voice data or the driver's face information matches successfully and the target user has verified the driver's identity, it indicates that the first voice data or face information that failed to match meets the conditions for voice-controlled vehicle start-up. Based on this, the server promptly updates this information to the target voice data and target face information provided by the target user, thereby improving the accuracy and generalization ability of subsequent matching.
[0087] In other embodiments, if the server fails to match the first voice data with the target voice data and also fails to match the face information with the target face information, it sends a third notification message indicating a matching failure to the in-vehicle terminal. This third notification message includes the reason for the matching failure. Based on the third notification message, the in-vehicle terminal triggers an anti-theft alarm and notifies the target user. In this way, the mechanism for voice-controlled vehicle start is tightly integrated with the vehicle's security systems (such as anti-theft systems and driving safety assistance systems). If an attempt to start the vehicle is detected by a non-owner's voice, not only is start-up refused, but an anti-theft alarm is also triggered, and the vehicle owner is notified, ensuring vehicle security.
[0088] 206. The vehicle terminal acquires the second voice data and performs voice recognition on the second voice data. If the second voice data successfully matches the second keyword in the target keywords, the second voice data is sent to the server.
[0089] In this embodiment, after the vehicle terminal starts the vehicle, it acquires second voice data and performs voice recognition on the second voice data. Based on the aforementioned step 202, the target keyword instructs the vehicle to perform various operations. In this step, the target user also has the authority to control the vehicle to shut down, such as the vehicle owner. The target keyword includes a second keyword, which instructs the vehicle to shut down. For example, the second keyword could be "shut down the vehicle," "vehicle engine off," or simply "engine off," etc. In this application, shutting down the vehicle instructs the vehicle engine to change from a running state to a stationary state.
[0090] Schematic illustration: The in-vehicle terminal performs speech recognition on the second voice data to convert it into text. The recognized text is then matched against target keywords to determine if the recognized text contains the second keyword provided by the target user. If it does (e.g., the recognized text includes the second keyword), a successful match is confirmed, and the second voice data is sent to the server. It should be understood that since turning off the vehicle does not necessarily require the driver to be in the driver's seat, this step may disregard video footage from the driver's seat. However, in some scenarios, video footage from the driver's seat can be acquired, and the second voice data can be sent to the server only if the video footage indicates that a driver is present in the driver's seat.
[0091] 207. The server matches the second voice data with the target voice data provided by the target user. If the second voice data and the target voice data match successfully, the server sends a fourth notification message of successful matching to the vehicle terminal.
[0092] In this embodiment, the server matches the second voice data with the target voice data in the same way as in step 203, and therefore will not be described again. If the second voice data and the target voice data match successfully, the server generates a fourth notification message indicating successful matching, and sends the fourth notification message to the vehicle terminal via the network using an encrypted transmission mechanism. Indicatively, this fourth notification message includes an identification field (such as a status code indicating successful matching) and additional information (such as a driver authentication identifier, authorization information, etc.) to facilitate subsequent operations by the vehicle terminal.
[0093] 208. The vehicle terminal controls the vehicle to shut down based on the fourth notification message.
[0094] In this embodiment, after receiving the fourth notification message, the vehicle terminal verifies its legality and integrity. For example, it checks whether the message's source is legitimate and whether it has been tampered with. The reliability of the message is ensured by verifying its digital signature or checksum. After confirming the validity of the fourth notification message, the vehicle terminal communicates with the vehicle's electronic control unit (ECU) and sends a shutdown command to the relevant ECU responsible for shutting down the vehicle (such as the engine control unit of a gasoline vehicle or the motor control unit of an electric vehicle). Upon receiving the shutdown command, these ECUs operate according to the vehicle's predetermined shutdown procedure. Specifically, when shutting down the vehicle, for gasoline vehicles, the engine stops running, and the fuel supply and ignition systems cease operation; for electric vehicles, the drive motor stops running, the high-voltage battery pack stops supplying power to the motor, and various control units of the vehicle enter a low-power or sleep state.
[0095] After steps 206 to 208, when the vehicle terminal recognizes the keyword for turning off the vehicle based on the voice data, it sends voice data to the server for further security verification. When the server determines that the voice data and the relevant information provided by the target user are successfully matched, it notifies the vehicle terminal to control the vehicle to turn off. This ensures the security of vehicle closing while achieving voice control, avoids situations where users cannot turn off the vehicle if they forget their car keys or the car keys are faulty, and improves the user experience.
[0096] In some embodiments, after controlling the vehicle to start, the in-vehicle terminal acquires a second video view of the vehicle, including video views of the vehicle's interior and exterior. When the second video view indicates that the vehicle is in a target scenario, the in-vehicle terminal acquires third voice data and performs voice recognition on the third voice data. If the third voice data successfully matches the second keyword in the target keywords, the terminal controls the vehicle to shut down. The vehicle is equipped with multiple cameras to capture video views of the vehicle's interior and exterior. The target scenario refers to a situation requiring emergency stopping, such as a traffic accident or other sudden event. Through this method, after the vehicle starts, the in-vehicle terminal acquires real-time video views of the vehicle. When the vehicle is identified as being in a target scenario, and the voice data matches the second keyword, there is no need for security verification through a server. This means the in-vehicle terminal can directly control the vehicle to shut down, avoiding the time delay caused by information interaction between the in-vehicle terminal and the server in the target scenario, which could affect driving safety.
[0097] In summary, in the vehicle control method provided in this application, when the in-vehicle terminal recognizes a keyword for starting the vehicle based on voice data, and if a driver is present in the driver's seat, it sends voice data and a video feed from the driver's seat to the server for further security verification. Once the server confirms that the voice data and the driver information match the relevant information provided by the target user, it notifies the in-vehicle terminal to start the vehicle. This ensures the security of vehicle starting while enabling voice control, preventing situations where the user cannot start the vehicle due to forgetting their car key or a malfunctioning key, thus improving the user experience. Furthermore, after the vehicle starts, when the in-vehicle terminal recognizes a keyword for turning off the vehicle based on voice data, it sends voice data to the server for further security verification. Once the server confirms that the voice data matches the relevant information provided by the target user, it notifies the in-vehicle terminal to turn off the vehicle. This ensures the security of vehicle closing while enabling voice control, preventing situations where the user cannot close the vehicle due to forgetting their car key or a malfunctioning key, further improving the user experience.
[0098] See Figure 3 This application provides a vehicle control device, which is configured on the vehicle terminal of a vehicle control system. The vehicle control system also includes a server. The device includes an acquisition module 301, a transmission module 302, and a control module 303.
[0099] The acquisition module 301 is used for the first voice data and the first video image of the driver's seat in the vehicle;
[0100] The sending module 302 is used to perform speech recognition on the first voice data. If the first voice data matches the first keyword provided by the target user and the first video screen indicates that there is a driver in the driver's seat, the first voice data and the first video screen are sent to the server. The first keyword indicates that the vehicle should be started. The server is used to match the first voice data with the target voice data provided by the target user and match the driver's facial information in the first video screen with the target facial information provided by the target user. If the first voice data matches the target voice data and the facial information matches the target facial information, the server sends a first notification message of successful matching to the vehicle terminal.
[0101] The control module 303 is used to control the vehicle to start based on the first notification message.
[0102] In some embodiments, the apparatus further includes:
[0103] The receiving module is used to receive a second notification message of matching failure. The second notification message is sent by the server to the vehicle terminal when the first voice data fails to match the target voice data but the face information matches the target face information successfully, or when the first voice data matches the target voice data successfully but the face information fails to match the target face information.
[0104] The control module 303 is also used to send a security verification message to the target user based on the second notification message, and control the vehicle to start when the target user performs a confirmation operation on the security verification message.
[0105] In some embodiments, the apparatus further includes:
[0106] The security module is used to trigger an anti-theft alarm and notify the target user based on a third notification message sent by the server indicating a failed match. The third notification message is sent by the server to the vehicle terminal when the first voice data fails to match the target voice data and the face information fails to match the target face information.
[0107] In some embodiments, after the vehicle is started, the acquisition module 301 is further configured to:
[0108] Acquire the second voice data and perform speech recognition on the second voice data. If the second voice data successfully matches the second keyword in the target keyword, send the second voice data to the server. The second keyword indicates that the vehicle should be turned off.
[0109] The control module 303 is also used to control the vehicle to shut down based on a fourth notification message indicating a matching failure. The fourth notification message is sent by the server to the vehicle terminal when the second voice data and the target voice data in the voice database are matched successfully.
[0110] In some embodiments, after the vehicle is started, the acquisition module 301 is further configured to:
[0111] Acquire a second video view of the vehicle, which includes video views of the vehicle's interior and exterior.
[0112] The control module 303 is also used to acquire third voice data and perform voice recognition on the third voice data when the second video screen indicates that the vehicle is in the target scene. When the third voice data successfully matches the second keyword in the target keywords, the control module 303 controls the vehicle to turn off, and the second keyword indicates that the vehicle should be turned off.
[0113] In this device, when the in-vehicle terminal recognizes the keyword used to start the vehicle based on voice data, and if there is a driver in the current driver's seat, it sends the voice data and a video image of the driver's seat to the server for further security verification. When the server confirms that the voice data and the driver both match the relevant information provided by the target user, it notifies the in-vehicle terminal to control the vehicle to start. This not only enables voice control of vehicle starting but also ensures the security of vehicle starting and avoids situations where users cannot start the vehicle if they forget their car keys or the car keys malfunction, thus improving the user experience.
[0114] It should be noted that the vehicle control device provided in the above embodiments is only illustrated by the division of the above functional modules when controlling the vehicle. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure can be divided into different functional modules to complete all or part of the functions described above. In addition, the vehicle control device and the vehicle control method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0115] See Figure 4 This application provides a vehicle control device, which is configured on the server of a vehicle control system. The vehicle control system also includes an on-board terminal. The device includes a receiving module 401, a matching module 402, and a sending module 403.
[0116] The receiving module 401 is used to receive the first voice data and the first video image of the driver's seat in the vehicle sent by the vehicle terminal;
[0117] The matching module 402 is used to match the first voice data with the target voice data provided by the target user, and to match the driver's facial information in the first video frame with the target facial information provided by the target user.
[0118] The sending module 403 is used to send a first notification message of successful matching to the vehicle terminal when the first voice data and the target voice data are successfully matched and the face information is successfully matched with the target face information.
[0119] The vehicle terminal is used to perform voice recognition on the first voice data. If the first keyword is successfully matched with the target keyword provided by the target user and the first video screen indicates that there is a driver in the driver's seat, the first voice data and the first video screen are sent to the server. The first keyword indicates that the vehicle should be started. The vehicle is started based on the first notification message.
[0120] In some embodiments, the matching module 402 is configured to:
[0121] Feature extraction is performed on the first speech data to obtain the timbre features of the first speech data, and the first matching degree between the timbre features and the timbre features of the target speech data is obtained. If the first matching degree is greater than or equal to the first threshold, it is determined that the first speech data and the target speech data are successfully matched.
[0122] Facial recognition is performed on the first video frame to obtain the driver's facial information. A second matching degree is obtained between the facial information and the target facial information. If the second matching degree is greater than or equal to a second threshold, it is determined that the facial information and the target facial information are successfully matched.
[0123] In some embodiments, the sending module 403 is further configured to: send a second notification message of matching failure to the vehicle terminal when the first voice data fails to match the target voice data but the face information matches the target face information successfully, or when the first voice data matches the target voice data successfully but the face information fails to match the target face information successfully.
[0124] In some embodiments, when the target user performs an acknowledgment operation on the security verification message, the apparatus further includes an update module for any of the following:
[0125] If the first voice data fails to match the target voice data but the face information matches the target face information successfully, the target voice data is updated based on the matching result between the first voice data and the target voice data.
[0126] If the first voice data and the target voice data match successfully but the face information and the target face information fail to match, the target face information is updated based on the matching result between the face information and the target face information.
[0127] In some embodiments, the sending module 403 is further configured to: send a third notification message of matching failure to the vehicle terminal when the first voice data fails to match the target voice data and the face information fails to match the target face information.
[0128] In this device, when the in-vehicle terminal recognizes the keyword used to start the vehicle based on voice data, and if there is a driver in the current driver's seat, it sends the voice data and a video image of the driver's seat to the server for further security verification. When the server confirms that the voice data and the driver both match the relevant information provided by the target user, it notifies the in-vehicle terminal to control the vehicle to start. This not only enables voice control of vehicle starting but also ensures the security of vehicle starting and avoids situations where users cannot start the vehicle if they forget their car keys or the car keys malfunction, thus improving the user experience.
[0129] It should be noted that the vehicle control device provided in the above embodiments is only illustrated by the division of the above functional modules when controlling the vehicle. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure can be divided into different functional modules to complete all or part of the functions described above. In addition, the vehicle control device and the vehicle control method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0130] refer to Figure 5 This application also provides an electronic device. Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 500 can vary considerably depending on its configuration or performance, and may include one or more Central Processing Units (CPUs) 501 and one or more memories 502. The memories 502 store at least one computer program, which is loaded and executed by the processor 501 to implement the steps performed by the vehicle terminal or server in the vehicle control method provided in the above-described method embodiment. Of course, the electronic device may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The electronic device may also include other components for implementing device functions, which will not be elaborated here.
[0131] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0132] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0133] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A vehicle control method, characterized in that, Applied to a vehicle control system, the vehicle control system including an on-board terminal and a server, the method includes: The vehicle-mounted terminal acquires first voice data and a first video image from the driver's seat in the vehicle. The vehicle terminal performs voice recognition on the first voice data. If the first voice data matches the first keyword in the target keywords provided by the target user and the first video screen indicates that there is a driver in the driver's seat, the first voice data and the first video screen are sent to the server. The first keyword indicates that the vehicle should be started. The server matches the first voice data with the target voice data provided by the target user, and matches the driver's facial information in the first video frame with the target facial information provided by the target user. If the first voice data and the target voice data are successfully matched and the facial information is successfully matched with the target facial information, the server sends a first notification message of successful matching to the vehicle terminal. The vehicle terminal controls the vehicle to start based on the first notification message; The server sends a second notification message of matching failure to the vehicle terminal when the first voice data fails to match the target voice data but the face information successfully matches the target face information, or when the first voice data successfully matches the target voice data but the face information fails to match the target face information. Based on the second notification message, the vehicle terminal sends a security verification message to the target user, and controls the vehicle to start when the target user confirms the security verification message.
2. The method according to claim 1, characterized in that, The server matches the first voice data with the target voice data provided by the target user in the voice database, and matches the driver's facial information in the first video frame with the target facial information provided by the target user, including: The server extracts features from the first speech data to obtain the timbre features of the first speech data, obtains a first matching degree between the timbre features and the timbre features of the target speech data, and determines that the first speech data and the target speech data are successfully matched if the first matching degree is greater than or equal to a first threshold. The server performs facial recognition on the first video frame to obtain the driver's facial information, acquires a second matching degree between the facial information and the target facial information, and determines that the facial information and the target facial information are successfully matched if the second matching degree is greater than or equal to a second threshold.
3. The method according to claim 1, characterized in that, When the target user performs a confirmation operation on the security verification message, the method further includes any one of the following: If the first voice data fails to match the target voice data but the face information successfully matches the target face information, the server updates the target voice data based on the matching result between the first voice data and the target voice data. If the first voice data successfully matches the target voice data but the face information fails to match the target face information, the server updates the target face information based on the matching result between the face information and the target face information.
4. The method according to claim 1, characterized in that, The method further includes: If the server fails to match the first voice data with the target voice data and fails to match the face information with the target face information, it sends a third notification message of matching failure to the vehicle terminal. The vehicle terminal triggers an anti-theft alarm and notifies the target user based on the third notification message.
5. The method according to claim 1, characterized in that, After the vehicle is started, the method further includes: The vehicle terminal acquires the second voice data and performs voice recognition on the second voice data. If the second voice data successfully matches the second keyword in the target keyword, the second voice data is sent to the server, and the second keyword indicates that the vehicle should be turned off. The server matches the second voice data with the target voice data in the voice database. If the second voice data and the target voice data are successfully matched, the server sends a fourth notification message of successful matching to the vehicle terminal. The vehicle terminal controls the vehicle to shut down based on the fourth notification message.
6. The method according to claim 1, characterized in that, After the vehicle is started, the method further includes: The vehicle terminal acquires a second video frame of the vehicle, the second video frame including a video frame inside the vehicle and a video frame outside the vehicle; When the vehicle terminal indicates that the vehicle is in the target scene in the second video screen, it acquires third voice data and performs voice recognition on the third voice data. If the third voice data successfully matches the second keyword in the target keywords, it controls the vehicle to shut down, and the second keyword indicates that the vehicle should be shut down.
7. A vehicle control system, characterized in that, The vehicle control system includes an on-board terminal and a server; The vehicle terminal is configured to: acquire first voice data and a first video image of the driver's seat in the vehicle; perform voice recognition on the first voice data; and, if the first voice data successfully matches the first keyword in the target keywords provided by the target user and the first video image indicates that there is a driver in the driver's seat, send the first voice data and the first video image to the server, wherein the first keyword indicates that the vehicle is started. The server is configured to match the first voice data with the target voice data provided by the target user, and match the driver's facial information in the first video frame with the target facial information provided by the target user. When the first voice data and the target voice data are successfully matched and the facial information is successfully matched with the target facial information, the server sends a first notification message of successful matching to the vehicle terminal. The vehicle terminal is also used to control the vehicle to start based on the first notification message; The server is further configured to send a second notification message of matching failure to the vehicle terminal when the first voice data fails to match the target voice data but the face information successfully matches the target face information, or when the first voice data successfully matches the target voice data but the face information fails to match the target face information. The vehicle terminal is further configured to send a security verification message to the target user based on the second notification message, and control the vehicle to start when the target user performs a confirmation operation on the security verification message.
8. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores at least one computer program, which is loaded and executed by the processor to implement the steps performed by the vehicle terminal in the vehicle control method as described in any one of claims 1-6, or to implement the steps performed by the server in the vehicle control method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to implement the steps performed by the vehicle terminal in the vehicle control method as described in any one of claims 1-6, or to implement the steps performed by the server in the vehicle control method as described in any one of claims 1-6.
Citation Information
Patent Citations
Control method for automobile movement, and automobile
CN110843725A
Vehicle self-starting method and device, terminal equipment and storage medium
CN115139977A