Vehicle control method and vehicle
By combining gesture interaction and voice interaction, vehicle control instructions are obtained and final instructions are determined based on clarity in conflict, the problem of vehicle control inaccuracy caused by noise interference by a single voice interaction is solved, and higher control accuracy is achieved.
Patent Information
- Application Number
- CN202411994893.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-02
Smart Images

Figure CN119911288A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of vehicle technology, and in particular to a vehicle control method and a vehicle. Background Art
[0002] In related technologies, vehicle control mainly relies on hardware interactions such as central control screens, physical buttons, and touch screens. However, such hardware interactions may distract the driver during driving and reduce driving safety. Therefore, in order to improve driving experience and driving safety, most vehicles are currently using voice interaction to control the vehicle. However, voice interaction is easily affected by noise or sound interruptions, resulting in inaccurate voice recognition, which in turn causes inaccurate vehicle control. Summary of the invention
[0003] In view of this, the present disclosure provides a vehicle control method and a vehicle to solve the problem of inaccurate vehicle control.
[0004] In a first aspect, the present disclosure provides a vehicle control method, the method comprising:
[0005] Acquire gesture images and voice information of passengers in the target vehicle;
[0006] Performing gesture recognition on the gesture image to obtain a target gesture control instruction;
[0007] Performing voice recognition on the voice information to obtain a target voice control instruction;
[0008] If there is a conflict between the target gesture control instruction and the target voice control instruction, obtaining a first clarity of the target gesture control instruction and a second clarity of the target voice control instruction;
[0009] Based on the first clarity and the second clarity, a vehicle control instruction is determined from the target gesture control instruction and the target voice control instruction to control the target vehicle.
[0010] In a second aspect, the present disclosure provides a vehicle, the vehicle comprising:
[0011] Vehicle body;
[0012] A vehicle control system, wherein the vehicle control system is used to execute the above-mentioned vehicle control method.
[0013] The vehicle control method disclosed in the present invention combines two control modes, gesture interaction and voice interaction, to control the target vehicle. Therefore, it can avoid the problem of inaccurate vehicle control caused by inaccurate single voice interaction to a certain extent. On this basis, if there is a conflict between the identified target gesture control instruction and the target voice control instruction, the vehicle control instruction is determined based on the clarity of the target gesture control instruction and the target voice control instruction to control the target vehicle. Therefore, it can avoid the problem of vehicle control conflict caused by the simultaneous use of gesture interaction and voice interaction, thereby effectively improving the accuracy of vehicle control. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the specific embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0015] Figure 1 is a flow chart of a vehicle control method according to an embodiment of the present disclosure;
[0016] Figure 2 is a structural block diagram of a vehicle control device according to an embodiment of the present disclosure;
[0017] Figure 3 is a structural block diagram of a vehicle controller according to an embodiment of the present disclosure;
[0018] Figure 4 It is a structural block diagram of a vehicle control system according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present disclosure.
[0020] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0021] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as a vehicle, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.
[0022] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the vehicle.
[0023] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet the relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0024] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.
[0025] In related technologies, vehicle control mainly relies on hardware interactions such as central control screens, physical buttons, and touch screens. However, such hardware interactions may distract the driver during driving and reduce driving safety. Therefore, in order to improve driving experience and driving safety, most vehicles are currently beginning to use voice interaction to control the vehicle.
[0026] The voice-interactive vehicle control method usually recognizes the commands issued by the driver or passenger to the cockpit based on the voice recognition function, and then adjusts and controls the cockpit. However, in actual use, controlling the vehicle only by voice recognition is susceptible to noise interference. For example, when the vehicle is playing navigation or music, waking up the voice recognition will lower the sound of the navigation and music, thereby affecting the driving experience. When the car is noisy, the accuracy of voice recognition will also be affected by excessive noise, resulting in inaccurate voice recognition. In addition, sound interruptions also lead to inaccurate voice recognition, resulting in inaccurate vehicle control.
[0027] In view of this, according to an embodiment of the present disclosure, a vehicle control method embodiment is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0028] In this embodiment, a vehicle control method is provided, which can be used in a vehicle control system of a target vehicle. Figure 1 is a flow chart of a vehicle control method according to an embodiment of the present disclosure, such as Figure 1 As shown, the process includes the following steps:
[0029] Step S101, acquiring gesture images and voice information of passengers in the target vehicle.
[0030] Specifically, at least one image acquisition device is deployed in the cabin of the target vehicle, wherein the image acquisition device includes a camera, and the gesture image of the occupant is captured by the image acquisition device.
[0031] Furthermore, the target deep learning algorithm is used to segment and locate the image captured by the image acquisition device, and the gesture image and the background image are segmented from the captured image. Therefore, even in the case of complex gestures (such as bending fingers and rotating palms), the gesture features of the gesture image can be accurately identified. Among them, the preset deep learning algorithm can be trained with the gestures and backgrounds of a large number of gesture samples to obtain the target deep learning algorithm, so that the target deep learning algorithm can learn the subtle differences between the gestures and the background, and achieve accurate gesture segmentation and positioning.
[0032] Optionally, the target deep learning algorithm includes a deep learning algorithm for target detection and instance segmentation (Mask Region-based Convolutional Neural Network, Mask R-CNN).
[0033] Specifically, at least one voice collection device is deployed in the cockpit of the target vehicle. The voice collection device includes a sound receiving device, which uses microphone array technology to capture voice input from all directions. The sound receiving device has the characteristics of high sensitivity and low noise, so it can ensure that the voice of the occupants can be clearly recorded in various environments in the car. In addition, the voice information of the occupants can be captured through the voice collection device.
[0034] Step S102: performing gesture recognition on the gesture image to obtain a target gesture control instruction.
[0035] Specifically, feature extraction is performed on the gesture image to obtain the gesture features of the gesture image, and the gesture features are matched with a preset gesture library to obtain a target gesture control instruction.
[0036] Step S103, performing voice recognition on the voice information to obtain a target voice control instruction.
[0037] Specifically, a signal processing algorithm is used to perform noise reduction processing on the voice information to reduce the influence of the noise in the target vehicle on the voice recognition. Then, the voice information after noise reduction is subjected to voice recognition to obtain the target voice control instruction.
[0038] Step S104: if there is a conflict between the target gesture control instruction and the target voice control instruction, a first clarity of the target gesture control instruction and a second clarity of the target voice control instruction are obtained.
[0039] It should be noted that if there is no conflict between the target gesture control instruction and the target voice control instruction, corresponding vehicle control instructions are generated based on the target gesture control instruction and the target voice control instruction respectively to control the target vehicle. For example, if the target gesture control instruction represents opening the sunroof, and the target voice control instruction represents playing music, then a vehicle control instruction for opening the sunroof is generated based on the target gesture control instruction, and a vehicle control instruction for playing music is generated based on the target voice control instruction.
[0040] Step S105 , based on the first clarity and the second clarity, determining a vehicle control instruction from the target gesture control instruction and the target voice control instruction to control the target vehicle.
[0041] Specifically, the target vehicle is equipped with different vehicle functions, such as navigation function, entertainment function, air conditioning control function, seat adjustment function, window lifting function, etc. The target vehicle function is determined based on the vehicle control instruction, and the target vehicle is controlled based on the target vehicle function.
[0042] It should be noted that if only the gesture image is collected, the target gesture control instruction is determined as the vehicle control instruction. If only the voice information is collected, the target voice control instruction is determined as the vehicle control instruction.
[0043] The vehicle control method provided in this embodiment combines gesture interaction and voice interaction to control the target vehicle. Therefore, it can avoid the problem of inaccurate vehicle control caused by inaccurate single voice interaction to a certain extent. On this basis, if the identified target gesture control instruction conflicts with the target voice control instruction, the vehicle control instruction is determined based on the clarity of the target gesture control instruction and the target voice control instruction to control the target vehicle. Therefore, it can avoid the problem of vehicle control conflict caused by the simultaneous use of gesture interaction and voice interaction, thereby effectively improving the accuracy of vehicle control.
[0044] In some optional implementations, the above step S101 includes:
[0045] Step a1, obtaining first environmental information of the target vehicle, where the first environmental information includes lighting information.
[0046] It should be noted that, in addition to the illumination information, the first environmental information also includes other environmental information related to the image captured by the image acquisition device.
[0047] Step a2: adjusting the image acquisition parameters of the image acquisition device based on the first environmental information.
[0048] Optionally, the image acquisition parameters include acquisition parameters such as exposure time and gain.
[0049] Step a3, capturing gesture images through the adjusted image capture device.
[0050] The vehicle control method provided in this embodiment adjusts the image acquisition parameters of the image acquisition device based on the first environmental information including the illumination information, and uses the adjusted image acquisition device to acquire gesture images. Therefore, it is possible to capture clear gestures under various lighting conditions such as strong light, weak light, and backlight, ensuring the maximum utilization of the dynamic range of the gesture image, and providing high-quality gesture images for gesture recognition.
[0051] In some optional implementations, the above step S102 includes:
[0052] Step b1, obtaining preset gesture data, where the preset gesture data includes at least one first gesture feature corresponding to a gesture control instruction.
[0053] Specifically, a gesture sample and a corresponding gesture control instruction are obtained, a feature is extracted from the gesture sample to obtain a first gesture feature, and preset gesture data is constructed based on the matched first gesture feature and the gesture control instruction.
[0054] Step b2: extracting features from the gesture image based on the gesture recognition model to obtain a second gesture feature.
[0055] Specifically, the gesture recognition model is configured with a target feature extraction algorithm, such as a convolutional neural network (CNN). The gesture recognition model extracts features from the gesture image based on the feature extraction algorithm to obtain a second gesture feature. The target feature extraction algorithm learns the feature differences between different gestures by training a large number of gesture samples.
[0056] The second gesture feature includes at least one of gesture shape, color, and texture, or includes other gesture features, which are not limited here.
[0057] Step b3, matching the second gesture feature with the preset gesture data to determine the target gesture control instruction.
[0058] Specifically, matching is performed based on the second gesture feature and the first gesture feature in the preset gesture data to determine the matched first gesture feature, and based on the gesture control instruction corresponding to the matched first gesture feature, a target gesture control instruction is determined.
[0059] The vehicle control method provided in this embodiment matches the second gesture feature of the gesture image with the preset gesture data to determine the target gesture control instruction, thereby ensuring the accuracy of gesture recognition to improve the accuracy of the target gesture control instruction.
[0060] In some optional implementations, the vehicle control method further includes: acquiring a usage frequency of at least one gesture control instruction; and adjusting parameters of a gesture recognition model based on the usage frequency.
[0061] Specifically, historical interaction data of passengers is collected and analyzed to determine the usage frequency of each gesture control command.
[0062] In addition, data mining and statistical analysis algorithms can be used to analyze historical user interaction data, understand user preferences and common operations, and provide analytical data for gesture recognition model optimization. Based on the analytical data, deep learning algorithms are used to optimize the gesture recognition model to improve performance and can be customized based on user feedback.
[0063] The vehicle control method provided in this embodiment adjusts the parameters of the gesture recognition model based on the usage frequency of each gesture control instruction. Therefore, the gesture recognition model can understand the usage habits and preferences of the passengers and improve the gesture recognition performance of the gesture recognition model.
[0064] In some optional implementations, the above step S103 includes:
[0065] Step b1, obtaining second environmental information of the target vehicle, where the second environmental information includes background noise.
[0066] Optionally, the second environmental information also includes other environmental information related to speech recognition.
[0067] Optionally, the background noise includes at least one of engine hum, wind noise, and tire noise. In addition, the background noise may also include other noises that interfere with speech recognition.
[0068] Step b2: Analyze the second environment information to determine the target noise reduction parameters.
[0069] Specifically, the target noise reduction parameters are determined based on the analysis of the second environment information and the voice information.
[0070] Step b3, performing noise reduction processing on the speech information based on the target noise reduction parameter to obtain the noise-reduced speech information.
[0071] Specifically, a signal processing algorithm and a target noise reduction parameter are used to perform noise reduction processing on the speech information to obtain the noise-reduced speech information.
[0072] Step b4, performing speech recognition on the noise-reduced speech information based on the speech recognition model to obtain the target voice control instruction.
[0073] The vehicle control method provided in this embodiment analyzes the second environmental information and determines the target noise reduction parameters. Then, based on the target noise reduction parameters, the voice information is subjected to noise reduction processing to obtain the noise-reduced voice information. Therefore, the background noise can be identified and eliminated. Then, voice recognition is performed based on the noise-reduced voice information. Therefore, the influence of the background noise in the target vehicle on voice recognition can be reduced, ensuring that a stable voice recognition effect can be maintained in different noise environments.
[0074] In some optional implementations, the above step b4 includes:
[0075] Step b41, performing speech recognition on the noise-reduced speech information based on the speech recognition model to obtain speech features.
[0076] Among them, the speech recognition model uses a speech recognition algorithm to convert the noise-reduced speech information into text instructions to obtain speech features.
[0077] Step b42, obtaining historical interaction information.
[0078] Step b43, performing context analysis on the voice features based on the historical interaction information to obtain the target voice control instruction.
[0079] Specifically, natural language processing technology and historical interaction information are used to analyze the implicit information and contextual relationships in the voice features, so as to more accurately understand the real needs of the passengers and obtain the target voice control instructions. For example, when the passenger says "I'm hot", natural language processing technology and historical interaction information are used to identify that this is a preset voice control instruction to turn on the air conditioner, so the target voice control instruction to turn on the air conditioner is generated.
[0080] The vehicle control method disclosed in the present invention also supports a preset voice control command function, and users (including passengers) can set custom preset voice control commands for commonly used cockpit functions or services according to their preferences and habits. Context analysis is performed based on historical interaction information, preset voice control commands, and voice features, and preset voice control commands that match the voice features are determined in the preset voice control commands to obtain the target voice control commands. This personalized setting not only provides flexibility in vehicle control, but also makes voice interaction closer to the actual needs of users. For example, a user can change the voice command of "open the sunroof" to "I want to get some fresh air", and the vehicle control system can also recognize and understand this preset voice control command, and automatically perform the corresponding vehicle control operation.
[0081] The vehicle control method provided in this embodiment performs context analysis on the recognized voice features based on historical interaction information to obtain the target voice control command. Therefore, it is possible to deeply understand the user intention represented by the voice information, and combined with the above-mentioned noise reduction processing, it can achieve adaptability to complex voice environments and in-depth understanding of user intentions, making voice interaction more natural and smooth, and improving the accuracy of vehicle control and user experience.
[0082] In some optional implementations, the vehicle control method further includes: acquiring the instruction type of the historical voice control instruction; and adjusting the parameters of the voice recognition model based on the instruction type.
[0083] Specifically, historical interaction data of the user is collected and analyzed to determine the instruction types of historical voice control instructions.
[0084] In addition, data mining and statistical analysis algorithms can be used to analyze historical user interaction data, understand user preferences and common operations, and provide analytical data for speech recognition model optimization. Based on the analytical data, deep learning algorithms are used to optimize the speech recognition model to improve performance, and can be customized based on user feedback.
[0085] The vehicle control method provided in this embodiment adjusts the parameters of the speech recognition model based on the instruction type of historical voice control instructions. Therefore, the speech recognition model can understand the user's usage habits and preferences to improve the speech recognition performance of the speech recognition model.
[0086] In some optional embodiments, the voice information includes the first voice information and the second voice information, the target voice control instruction includes the first voice control instruction and the second voice control instruction, the first voice control instruction is obtained by voice recognition of the first voice information, and the second voice control instruction is obtained by voice recognition of the second voice information. The above-mentioned vehicle control method also includes: if there is a conflict between the first voice control instruction and the second voice control instruction, based on the first instruction information of the first voice control instruction and the second instruction information of the second voice control instruction, the vehicle control instruction is determined in the first voice control instruction and the second voice control instruction to control the target vehicle. Wherein, the first instruction information includes the signal strength of the first voice information, the first user intention represented by the first voice information, and at least one of the second user intention, and the second user intention of the first instruction information is obtained by context analysis based on the voice features of the first voice information and historical interaction information. The second instruction information includes the signal strength of the second voice information, the first user intention represented by the second voice information, and at least one of the second user intention, and the second user intention of the second instruction information is obtained by context analysis based on the voice features of the second voice information and historical interaction information.
[0087] The first user intent is the user intent directly expressed by the voice information, and the second user intent is the user intent expressed by the voice information understood by the vehicle control system based on the historical interaction information. For example, if the voice information is "I am hot", the first user intent indicates that the user is hot, and the second user intent indicates that the user wants to turn on the air conditioner.
[0088] Therefore, the signal strength, the first user intention and the second user intention can be comprehensively considered to ensure the accurate execution of the user intention and effectively avoid the problem of confused voice interaction caused by the user giving multiple voice messages at the same time, thereby effectively improving the stability of the vehicle control system and the user experience.
[0089] In some optional implementations, based on the first clarity and the second clarity, determining a vehicle control instruction from a target gesture control instruction and a target voice control instruction to control a target vehicle includes:
[0090] Step c1: if the first clarity is consistent with the second clarity, determining the priority between the target gesture control instruction and the target voice control instruction.
[0091] Specifically, if the first clarity is consistent with the second clarity, and both the first clarity and the second clarity are greater than a preset clarity threshold, the priority between the target gesture control instruction and the target voice control instruction is determined.
[0092] It is understandable that the priority between the target gesture control command and the target voice control command is determined only when both the target gesture control command and the target voice control command are clear. If the clarity of the target gesture control command and the target voice control command is low, feedback information is issued to inform the occupant to re-enter the voice information or make the gesture again.
[0093] Specifically, the priority between the target gesture control instruction and the target voice control instruction is determined according to a preset configuration or a user preference.
[0094] For example, if the preset configuration indicates that the priority of the gesture control instruction is higher than the priority of the voice control instruction, then it is determined that the priority of the target gesture control instruction is higher than the priority of the target voice control instruction.
[0095] Alternatively, the preset configuration includes priorities between gesture control commands and voice control commands in different operating scenarios. The current operating scenario is determined according to the operating environment information of the target vehicle. The priority between the target gesture control command and the target voice control command is determined by matching the current operating scenario with the preset configuration.
[0096] Step c2: determining the vehicle control instruction from the target gesture control instruction and the target voice control instruction based on the priority.
[0097] If the priority of the target gesture control command is higher than the priority of the target voice control command, the target gesture control command is determined as the vehicle control command.
[0098] If the priority of the target voice control command is higher than the priority of the target gesture control command, the target voice control command is determined as the vehicle control command.
[0099] Step c3, controlling the target vehicle based on the vehicle control instruction.
[0100] In practical applications, signal processing technology is used to synchronously analyze the target voice control command and the target gesture control command to determine the vehicle control command from the two.
[0101] In the vehicle control method provided in this embodiment, if the first clarity is consistent with the second clarity, the vehicle control instruction is determined based on the priority between the target gesture control instruction and the target voice control instruction. Therefore, when gesture interaction and voice interaction are simultaneously acting and causing confusion, a control instruction with a higher priority can be selected as the vehicle control instruction according to the current operating scenario, thereby improving the stability of vehicle control.
[0102] In some optional embodiments, after controlling the target vehicle, the above-mentioned vehicle control method also includes: obtaining a target control type of a vehicle control instruction and feedback preferences of multiple control types; determining a target feedback preference based on the target control type and feedback preferences of multiple control types; and controlling the target vehicle based on the target feedback preference.
[0103] Specifically, the feedback preference includes feedback modes of different interaction modes, such as visual feedback mode (such as display device), auditory feedback mode (such as audio playback device) or tactile feedback mode (such as vibrating seat). Therefore, different feedback modes can be used to provide instant feedback of control results.
[0104] If the target feedback preference is voice feedback, the speech synthesis model is used to generate voice feedback, and the audio playback device of the target vehicle is controlled to play the voice feedback to inform the occupants of the operation results or vehicle status. Among them, the speech synthesis model uses speech synthesis technology, and in the subsequent process, the speech synthesis model is continuously optimized to achieve a smooth voice interaction experience.
[0105] The vehicle control method provided in this embodiment controls the target vehicle based on the target feedback preference matching the target control type after controlling the target vehicle. Therefore, it is possible to provide personalized feedback methods according to different operation types and user preferences to enhance the user experience. For example, when executing navigation instructions, the display device will display route information and navigation prompts in real time. When adjusting the air-conditioning temperature, the audio playback device will emit a sound prompting the current temperature, so that the user can feel the interactive experience and interactive feedback in a timely manner.
[0106] In some optional implementations, after controlling the target vehicle, the vehicle control method further includes: acquiring a control state of the target vehicle; and triggering an alarm mechanism of the target vehicle if the control state indicates that the vehicle control instruction is executed abnormally.
[0107] Specifically, the vehicle control system is equipped with a status monitoring function, which can monitor the control status of each device in the cabin in real time, such as whether the air conditioner is turned on, whether the window is closed, etc., to ensure the normal operation of various vehicle functions. Once the vehicle control command execution is abnormal or the equipment fails, the alarm mechanism of the target vehicle is triggered.
[0108] In the vehicle control method provided in this embodiment, if the control state of the target vehicle indicates that the vehicle control instruction is executed abnormally, the alarm mechanism of the target vehicle is triggered, thereby ensuring the normal operation of the vehicle control.
[0109] In some optional embodiments, the occupant includes a driver. The vehicle control method further includes: acquiring a facial image of the driver; performing facial recognition based on the facial image to determine the driver's attention state; if the attention state indicates that the driver is not paying attention, controlling the target vehicle to issue a reminder message, the reminder message is used to control the target vehicle to remind the driver.
[0110] Specifically, the driver's facial image is collected by an image acquisition device. The driver's attention state is determined by using facial recognition technology and performing facial recognition on the facial image. If the attention state indicates that the driver is not paying attention, the target vehicle is controlled to issue a reminder message, such as by sound, vision, or trigger feedback to remind the driver.
[0111] Furthermore, facial recognition technology and driving safety monitoring systems can be integrated to monitor the driver's attention status and driving environment in real time to ensure driving safety. When distracted driving or potential danger is detected, the safety monitoring module can remind the driver through sound, vision or touch to reduce the risk of accidents.
[0112] The vehicle control method provided in this embodiment performs facial recognition based on facial images to determine the driver's attention state; if the attention state indicates that the driver is not paying attention, the target vehicle is controlled to issue a reminder message. Therefore, the driver or other passengers can be reminded in time when the driver is not paying attention, thereby improving driving safety.
[0113] Furthermore, the vehicle control method disclosed in the present invention supports customizing preset gesture data and preset voice control instructions according to personal preferences to enhance the interactive experience. At the same time, it supports users to save and load different configuration methods to adapt to different usage scenarios and usage requirements.
[0114] In actual application, performance evaluation algorithms can also be used to regularly monitor vehicle control and evaluate vehicle control performance, such as recognition accuracy, response time, etc., to ensure optimal vehicle control and adjust vehicle control strategies.
[0115] In some optional embodiments, the target vehicle is equipped with a communication module, which communicates with other electronic devices (such as mobile phones, smart watches, etc.) to achieve data synchronization and remote control functions between the target vehicle and other electronic devices. For example, the target vehicle can be remotely started through an application on a mobile phone, and the cabin settings can be adjusted.
[0116] The communication module can be connected to other electronic devices through wireless network or Bluetooth. While controlling the cockpit through other electronic devices, gesture operations of other electronic devices can also be accessed, such as "pinch" in smart watches, smart rings and other electronic devices, further increasing the interactive mode of gestures.
[0117] The vehicle control method disclosed in the present invention can significantly improve user experience and driving safety by integrating gesture and voice interaction. In a complex and changeable driving environment, users can flexibly choose gesture interaction or voice interaction according to actual conditions to achieve diversified interactive control of the vehicle cockpit and avoid the problem of inaccurate recognition caused by noise interference or sound interruption in single voice interaction. At the same time, it can detect and synchronously process gesture images and voice information in real time, automatically determine which interaction method has a higher priority according to preset configuration or user preference, and select the most appropriate response method in case of conflict to ensure the accurate execution of user intentions.
[0118] The vehicle control method disclosed in the present invention can achieve seamless connection with other electronic devices such as mobile phones and smart watches through a cross-platform communication module integrated in the vehicle control system, facilitate users to synchronize data and perform remote control, enhance the interconnection capabilities of the vehicle cabin, and improve the interactive experience.
[0119] The vehicle control method disclosed in the present invention uses an intelligent recommendation module to intelligently recommend suitable cabin functions and services based on user interaction history and preferences, thereby providing users with a more considerate and personalized service experience, while also helping users quickly find required functions and improve usage efficiency.
[0120] The gesture interaction and voice interaction in the vehicle control method disclosed in the present invention have excellent environmental adaptability. Through deep learning algorithms, gesture interaction can accurately segment, locate and recognize complex gesture movements. Voice interaction uses a voice recognition engine (refer to the above-mentioned voice recognition module) and noise suppression technology to ensure high-accuracy voice recognition in various environments in the car. Therefore, the gesture interaction and voice interaction disclosed in the present invention can automatically detect and adapt to different driving environments, lighting conditions and noise levels, ensuring a stable and accurate interactive experience in various situations. At the same time, the vehicle control system uses deep learning algorithms to continuously optimize the accuracy of gesture recognition and voice recognition to ensure rapid response and accurate execution of user commands.
[0121] The vehicle control method disclosed herein allows users to customize the interaction mode according to personal preferences, further enhancing the usability and personalized experience of the vehicle control system, and providing a more convenient, flexible, safe and personalized cockpit control solution.
[0122] In this embodiment, a vehicle control device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0123] This embodiment provides a vehicle control device, such as Figure 2 As shown, including:
[0124] The information acquisition module 201 is used to obtain gesture images and voice information of passengers in the target vehicle;
[0125] The gesture interaction module 202 is used to perform gesture recognition on the gesture image to obtain a target gesture control instruction;
[0126] The voice interaction module 203 is used to perform voice recognition on the voice information to obtain the target voice control instruction;
[0127] a conflict handling module 204, configured to obtain a first clarity of the target gesture control instruction and a second clarity of the target voice control instruction if there is a conflict between the target gesture control instruction and the target voice control instruction;
[0128] The vehicle control module 205 is used to determine a vehicle control instruction from a target gesture control instruction and a target voice control instruction based on the first clarity and the second clarity, so as to control the target vehicle.
[0129] In some optional embodiments, the vehicle control module 205 includes:
[0130] a priority calculation unit, configured to determine the priority between the target gesture control instruction and the target voice control instruction if the first clarity is consistent with the second clarity;
[0131] an instruction determination unit, configured to determine a vehicle control instruction from among a target gesture control instruction and a target voice control instruction based on a priority;
[0132] The vehicle control unit is used to control the target vehicle based on the vehicle control instruction.
[0133] In some optional implementations, the information collection module 201 includes:
[0134] A first information acquisition unit, used to acquire first environmental information of the target vehicle, the first environmental information including illumination information;
[0135] An acquisition parameter adjustment unit, configured to adjust an image acquisition parameter of the image acquisition device based on the first environmental information;
[0136] The image acquisition unit is used to acquire gesture images through the adjusted image acquisition device.
[0137] In some optional implementations, the gesture interaction module 202 includes:
[0138] A gesture data acquisition unit, used to acquire preset gesture data, wherein the preset gesture data includes at least one first gesture feature corresponding to a gesture control instruction;
[0139] A gesture feature extraction unit, used to extract features from the gesture image based on a gesture recognition model to obtain a second gesture feature;
[0140] The first instruction operation unit is used to match the second gesture feature with the preset gesture data to determine the target gesture control instruction.
[0141] In some optional embodiments, the vehicle control device of the present disclosure further includes:
[0142] A usage frequency acquisition module, used to acquire a usage frequency of at least one gesture control instruction;
[0143] The first model optimization module is used to adjust the parameters of the gesture recognition model based on the usage frequency.
[0144] In some optional implementations, the voice interaction module 203 includes:
[0145] A second information acquisition unit, used to acquire second environmental information of the target vehicle, the second environmental information including background noise;
[0146] A noise reduction parameter calculation unit, used to analyze the second environment information and determine a target noise reduction parameter;
[0147] A noise reduction processing unit, used to perform noise reduction processing on the voice information based on the target noise reduction parameter to obtain the noise-reduced voice information;
[0148] The second instruction operation unit is used to perform speech recognition on the noise-reduced speech information based on the speech recognition model to obtain the target speech control instruction.
[0149] In some optional implementations, the second instruction operation unit includes:
[0150] A speech feature extraction subunit is used to perform speech recognition on the noise-reduced speech information based on a speech recognition model to obtain speech features;
[0151] A historical information acquisition subunit, used to acquire historical interaction information;
[0152] The control instruction operation subunit is used to perform context analysis on the voice features based on the historical interaction information to obtain the target voice control instruction.
[0153] In some optional embodiments, the vehicle control device of the present disclosure further includes:
[0154] An instruction type acquisition module is used to acquire the instruction type of historical voice control instructions;
[0155] The second model optimization module is used to adjust the parameters of the speech recognition model based on the instruction type.
[0156] In some optional embodiments, the vehicle control device of the present disclosure further includes:
[0157] A feedback information acquisition module, used to acquire a target control type of a vehicle control instruction and feedback preferences of multiple control types after controlling the target vehicle;
[0158] determining a target feedback preference based on the target control type and the feedback preferences of the multiple control types;
[0159] The target vehicle is controlled based on the target feedback preference.
[0160] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0161] The vehicle control device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0162] The embodiment of the present disclosure further provides a vehicle, including: a vehicle body, such as an audio playback device, an air conditioner, a vehicle window, a voice acquisition device, an image acquisition device, and the like.
[0163] See also Figure 3 , Figure 3 is a structural block diagram of a vehicle controller provided by an optional embodiment of the present disclosure, such as Figure 3 As shown, the vehicle controller includes: one or more processors 301, a memory 302, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process instructions executed in the vehicle, including instructions stored in or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple vehicles can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 3 A processor 301 is taken as an example.
[0164] The processor 301 may be a central processing unit, a network processor or a combination thereof. The processor 301 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable logic gate array, a general purpose array logic or any combination thereof.
[0165] The memory 302 stores instructions executable by at least one processor 301 , so that the at least one processor 301 executes the method shown in the above embodiment.
[0166] The memory 302 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and applications required for at least one function; the data storage area may store data created according to the use of the vehicle, etc. In addition, the memory 302 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 302 may optionally include a memory remotely arranged relative to the processor 301, and these remote memories may be connected to the vehicle via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0167] The memory 302 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 302 may also include a combination of the above types of memory.
[0168] The vehicle controller also includes an input device 303 and an output device 304. The processor 301, the memory 302, the input device 303 and the output device 304 can be connected via a bus or other means. Figure 3 The example of connecting through bus is taken in the following.
[0169] The input device 303 can receive input digital or character information, and generate key signal input related to the user settings and function control of the vehicle, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator bar, one or more mouse buttons, a trackball, a joystick, etc. The output device 304 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device can be a touch screen.
[0170] See also Figure 4 , Figure 43 is a structural block diagram of a vehicle control system provided by an optional embodiment of the present disclosure. The processor 301 is deployed with a vehicle control system, and the vehicle control system is used to execute the above-mentioned vehicle control method. Among them, the vehicle control system includes a recommendation module, a learning and optimization module, a voice interaction module, a gesture interaction module, a fusion processing module, a cockpit application module, a communication module, a personalized configuration module, a safety monitoring module, and a user feedback module.
[0171] The gesture interaction module is used to capture the gestures of passengers in real time. The module pre-processes, extracts features and classifies the captured gesture images through deep learning algorithms to obtain gesture features. The module also converts the gesture features into control signals that can be recognized by the vehicle control system in combination with the preset gesture data to obtain the target gesture control instructions.
[0172] The voice interaction module integrates a voice recognition model and voice synthesis technology. The voice interaction module is used to receive the user's voice input and obtain voice information. In addition, the voice information is converted into text instructions using natural language processing technology to obtain the target voice control instructions. In addition, the voice synthesis technology is used to provide clear voice feedback to inform the vehicle control results or vehicle status.
[0173] The fusion processing module is used to detect the existence of the target gesture control instruction and the target voice control instruction in real time. And, in the case of a conflict between the target gesture control instruction and the target voice control instruction, determine the priority of the target gesture control instruction and the target voice control instruction according to a preset configuration or user preference. And, determine the vehicle control instruction from the target gesture control instruction and the target voice control instruction based on the priority.
[0174] The cockpit application module is used to execute various vehicle functions according to the vehicle control instructions output by the fusion processing module. At the same time, the cockpit application module also has a status monitoring function, which is used to monitor the operating status of each device in the cockpit in real time to ensure the normal function of the vehicle.
[0175] The user feedback module is used to provide instant feedback of the control results through visual, auditory or tactile feedback. In addition, according to different control types and feedback preferences, a target feedback preference is determined. The target vehicle is controlled according to the target feedback preference to provide a personalized feedback method.
[0176] The safety monitoring module is used to monitor the driver's attention status and driving environment in real time through facial recognition technology and driving safety monitoring system to ensure driving safety. When it detects that the driver is not paying attention or there is danger, it will send a reminder message through sound, visual or tactile feedback to remind the driver.
[0177] The personalized configuration module is used to support custom settings of preset gesture data and preset voice control commands to enhance the interactive experience. In addition, it supports users to save and load different configuration schemes to adapt to different usage scenarios and needs.
[0178] The learning and optimization module is used to optimize the gesture recognition model and the speech recognition model through machine learning algorithms and historical user interaction data, thereby improving the recognition accuracy and response speed of the gesture recognition model and the speech recognition model.
[0179] The communication module is used to connect with other electronic devices to achieve data synchronization and remote control functions.
[0180] The recommendation module is used to recommend suitable vehicle functions based on historical interaction data and user preferences, providing users with a personalized service experience.
[0181] The embodiments of the present disclosure also provide a computer-readable storage medium. The above-mentioned method according to the embodiments of the present disclosure can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium and downloaded through a network, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.
[0182] A part of the present disclosure may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present disclosure through the operation of the computer. Those skilled in the art should understand that the existence of computer program instructions in computer-readable media includes, but is not limited to, source files, executable files, installation package files, etc., and accordingly, the way in which computer program instructions are executed by a computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.
[0183] Although the embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A vehicle control method, characterized in that: The method comprises: Acquire gesture images and voice information of passengers in the target vehicle; Performing gesture recognition on the gesture image to obtain a target gesture control instruction; Performing voice recognition on the voice information to obtain a target voice control instruction; If there is a conflict between the target gesture control instruction and the target voice control instruction, obtaining a first clarity of the target gesture control instruction and a second clarity of the target voice control instruction; Based on the first clarity and the second clarity, a vehicle control instruction is determined from the target gesture control instruction and the target voice control instruction to control the target vehicle.
2. The vehicle control method according to claim 1, characterized in that: The determining a vehicle control instruction from the target gesture control instruction and the target voice control instruction based on the first clarity and the second clarity to control the target vehicle includes: If the first clarity is consistent with the second clarity, determining the priority between the target gesture control instruction and the target voice control instruction; Determining the vehicle control instruction from among the target gesture control instruction and the target voice control instruction based on the priority; The target vehicle is controlled based on the vehicle control instruction.
3. The vehicle control method according to claim 1, characterized in that: The step of acquiring a gesture image of an occupant in a target vehicle includes: Acquire first environmental information of the target vehicle, where the first environmental information includes lighting information; adjusting image acquisition parameters of an image acquisition device based on the first environmental information; The gesture image is captured by the adjusted image capture device.
4. The vehicle control method according to claim 1, characterized in that: The performing gesture recognition on the gesture image to obtain a target gesture control instruction includes: Acquire preset gesture data, where the preset gesture data includes at least one first gesture feature corresponding to a gesture control instruction; Extracting features from the gesture image based on a gesture recognition model to obtain a second gesture feature; The second gesture feature and the preset gesture data are matched to determine the target gesture control instruction.
5. The vehicle control method according to claim 4, characterized in that: The method further comprises: Acquire a usage frequency of the at least one gesture control instruction; Parameters of the gesture recognition model are adjusted based on the usage frequency.
6. The vehicle control method according to claim 1, characterized in that: The performing voice recognition on the voice information to obtain a target voice control instruction includes: Acquire second environmental information of the target vehicle, where the second environmental information includes background noise; Analyze the second environmental information to determine target noise reduction parameters; Performing noise reduction processing on the voice information based on the target noise reduction parameter to obtain noise-reduced voice information; Speech recognition is performed on the noise-reduced speech information based on a speech recognition model to obtain the target speech control instruction.
7. The vehicle control method according to claim 6, characterized in that: The performing speech recognition on the noise-reduced speech information based on the speech recognition model to obtain the target speech control instruction includes: Performing speech recognition on the noise-reduced speech information based on the speech recognition model to obtain speech features; Get historical interaction information; The voice feature is contextually analyzed based on the historical interaction information to obtain the target voice control instruction.
8. The vehicle control method according to claim 6 or 7, characterized in that: The method further comprises: Get the command type of historical voice control commands; Parameters of the speech recognition model are adjusted based on the instruction type.
9. The vehicle control method according to claim 1, characterized in that: After controlling the target vehicle, the method further includes: Acquire a target control type of the vehicle control command and feedback preferences of multiple control types; determining a target feedback preference based on the target control type and the feedback preferences of the plurality of control types; The target vehicle is controlled based on the target feedback preference.
10. A vehicle, characterized in that: The vehicle comprises: Vehicle body; A vehicle control system, the vehicle control system being configured to execute the vehicle control method according to any one of claims 1 to 9.
Citation Information
Cited By
Vehicle control method, device and equipment
CN120902762A