Vehicle intelligent control method and device, vehicle and storage medium

CN119559944BActive Publication Date: 2026-09-22DONGFENG MOTOR CO LTD DONGFENG NISSAN PASSENGER VEHICLE CO
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411740980.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-29
Publication Date
2026-09-22
Estimated Expiration
2044-11-29

AI Technical Summary

Technical Problem

[0003]本申请的主要目的在于提供一种车辆智能控制方法、装置、车辆及存储介质,旨在解决现有车辆控制需要结合TOF摄像头导致成本较高的技术问题

Benefits of technology

[0044]本申请提出的一个或多个技术方案,车辆智能控制方法应用于车辆,所述车辆上设有语音采集模块和图像采集模块,方法包括:通过所述语音采集模块确定车内的语音信息,并根据所述语音信息确定声源位置和语音指令;在所述语音指令为目标指令时,基于所述声源位置获取所述图像采集模块采集的车内图像数据;基于所述车内图像数据使用预设识别函数进行目标手势识别,得到识别结果;在所述识别结果为识别到目标手势时,根据所述目标手势、所述语音指令以及所述声源位置对车辆进行控制,本方案无需新增额外的深度信息,通过复用图像采集模块进行车内用户的手势识别,降低车辆内识别的硬件成本,提高识别的效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119559944B_ABST
    Figure CN119559944B_ABST
Patent Text Reader

Abstract

The application discloses a vehicle intelligent control method and device, a vehicle and a storage medium, and relates to the technical field of vehicle control. The method comprises the following steps: determining voice information in the vehicle through a voice acquisition module, and determining a sound source position and a voice instruction according to the voice information; when the voice instruction is a target instruction, acquiring vehicle image data collected by an image acquisition module based on the sound source position; performing target gesture recognition on the vehicle image data based on a preset recognition function to obtain a recognition result; and when the recognition result is a recognized target gesture, controlling the vehicle according to the target gesture, the voice instruction and the sound source position. The present scheme does not need to add additional depth information, gesture recognition of a user in the vehicle is performed by reusing the image acquisition module, the hardware cost of recognition in the vehicle is reduced, and the recognition effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle control technology, and in particular to vehicle intelligent control methods, devices, vehicles and storage media. Background Technology

[0002] Current in-vehicle cockpit interactions primarily rely on voice control, button control, and touchscreen control. This design differs from real-world social interaction, resulting in a lack of responsiveness in the human-computer interaction process. Therefore, to improve the user experience, current vehicle control utilizes TOF (Time of Fly) cameras for 3D gesture recognition, combining the recognized information with voice commands for intelligent vehicle control. However, this control method requires additional TOF cameras, increasing hardware costs. Summary of the Invention

[0003] The main purpose of this application is to provide a vehicle intelligent control method, device, vehicle and storage medium, which aims to solve the technical problem that the existing vehicle control requires the combination of TOF camera, resulting in high cost.

[0004] To achieve the above objectives, this application proposes a vehicle intelligent control method, which is applied to a vehicle equipped with a voice acquisition module and an image acquisition module. The method includes:

[0005] The voice acquisition module determines the voice information inside the vehicle, and determines the location of the sound source and the voice command based on the voice information;

[0006] When the voice command is a target command, the in-vehicle image data acquired by the image acquisition module is obtained based on the location of the sound source;

[0007] Based on the in-vehicle image data, a preset recognition function is used to perform target gesture recognition to obtain the recognition result;

[0008] When the recognition result indicates that the target gesture has been recognized, the vehicle is controlled based on the target gesture, the voice command, and the location of the sound source.

[0009] In one embodiment, the step of controlling the vehicle based on the target gesture, the voice command, and the sound source location when the recognition result indicates that a target gesture has been recognized includes:

[0010] When the recognition result indicates that the target gesture has been recognized, the gesture direction is determined based on the target gesture.

[0011] The target component is determined based on the direction of the gesture and the location of the sound source;

[0012] The target component is controlled according to the voice command.

[0013] In one embodiment, the step of controlling the target component according to the voice command includes:

[0014] Determine key instruction words based on the voice commands;

[0015] Obtain the status of vehicle components;

[0016] Determine the target action for the target component based on the key instruction words and the vehicle component status;

[0017] The vehicle is controlled according to the target action.

[0018] In one embodiment, before the step of performing target gesture recognition using a preset recognition function based on the in-vehicle image data to obtain the recognition result, the method further includes:

[0019] Historical hand data is acquired, and image learning is performed using the historical hand data to construct a two-dimensional hand model, which is then transformed into a first function.

[0020] Historical hand gesture data is obtained based on the first function and the historical hand data, and finger pointing gesture learning is performed based on the historical hand gesture data to construct a two-dimensional gesture recognition model. The two-dimensional gesture recognition model is then transformed into a second function.

[0021] Historical finger pointing gesture data is obtained based on the second function and the historical gesture posture data, and pointing direction learning is performed based on the historical finger pointing gesture data to construct a two-dimensional pointing model, and the two-dimensional pointing model is transformed into a third function;

[0022] The preset identification function is obtained through the third function.

[0023] In one embodiment, before the step of acquiring the in-vehicle image data acquired by the image acquisition module based on the sound source location when the voice command is a target command, the method further includes:

[0024] The voice command is parsed to determine whether it is a vehicle control command;

[0025] When the voice command is the vehicle control command, determine whether vehicle component information exists in the vehicle control command;

[0026] If vehicle component information is not present in the vehicle control command, determine whether the vehicle control command contains a preset pronoun;

[0027] When the preset pronoun exists in the vehicle control command, the voice command is determined to be the target command.

[0028] In one embodiment, after the step of performing target gesture recognition using a preset recognition function based on the in-vehicle image data to obtain the recognition result, the method further includes:

[0029] When the recognition result is that the target gesture is not recognized, a voice prompt message is generated;

[0030] The voice acquisition module provides the voice prompts and returns the voice information received in the vehicle, and determines the sound source location and voice command based on the voice information.

[0031] In one embodiment, the voice acquisition module includes a microphone array, which includes a plurality of microphones;

[0032] The steps for determining the location of a sound source based on the speech information include:

[0033] The time difference between each microphone in the microphone array is obtained based on the sound source localization strategy and the speech information;

[0034] Calculate the angle and distance between the sound source and the microphone array based on the time difference;

[0035] The location of the sound source is determined by the angle and the distance.

[0036] Furthermore, to achieve the above objectives, this application also proposes a vehicle intelligent control device, which includes:

[0037] The determination module is used to determine the voice information inside the vehicle through the voice acquisition module, and to determine the location of the sound source and the voice command based on the voice information;

[0038] The acquisition module is used to acquire in-vehicle image data based on the location of the sound source when the voice command is the target command;

[0039] The recognition module is used to perform target gesture recognition based on the in-vehicle image data using a preset recognition function to obtain the recognition result;

[0040] The control module is used to control the vehicle based on the target gesture, the voice command, and the location of the sound source when the recognition result indicates that the target gesture has been recognized.

[0041] In addition, to achieve the above objectives, this application also proposes a vehicle comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the vehicle intelligent control method described above.

[0042] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the vehicle intelligent control method described above.

[0043] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the vehicle intelligent control method described above.

[0044] This application proposes one or more technical solutions for a vehicle intelligent control method applied to a vehicle. The vehicle is equipped with a voice acquisition module and an image acquisition module. The method includes: determining in-vehicle voice information through the voice acquisition module, and determining the sound source location and voice command based on the voice information; when the voice command is a target command, acquiring in-vehicle image data acquired by the image acquisition module based on the sound source location; performing target gesture recognition using a preset recognition function based on the in-vehicle image data to obtain a recognition result; when the recognition result indicates that a target gesture has been recognized, controlling the vehicle based on the target gesture, the voice command, and the sound source location. This solution does not require additional depth information, and by reusing the image acquisition module for in-vehicle user gesture recognition, it reduces the hardware cost of in-vehicle recognition and improves the recognition effect. Attached Figure Description

[0045] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0046] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0047] Figure 1 This is a flowchart illustrating an embodiment of the vehicle intelligent control method of this application.

[0048] Figure 2 This is a schematic diagram of the distribution of in-vehicle components in one embodiment of the vehicle intelligent control method of this application;

[0049] Figure 3 This is a schematic diagram of the coordinates of a sound source inside a vehicle in one embodiment of the vehicle intelligent control method of this application;

[0050] Figure 4 This is a flowchart illustrating Embodiment 2 of the vehicle intelligent control method of this application;

[0051] Figure 5 This is a flowchart illustrating Embodiment 3 of the vehicle intelligent control method of this application;

[0052] Figure 6 This is a flowchart illustrating Embodiment 4 of the vehicle intelligent control method of this application;

[0053] Figure 7 This is a simplified flowchart of an embodiment of the vehicle intelligent control method of this application;

[0054] Figure 8 This is a schematic diagram of the module structure of the vehicle intelligent control device according to an embodiment of this application;

[0055] Figure 9 This is a schematic diagram of the vehicle structure of the hardware operating environment involved in the vehicle intelligent control method in this application embodiment.

[0056] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0057] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0058] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0059] The main solution of this application embodiment is as follows: the voice information inside the vehicle is determined by the voice acquisition module, and the location of the sound source and the voice command are determined based on the voice information; when the voice command is a target command, the in-vehicle image data is acquired by the image acquisition module based on the location of the sound source; the target gesture is recognized by a preset recognition function based on the in-vehicle image data, and the recognition result is obtained; when the recognition result is that the target gesture is recognized, the vehicle is controlled according to the target gesture, the voice command and the location of the sound source.

[0060] In existing technology, when a user activates the voice assistant, the array microphones in the cabin identify the voice initiator (user) using a sound source localization algorithm. The Time-of-Flight (TOF) camera then activates and focuses on identifying the voice initiator. After the user completes the voice command, the TOF camera emits continuous external light pulses and receives the light reflected back from the object using a sensor. By detecting the time of flight (round trip) of the light pulses, the distance to the target object is obtained. Then, an algorithm combines information from multiple distance points to construct 3D scene data and accurately identify the direction of the gesture. This requires adding a Time-of-Flight camera, increasing hardware costs, particularly the additional TOF sensor hardware. Therefore, it lacks the capability for OTA (Over-The-Air) upgrades for existing vehicle models.

[0061] This application provides a solution that eliminates the need for a new TOF camera, instead reusing an OMS (Occupant Monitor System) camera for gesture recognition. Similar functionality is achieved through software without increasing hardware costs, and existing vehicles can be upgraded via OTA (Over-The-Air).

[0062] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device capable of performing the above functions, such as an in-vehicle infotainment system. The following description uses an in-vehicle infotainment system as an example to illustrate this embodiment and the subsequent embodiments.

[0063] Based on this, embodiments of this application provide a vehicle intelligent control method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the vehicle intelligent control method of this application.

[0064] In this embodiment, the vehicle intelligent control method is applied to a vehicle, which is equipped with a voice acquisition module and an image acquisition module. The vehicle intelligent control method includes steps S10 to S40:

[0065] Step S10: Determine the voice information inside the vehicle through the voice acquisition module, and determine the location of the sound source and the voice command based on the voice information.

[0066] It should be noted that the vehicle mainly includes an in-vehicle infotainment system, a voice acquisition module, and an in-cabin monitoring system (OMS). The in-vehicle infotainment system may include an OMS control module, a voice recognition module, and a vehicle control module, etc. The OMS can be installed on the vehicle's windshield or in the vehicle's central control system. The voice acquisition module includes multiple microphones to collect the voice information of the users inside the vehicle. The voice acquisition module may include a microphone array, which includes multiple microphones, such as... Figure 2 As shown, Figure 2This is a schematic diagram of the distribution of components inside the vehicle, including microphones (microphone 1, microphone 2, microphone 3 and microphone 4) located in various positions inside the vehicle, an OMS camera, a vehicle infotainment system, and a T-BOX (wireless gateway) located on the center console.

[0067] In practice, once the vehicle is started, all microphones in the voice acquisition module are activated and continuously collect audio signals until the vehicle is powered off. Each microphone collects audio signals and sends them to the voice recognition module via A2B (Automotive AudioBus) technology. When the user activates the voice recognition function, the voice recognition module determines the location of the sound source through internal links and sends the location to the OMS control module.

[0068] Therefore, voice information inside the vehicle can be collected through a microphone and analyzed to determine the location of the sound source and voice commands.

[0069] In one feasible implementation, step S10, "determining the location of the sound source based on the voice information," may include steps A11 to A13:

[0070] Step A11: Obtain the time difference between each microphone in the microphone array based on the sound source localization strategy and the voice information.

[0071] It is understandable that the sound source localization strategy is the TDOA (Time Difference of Arrival) sound source localization technology, which obtains the time difference between the arrival of the voice signal at each microphone by using TDOA and the collected voice information.

[0072] Step A12: Calculate the angle and distance between the sound source and the microphone array based on the time difference.

[0073] In practice, the angle and distance between the sound source inside the vehicle and the microphone array can be calculated based on the time difference. Specifically, this may include the angle and distance between the sound source and each microphone.

[0074] Step A13: Determine the location of the sound source using the angle and the distance.

[0075] It should be understood that the location of the sound source inside the vehicle can be determined based on the specific angle and distance, such as... Figure 3 As shown, Figure 3 This is a schematic diagram of the sound source coordinates inside the vehicle. This embodiment uses a four-seater vehicle as an example for illustration. The sound source coordinates are X_11, X_12, X_13, and X_14.

[0076] After sound source localization, image data can be collected by driving OMS through sound source localization.

[0077] Step S20: When the voice command is the target command, obtain the in-vehicle image data collected by the image acquisition module based on the sound source location.

[0078] Understandably, after receiving a voice command, semantic analysis can be performed on the voice command to determine the specific meaning of the command. If the target command is a vehicle control command, contains pronouns, and does not contain specific information about in-vehicle components, the voice command can be analyzed to determine whether it is the target command.

[0079] When the voice command is not the target command, the voice command can be sent out directly, and the command can be forwarded and verified, thereby controlling the vehicle through the voice command.

[0080] If the voice command is a target command, that is, in addition to voice control, it is also necessary to combine the user's actions, so that the vehicle can be controlled based on the combined actions and voice. Therefore, the in-vehicle image data collected by the image acquisition module can be obtained based on the location of the sound source.

[0081] It should be noted that the in-vehicle image data refers to the image data acquired by the image acquisition module at the location of the sound source. The image data acquired by the image acquisition module can be filtered according to the location of the sound source to obtain the in-vehicle image data.

[0082] Step S30: Based on the in-vehicle image data, use a preset recognition function to perform target gesture recognition and obtain the recognition result.

[0083] It should be understood that the preset recognition function can be a function that recognizes the hands, gestures, and directions of users inside the vehicle. After determining the in-vehicle image data corresponding to the location of the sound source, it can further determine whether the user in the image makes a gesture and the specific information of the gesture, thereby obtaining the recognition result.

[0084] It should be noted that the recognition result may be that the target gesture was recognized or not. The target gesture is a valid gesture, which is a gesture with a specific direction or a general direction.

[0085] It should be understood that if the target gesture is not recognized, it needs to be handled promptly, such as by providing a reminder or attempting to recognize it again, to avoid recognition errors.

[0086] In one feasible implementation, step S30 is followed by steps S31 to S32:

[0087] Step S31: When the recognition result is that the target gesture is not recognized, generate a voice prompt message.

[0088] Understandably, if the recognition result is that no valid gesture was recognized, it may be because the user did not make a gesture or the user made a gesture but it was not captured by the camera. Therefore, a voice prompt message can be generated, such as: "Your gesture was not recognized. Please make the gesture again" or "Your gesture was not recognized. Please confirm whether to continue."

[0089] Step S32: Prompt the voice prompt information through the voice acquisition module, and return the voice information that the voice acquisition module responds to in the vehicle, and determine the sound source location and voice command based on the voice information.

[0090] In practice, voice prompts can be sent to a microphone array or the microphone corresponding to the sound source location within the microphone array to determine if the user wishes to continue. If the user continues, the user's audio data can be collected again to obtain the sound source location and voice command, and further determine whether a valid gesture has been recognized. If the user requests to stop, the interaction ends.

[0091] Step S40: When the recognition result indicates that the target gesture has been recognized, the vehicle is controlled according to the target gesture, the voice command, and the location of the sound source.

[0092] It should be noted that if the recognition result is that a valid gesture of the user is recognized, the vehicle can be controlled based on the recognized valid gesture, the user's voice command, and the location of the sound source. This control can be applied to specific components in the vehicle or to components near the user at the location of the sound source.

[0093] For example, the user's intention can be determined by recognizing the user's valid gestures, and the target control component can be determined based on the user's voice command, the user's intention, and the location of the sound source, thereby controlling the target control component.

[0094] This embodiment provides a vehicle intelligent control method applied to a vehicle equipped with a voice acquisition module and an image acquisition module. The method includes: determining in-vehicle voice information through the voice acquisition module, and determining the sound source location and voice command based on the voice information; when the voice command is a target command, acquiring in-vehicle image data acquired by the image acquisition module based on the sound source location; performing target gesture recognition using a preset recognition function based on the in-vehicle image data to obtain a recognition result; when the recognition result indicates that a target gesture has been recognized, controlling the vehicle based on the target gesture, the voice command, and the sound source location. This solution does not require additional depth information, and by reusing the image acquisition module for in-vehicle user gesture recognition, it reduces the hardware cost of in-vehicle recognition and improves the recognition effect.

[0095] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 4 Step S40 includes steps S401 to S403:

[0096] Step S401: When the recognition result is that the target gesture has been recognized, determine the gesture direction according to the target gesture.

[0097] It should be noted that if the recognition result indicates that the target gesture has been detected, it proves that the user has a gesture command, and the gesture direction is obtained based on the user's valid gesture. The gesture direction is determined by analyzing the user's gesture direction through in-vehicle image data. Specifically, when recognizing the user's gesture using a preset recognition function, the user's gesture direction can be directly identified.

[0098] In practice, the direction of the gesture can be left, right, up, or down at the location of the sound source.

[0099] Step S402: Determine the target component based on the gesture direction and the sound source location.

[0100] In practice, the components inside the vehicle can be divided in advance to obtain a component set {Target_1}. {Target_1} consists of all identifiable vehicle components, including but not limited to: windows in various locations, sunroof, air conditioning, ambient lighting, trunk, fragrance, seat heating, seat ventilation, seat massage, etc.

[0101] It should be noted that after determining the direction of the gesture, the target component can be determined based on the location of the sound source and the direction of the gesture. The target component is the component to be controlled.

[0102] For example, if the hand gesture is to the left and the sound source is located at X_11, then the target component is determined to be the passenger side window. If the hand gesture is to the right and the sound source is located at X_21, then the target component is determined to be the rear left side window, as shown in Table 1 below. Table 1 is a list of components inside the vehicle.

[0103] Table 1

[0104] T_11 Driver's window T_51 trunk T_12 Passenger window T_61 Fragrance T_13 Rear left side window (driver's side rear window) T_71 Driver's seat heating T_14 Rear right-side window (passenger side rear window) T_72 passenger seat heating T_21 skylight T_73 Rear left seat heating T_31 Front air conditioning lower vent T_74 Rear right seat heating T_32 Rear air conditioning lower vent T_81 Driver's seat ventilation T_33 Driver's side air conditioning vents T_82 Passenger seat ventilation T_34 Passenger side air conditioning vent T_83 Rear left seat ventilation T_35 Rear left air conditioning vent T_84 Rear right seat ventilation T_36 Rear right air conditioning vent T_91 Driver's seat massage T_37 Car refrigerator electric door T_92 passenger seat massage T_41 Ambient Lighting T_93 Rear left seat massage T_42 Central reading light T_94 Rear right seat massage T_43 Rear reading lights

[0105] As shown in Tables 2 and 3, Tables 2 and 3 are tables of target components corresponding to gesture directions and sound source locations.

[0106] Table 2

[0107]

[0108] Table 3

[0109]

[0110] Step S403: Control the target component according to the voice command.

[0111] Understandably, specific components can be controlled based on the user's voice commands. For example, if the voice command is "on", the target component will be turned on; if the voice command is "off", the target component will be turned off.

[0112] In one feasible implementation, step S403 may include steps B11 to B14:

[0113] Step B11: Determine the key instruction words based on the voice command.

[0114] In practice, voice commands can be analyzed to identify key command words, such as "turn on," "open," "close," and "shut down."

[0115] Step B12: Obtain the status of vehicle components.

[0116] It should be understood that by obtaining the status of all components inside the vehicle, such as if a vehicle component is in an off state, if the user's intention is to turn off this component, no further action is required. Therefore, the user's instructions can be responded to more quickly based on the status of vehicle components.

[0117] Step B13: Determine the target action for the target component based on the key instruction words and the vehicle component status.

[0118] It should be noted that whether control of the target component is needed can be determined based on the specific command and the status of the vehicle component. For example, if the key command is "open," the target component is the sunroof, and the sunroof's status in the vehicle component status is "fully open," then the target action for the sunroof is no action, but a voice prompt can be generated stating that the sunroof is open. If the key command is "open," the target component is the sunroof, and the sunroof's status in the vehicle component status is "not fully open" or "closed," then the target action is to open the sunroof.

[0119] Step B14: Control the vehicle according to the target action.

[0120] In practice, after determining the target action, the vehicle can be controlled according to the target action to complete the interaction with the user.

[0121] Table 4

[0122]

[0123] As shown in Table 4, Table 4 is a relationship table between key command words, vehicle component status, gesture direction, sound source location, indicated component, target component, and target action. For example, if the key command words are "open" and "start," the gesture direction is left, the sound source location is X_11, and the target component is the passenger side window, then if the target component status is closed, the target action is to open the passenger side window; if the target component status is that the passenger side window is not fully open, the target action is to open the passenger side window; if the target component status is that the passenger side window is fully open and the passenger side air conditioning vent is open, the target action is a voice prompt that the passenger side window is open. If the target component is the passenger side air conditioning vent, and the target component status is that the passenger side window is fully open and the passenger side air conditioning vent is closed, the target action is to open the passenger side air conditioning vent.

[0124] In this embodiment, when the recognition result indicates that a target gesture has been detected, the gesture direction is determined based on the target gesture; the target component is determined based on the gesture direction and the sound source location; and the target component is controlled according to the voice command. By quickly determining the gesture direction through the target gesture, and then identifying the target component to be controlled based on the gesture direction and the sound source location, the control effect is improved.

[0125] Based on the first embodiment of this application, in the third embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 5 Before step S30, the vehicle intelligent control method further includes steps S21 to S24:

[0126] Step S21: Obtain historical hand data, perform image learning using the historical hand data, construct a two-dimensional hand model, and convert the two-dimensional hand model into a first function.

[0127] It should be noted that, to improve the accuracy and efficiency of gesture recognition, visual training can be performed using an OMS camera, enabling rapid recognition of two-dimensional hand models, two-dimensional gestures, and two-dimensional finger pointing. The main content of OMS visual training is the recognition of valid gestures.

[0128] Therefore, during training, historical hand data can be collected first. This historical hand data includes hand data from different age groups, genders, and ethnicities, allowing for image learning from this historical hand data. The learning objective is to be able to recognize 10,000 samples using an OMS camera with a recognition rate of 99.5%, thereby constructing a two-dimensional hand model. The recognition of this vision-trained two-dimensional hand model is then transformed into the function f_hand(), which is the first function.

[0129] Step S22: Obtain historical gesture posture data based on the first function and the historical hand data, and perform finger pointing gesture learning based on the historical gesture posture data to construct a two-dimensional gesture recognition model, and convert the two-dimensional gesture recognition model into a second function.

[0130] In practical implementation, different hand gesture data can be identified based on the function `f_hand()`, thus obtaining historical hand gesture data. This historical data includes various hand gestures. Image learning is then performed on these various hand gestures, with a focus on finger pointing gestures. The learning objective is to be able to recognize 10,000 samples using an OMS camera, achieving a recognition rate of 99.5% and a false recognition rate of less than 0.5%, thereby training and constructing a two-dimensional hand gesture recognition model. This two-dimensional hand gesture recognition model is then transformed into the second function `f_gesture()`.

[0131] Step S23: Obtain historical finger pointing gesture data based on the second function and the historical gesture posture data, and perform pointing direction learning based on the historical finger pointing gesture data to construct a two-dimensional pointing model, and convert the two-dimensional pointing model into a third function.

[0132] In practical implementation, the pointing direction of finger gestures can be learned based on the second function `f_gesture()`. This involves obtaining different finger pointing gesture data from historical gesture data and the second function, thereby learning the pointing direction of the finger gestures. The learning range is limited to four directions: up, down, left, and right. The learning objective is to achieve a 99% recognition rate and a false recognition rate of less than 0.5% on 10,000 samples obtained through an OMS camera, constructing a two-dimensional pointing model, and then converting the visually trained two-dimensional finger pointing model into the third function `f_2Ddirection()`.

[0133] Step S24: Obtain the preset identification function through the third function.

[0134] In practical implementation, the third function can be flashed into the OMS control module when the vehicle's infotainment system is upgraded via OTA through the T-BOX. When image data is acquired from the OMS camera, the acquired data is compared with the function f_2Ddirection(). If they match, it is judged as yes, and the next judgment logic is entered; if they do not match, it is judged as no, and a voice prompt "I did not see your gesture, please try again" is given.

[0135] This embodiment acquires historical hand data and performs image learning based on this data to construct a two-dimensional hand model, which is then transformed into a first function. Based on the first function and the historical hand data, historical gesture posture data is obtained, and finger pointing gestures are learned from this data to construct a two-dimensional gesture recognition model, which is then transformed into a second function. Based on the second function and the historical gesture posture data, historical finger pointing gesture data is obtained, and pointing direction is learned from this data to construct a two-dimensional pointing model, which is then transformed into a third function. A preset recognition function is obtained through the third function. Visual training is performed using an OMS camera, focusing on three main aspects: two-dimensional hand model recognition, two-dimensional gesture recognition, and two-dimensional finger pointing (up, down, left, right) recognition. This allows for rapid recognition of user gestures using the preset recognition function, improving the response speed of vehicle interior component control.

[0136] Based on the first embodiment of this application, in the fourth embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 6 Before step S20, the vehicle intelligent control method further includes steps S11 to S14:

[0137] Step S11: Parse the voice command to determine whether the voice command is a vehicle control command.

[0138] It should be noted that voice commands can be parsed according to certain parsing rules, such as keyword splitting and keyword extraction, to determine whether the voice command is a vehicle control command. The vehicle control commands are set {B_1}, which are voice control commands classified as vehicle control commands. The set {A_1} can be written into the voice recognition module in advance. The set {A_1} contains all recognizable valid functional voice commands (functional voice commands include: vehicle control, multimedia, telephone, navigation, AI chat, etc.). Among them, the subset {B_1} of {A_1} are voice control commands classified as vehicle control commands.

[0139] If the voice command is not a vehicle control command, then determine whether it is another voice function command. If so, then directly control other voice functions according to the voice command.

[0140] Step S12: When the voice command is the vehicle control command, determine whether there is vehicle component information in the vehicle control command.

[0141] It should be understood that if the voice command is a vehicle control command, it can be further determined whether there is vehicle component information in the vehicle control command. If vehicle component information exists, the specific component can be controlled directly according to the vehicle control command without the need for a subsequent gesture recognition process.

[0142] For example, if the voice command is "Please open the driver's side window", then the voice command contains vehicle component information for the driver's side window, and the driver's side window can be opened directly according to this voice command.

[0143] Step S13: When there is no vehicle component information in the vehicle control command, determine whether there is a preset referent in the vehicle control command.

[0144] Understandably, if the vehicle control command does not contain vehicle component information, it is necessary to further determine whether the vehicle control command contains preset pronouns. Preset pronouns include, but are not limited to, pronouns such as "this," "that," "here," "there," and "it." A set {Pr_1} can be written into the speech recognition module in advance, and Pr_1} contains all the preset pronouns that can be recognized.

[0145] When the voice command matches the vehicle control command result (Y), the logic 2 for determining whether vehicle component information exists in the command can proceed. If the voice control command matches the set {Target_1}, it is judged as N, and the specific vehicle component is directly controlled according to this voice control command. If the voice control command matches the set {Pr_1} but does not contain the set {Target_1}, it is judged as Y, and the next judgment logic is entered. Otherwise, it is considered invalid voice information, the interaction fails, and the voice terminal blocks the recognition of this command.

[0146] Step S14: When the preset pronoun exists in the vehicle control command, determine that the voice command is the target command.

[0147] In practice, if there is a preset pronoun in the vehicle control command, the voice command is determined to be the target command, and the next judgment logic can be entered, namely whether the target gesture is recognized.

[0148] This embodiment parses the voice command to determine whether it is a vehicle control command; if it is, it determines whether vehicle component information exists within the command; if no vehicle component information exists, it determines whether a preset pronoun exists; if the preset pronoun exists, it determines that the voice command is a target command. This allows for rapid determination of whether a voice command is a target command based on specific recognition rules. If the voice command is not a target command, subsequent gesture recognition is unnecessary, improving the response speed to user voice commands.

[0149] For example, to help understand the implementation flow of the vehicle intelligent control method obtained by combining this embodiment with the above embodiment one, please refer to... Figure 7 , Figure 7 A simplified flowchart of a vehicle intelligent control method is provided, specifically: 01: During voice wake-up; 02: Audio data collection; 03: Sound source localization based on the collected audio data; after successful sound source localization, 13: OMS driving; 14: Image data collection via OMS; simultaneously, after sound source localization, 04: Receiving commands; 05: Semantic parsing; 06: Entering judgment logic 1: Is it a vehicle control command?; 07: If not, identify it as another voice function; 08: If yes, enter judgment logic 2: Identify whether the vehicle control command lacks component information and whether it contains pronouns; 09: If component information exists or pronouns are absent, confirm the command; 10: Vehicle control command is sent out; 11: 12: Command forwarding and verification; 19: Command execution; 20: Voice prompt result; 21: Enter judgment logic 4: Does the user request to continue operation?; 22: If no, the interaction ends; if yes, return to step 02: Audio data collection steps; If no component information exists and a pronoun exists, then 15: Enter judgment logic 3; 16: Identify whether there is a valid gesture based on the collected image data; 17: If there is a valid gesture, perform gesture pointing guessing, and determine the command based on the combination of gesture and voice information, i.e., enter step 09, thereby performing the process of vehicle control command outgoing, command forwarding, verification and command execution; 18: If no valid gesture is identified, 19: Voice prompt: No valid gesture is identified; 20: Enter judgment logic 4.

[0150] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the vehicle intelligent control method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0151] This application also provides a vehicle intelligent control device, please refer to... Figure 8 The vehicle intelligent control device includes:

[0152] The determination module 10 is used to determine the voice information inside the vehicle through the voice acquisition module, and to determine the location of the sound source and the voice command based on the voice information.

[0153] The acquisition module 20 is used to acquire in-vehicle image data based on the location of the sound source when the voice command is the target command.

[0154] The recognition module 30 is used to perform target gesture recognition based on the in-vehicle image data using a preset recognition function to obtain the recognition result.

[0155] The control module 40 is used to control the vehicle based on the target gesture, the voice command, and the location of the sound source when the recognition result is that the target gesture has been recognized.

[0156] The vehicle intelligent control device provided in this application, employing the vehicle intelligent control method described in the above embodiments, can solve the technical problem of high cost caused by the need for existing vehicle control systems to incorporate TOF cameras. Compared with the prior art, the beneficial effects of the vehicle intelligent control device provided in this application are the same as those of the vehicle intelligent control method provided in the above embodiments, and other technical features in the vehicle intelligent control device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0157] In one embodiment, the control module 40 is further configured to, when the recognition result indicates that a target gesture has been recognized, determine the gesture direction based on the target gesture; determine the target component based on the gesture direction and the sound source location; and control the target component based on the voice command.

[0158] In one embodiment, the control module 40 is further configured to: determine key instruction words based on the voice command; acquire the status of vehicle components; determine a target action for the target component based on the key instruction words and the status of the vehicle components; and control the vehicle based on the target action.

[0159] In one embodiment, the recognition module 30 is further configured to acquire historical hand data, perform image learning using the historical hand data to construct a two-dimensional hand model, and convert the two-dimensional hand model into a first function; obtain historical gesture posture data based on the first function and the historical hand data, and perform finger pointing gesture learning based on the historical gesture posture data to construct a two-dimensional gesture recognition model, and convert the two-dimensional gesture recognition model into a second function; obtain historical finger pointing gesture data based on the second function and the historical gesture posture data, and perform pointing direction learning based on the historical finger pointing gesture data to construct a two-dimensional pointing model, and convert the two-dimensional pointing model into a third function; and obtain a preset recognition function through the third function.

[0160] In one embodiment, the acquisition module 20 is further configured to parse the voice command to determine whether the voice command is a vehicle control command; when the voice command is a vehicle control command, to determine whether vehicle component information exists in the vehicle control command; when vehicle component information does not exist in the vehicle control command, to determine whether a preset pronoun exists in the vehicle control command; and when the preset pronoun exists in the vehicle control command, to determine that the voice command is a target command.

[0161] In one embodiment, the recognition module 30 is further configured to generate voice prompt information when the recognition result is that the target gesture is not recognized; to prompt the voice prompt information through the voice acquisition module; and to return the steps of responding to the voice information in the vehicle through the voice acquisition module and determining the sound source location and voice command based on the voice information.

[0162] In one embodiment, the voice acquisition module includes a microphone array, which includes multiple microphones; the determining module 10 is further configured to determine the location of a sound source based on the voice information, including: obtaining the time difference between each microphone in the microphone array based on a sound source localization strategy and the voice information; calculating the angle and distance of the sound source from the microphone array based on the time difference; and determining the location of the sound source based on the angle and the distance.

[0163] This application provides a vehicle, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the vehicle intelligent control method in Embodiment 1 above.

[0164] The following is for reference. Figure 9 The diagram illustrates a structural schematic of a vehicle suitable for implementing embodiments of this application. The vehicle in these embodiments may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 9 The vehicle shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of this application.

[0165] like Figure 9As shown, the vehicle may include a processing unit 1001 (e.g., a central processing unit, a graphics processor, etc.) that can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for vehicle operation. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the vehicle to communicate wirelessly or wiredly with other devices to exchange data. Although vehicles with various systems are shown in the figures, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.

[0166] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0167] The vehicle provided in this application, employing the intelligent vehicle control method described in the above embodiments, can solve the technical problem of high costs caused by the need to combine existing vehicle control with a TOF camera. Compared with the prior art, the beneficial effects of the vehicle provided in this application are the same as those of the intelligent vehicle control method provided in the above embodiments, and other technical features of the vehicle are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0168] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0169] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0170] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the vehicle intelligent control method in the above embodiments.

[0171] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0172] The aforementioned computer-readable storage medium may be included in the vehicle or may exist independently and not installed in the vehicle.

[0173] The aforementioned computer-readable storage medium carries one or more programs that, when executed by a vehicle, cause the vehicle to: determine in-vehicle voice information via a voice acquisition module, and determine the location of the sound source and a voice command based on the voice information; when the voice command is a target command, acquire in-vehicle image data based on the sound source location via an image acquisition module; perform target gesture recognition using a preset recognition function based on the in-vehicle image data to obtain a recognition result; and when the recognition result indicates that a target gesture has been recognized, control the vehicle based on the target gesture, the voice command, and the sound source location.

[0174] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0175] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0176] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0177] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., computer programs) for executing the above-described vehicle intelligent control method. This solves the technical problem of high costs associated with existing vehicle control systems that require the use of TOF cameras. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the vehicle intelligent control method provided in the above embodiments, and will not be elaborated upon here.

[0178] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the vehicle intelligent control method described above.

[0179] The computer program product provided in this application can solve the technical problem that the existing vehicle control requires the use of a TOF camera, resulting in high costs. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the vehicle intelligent control method provided in the above embodiments, and will not be repeated here.

[0180] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A vehicle intelligent control method, characterized in that, The vehicle intelligent control method is applied to a vehicle, which is equipped with a voice acquisition module and an image acquisition module. The method includes: The voice acquisition module determines the voice information inside the vehicle, and determines the location of the sound source and the voice command based on the voice information; When the voice command is a target command, the in-vehicle image data collected by the image acquisition module is obtained based on the location of the sound source. The target command is a vehicle control command that contains a pronoun and does not contain specific in-vehicle component information. Based on the in-vehicle image data, a preset recognition function is used to perform target gesture recognition to obtain the recognition result; When the recognition result indicates that the target gesture has been recognized, the vehicle is controlled based on the target gesture, the voice command, and the location of the sound source. Before the step of obtaining the in-vehicle image data acquired by the image acquisition module based on the sound source location when the voice command is a target command, the method further includes: The voice command is parsed to determine whether it is a vehicle control command; When the voice command is the vehicle control command, determine whether vehicle component information exists in the vehicle control command; If vehicle component information is not present in the vehicle control command, determine whether the vehicle control command contains a preset pronoun; When the preset pronoun exists in the vehicle control command, the voice command is determined to be the target command; Before the step of performing target gesture recognition using a preset recognition function based on the in-vehicle image data to obtain the recognition result, the method further includes: Historical hand data is acquired, and image learning is performed using the historical hand data to construct a two-dimensional hand model, which is then transformed into a first function. Historical hand gesture data is obtained based on the first function and the historical hand data, and finger pointing gesture learning is performed based on the historical hand gesture data to construct a two-dimensional gesture recognition model. The two-dimensional gesture recognition model is then transformed into a second function. Historical finger pointing gesture data is obtained based on the second function and the historical gesture posture data, and pointing direction learning is performed based on the historical finger pointing gesture data to construct a two-dimensional pointing model, and the two-dimensional pointing model is transformed into a third function; The preset recognition function is obtained through the third function; The voice acquisition module includes a microphone array, which includes multiple microphones; The steps for determining the location of a sound source based on the speech information include: The time difference between each microphone in the microphone array is obtained based on the sound source localization strategy and the speech information; Calculate the angle and distance between the sound source and the microphone array based on the time difference; The location of the sound source is determined by the angle and the distance.

2. The method as described in claim 1, characterized in that, The step of controlling the vehicle based on the target gesture, the voice command, and the sound source location when the recognition result indicates that the target gesture has been recognized includes: When the recognition result indicates that the target gesture has been recognized, the gesture direction is determined based on the target gesture. The target component is determined based on the direction of the gesture and the location of the sound source; The target component is controlled according to the voice command.

3. The method as described in claim 2, characterized in that, The step of controlling the target component according to the voice command includes: Determine key instruction words based on the voice commands; Obtain the status of vehicle components; Determine the target action for the target component based on the key instruction words and the vehicle component status; The vehicle is controlled according to the target action.

4. The method as described in claim 1, characterized in that, After the step of performing target gesture recognition using a preset recognition function based on the in-vehicle image data to obtain the recognition result, the method further includes: When the recognition result is that the target gesture is not recognized, a voice prompt message is generated; The voice acquisition module provides the voice prompts and returns the voice information received in the vehicle, and determines the sound source location and voice command based on the voice information.

5. A vehicle intelligent control device, characterized in that, The device includes: The determination module is used to determine the voice information inside the vehicle through the voice acquisition module, and to determine the location of the sound source and the voice command based on the voice information; The acquisition module is used to acquire in-vehicle image data based on the location of the sound source when the voice command is the target command. The target command is a vehicle control command that contains a pronoun and does not contain specific in-vehicle component information. The recognition module is used to perform target gesture recognition based on the in-vehicle image data using a preset recognition function to obtain the recognition result; The control module is used to control the vehicle based on the target gesture, the voice command, and the location of the sound source when the recognition result indicates that the target gesture has been recognized. The acquisition module is further configured to parse the voice command and determine whether the voice command is a vehicle control command; when the voice command is a vehicle control command, determine whether there is vehicle component information in the vehicle control command; when there is no vehicle component information in the vehicle control command, determine whether there is a preset pronoun in the vehicle control command; when there is the preset pronoun in the vehicle control command, determine that the voice command is a target command. The recognition module is further configured to acquire historical hand data, perform image learning using the historical hand data to construct a two-dimensional hand model, and convert the two-dimensional hand model into a first function; obtain historical gesture posture data based on the first function and the historical hand data, and perform finger pointing gesture learning based on the historical gesture posture data to construct a two-dimensional gesture recognition model, and convert the two-dimensional gesture recognition model into a second function; obtain historical finger pointing gesture data based on the second function and the historical gesture posture data, and perform pointing direction learning based on the historical finger pointing gesture data to construct a two-dimensional pointing model, and convert the two-dimensional pointing model into a third function; and obtain a preset recognition function through the third function. The voice acquisition module includes a microphone array, which includes multiple microphones; The determining module is further configured to obtain the time difference between each microphone in the microphone array based on the sound source localization strategy and the voice information; calculate the angle and distance of the sound source from the microphone array based on the time difference; and determine the sound source location using the angle and the distance.

6. A vehicle, characterized in that, The vehicle includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the vehicle intelligent control method as described in any one of claims 1 to 4.

7. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the vehicle intelligent control method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Indoor quad-rotor unmanned aerial vehicle based on dual-core DSP gesture control

    CN112947589A

  • Vehicle control method, control device, control equipment, vehicle and storage medium

    CN116795202A