Vehicle Control Method, Device, Storage Medium, Controller and Vehicle

Through the combination of lip motion detection and line of sight detection, the problem of vehicle voice system misjudging voice commands in complex environments is solved, and more accurate vehicle voice control is achieved, improving user experience and security.

CN119479649BActive Publication Date: 2025-07-22BYD CO LTD
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202510045429.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-07-22
Estimated Expiration
2045-01-13

AI Technical Summary

Technical Problem

The existing vehicle voice system is prone to misjudgment of voice commands when users talk to others, resulting in misexecution of operations, and it is difficult for users to accurately describe vehicle control commands, resulting in low accuracy of voice control.

Method used

Through the combination of lip motion detection and line of sight detection, we first determine whether the voice command is issued by the target user, further correct the execution target of the voice command to ensure the accuracy of voice control.

Benefits of technology

Improve the accuracy and security of voice control, avoid the error of invalid commands, and enhance user experience and driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119479649B_ABST
    Figure CN119479649B_ABST
Patent Text Reader

Abstract

The present application relates to a vehicle control method, device, storage medium, controller and vehicle. The present application first obtains voice information in the vehicle, performs semantic recognition processing on the voice information to determine an initial semantic instruction, then performs lip movement detection on a target user in the vehicle to obtain a first detection result for determining the validity of the initial semantic instruction. When the first detection result determines that the initial semantic instruction is valid, line-of-sight detection is performed on the target user to obtain a second detection result, and the initial semantic instruction is corrected based on the second detection result to obtain a target semantic instruction. Finally, the target semantic instruction is executed to control the vehicle. In this way, the present application can improve the accuracy of voice control of the vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of vehicle control, and particularly to a vehicle control method, device, storage medium, controller, and vehicle. Background Art

[0002] Currently, in-vehicle voice systems usually have a long voice recognition function. The system can continuously recognize the user's voice within a certain period of time and perform operations corresponding to the voice commands to control the vehicle.

[0003] However, when the user is talking to someone else, the system is prone to misjudging the conversation content as a valid voice command, thus misexecuting the operations of the voice command. In addition, it is difficult for the user to accurately describe the vehicle control commands completely by voice. Therefore, the accuracy of the existing voice control of vehicles is relatively low. Summary of the Invention

[0004] Embodiments of this application provide a vehicle control method, which can improve the accuracy of vehicle voice control to at least partially solve the above technical problems.

[0005] To achieve the above object, according to the first aspect of this application, there is provided a vehicle control method, including:

[0006] Obtain voice information in the vehicle, and perform semantic recognition processing on the voice information to determine an initial semantic command;

[0007] Perform lip movement detection on the target user in the vehicle to obtain a first detection result for determining the validity of the initial semantic command;

[0008] In the case where the first detection result determines that the initial semantic command is valid, perform line-of-sight detection on the target user to obtain a second detection result, and correct the initial semantic command based on the second detection result to obtain a target semantic command;

[0009] Execute the target semantic command to control the vehicle.

[0010] Optionally, the performing semantic recognition processing on the voice information to determine an initial semantic command includes:

[0011] Perform noise reduction processing on the voice information to obtain first voice data after noise reduction;

[0012] Perform endpoint detection processing on the first voice data to determine second voice data including a start time point and an end time point;

[0013] Input the second voice data into a first processing model to perform character recognition on the second voice data to obtain character information;

[0014] Input the text information into a second processing model for semantic recognition to obtain the initial semantic instruction.

[0015] Optionally, perform gaze detection on the target user to obtain a second detection result, including:

[0016] Obtain the gaze information of the target user within a preset time range before and after the start time point;

[0017] Determine the fixation target of the target user based on the gaze information to obtain the second detection result.

[0018] Optionally, the determining the fixation target of the target user based on the gaze information to obtain the second detection result includes:

[0019] When the target user continuously fixates on the same in-vehicle interaction object for more than a preset duration within the preset time range, determine that the target is a valid fixation target;

[0020] When there is only one such valid fixation target, determine the valid fixation target as the only fixation target of the target user;

[0021] When all the in-vehicle interaction objects that the target user continuously fixates on within the preset time range do not exceed the preset duration, or there are multiple valid fixation targets, determine the in-vehicle interaction object or valid fixation target with the maximum continuous fixation duration of the target user as the only fixation target of the target user.

[0022] Optionally, the method further includes:

[0023] When the target user continuously fixates on the same area without the in-vehicle interaction object for more than the preset duration within the preset time range, determine whether there is another user in the area;

[0024] If there is another user in the area, the second detection result is to determine that the initial semantic instruction is invalid;

[0025] If there is no other user in the area, determine the in-vehicle interaction object with the maximum continuous fixation duration of the target user as the only fixation target of the target user to obtain the second detection result.

[0026] Optionally, the modifying the initial semantic instruction based on the second detection result to obtain a target semantic instruction includes:

[0027] Obtain the only fixation target determined based on the second detection result;

[0028] Set the control object of the initial semantic instruction as the only fixation target to obtain the target semantic instruction for controlling the only fixation target.

[0029] Optionally, the lip movement detection of the target user in the vehicle to obtain the first detection result for determining the validity of the initial semantic instruction includes:

[0030] Collect image information including the target user through an in-vehicle image acquisition device;

[0031] Perform target recognition on the image information to obtain the image of the lips of the target user in the image information;

[0032] Detect whether there is lip movement information in the image information of the target user according to the image of the lips of the target user to obtain the first detection result for determining the validity of the initial semantic instruction.

[0033] Optionally, the detecting whether there is the lip movement information in the image information of the target user according to the image of the lips of the target user to obtain the first detection result for determining the validity of the initial semantic instruction further includes:

[0034] Obtain the images of the lips in at least two pieces of the image information, and generate the first detection result according to the images of the lips in at least two pieces of the image information.

[0035] Optionally, the lip movement detection of the target user in the vehicle to obtain the first detection result for determining the validity of the initial semantic instruction further includes:

[0036] If there is the lip movement information in the image information of the target user, the first detection result determines that the initial semantic instruction is valid;

[0037] If there is no lip movement information in the image information of the target user, the first detection result determines that the initial semantic instruction is invalid.

[0038] According to a second aspect of the present application, there is provided a vehicle control device, including:

[0039] A semantic acquisition module, configured to acquire voice information in the vehicle and perform semantic recognition processing on the voice information to determine an initial semantic instruction;

[0040] A lip movement detection module, configured to perform lip movement detection on the target user in the vehicle to obtain a first detection result for determining the validity of the initial semantic instruction;

[0041] A semantic correction module, configured to perform a line-of-sight detection on the target user to obtain a second detection result when the first detection result determines that the initial semantic instruction is valid, and correct the initial semantic instruction based on the second detection result to obtain a target semantic instruction;

[0042] A vehicle control module, configured to execute the target semantic instruction to control the vehicle.

[0043] According to a third aspect of the present application, there is also provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described above are implemented.

[0044] According to a fourth aspect of the present application, there is also provided a controller, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described above are implemented.

[0045] According to a fifth aspect of the present application, there is also provided a vehicle, including the controller described above.

[0046] According to a sixth aspect of the present application, there is also provided a computer program product, including a computer program or instruction, and when the computer program or instruction is executed by a processor, the steps of the method described above are implemented.

[0047] In summary, in the embodiments of the present application, through the above technical solutions, lip movement detection is first used to determine whether the semantic instruction is issued by the target user, so as to avoid receiving and executing invalid semantic instructions issued by other users. After confirming that a valid semantic instruction is received, the line-of-sight detection of the target user can be performed to accurately correct the actual execution target of the voice instruction, which can not only improve the accuracy and integrity of the semantic instruction description, but also improve the accuracy and safety of voice control of the vehicle based on the accurate semantic instruction.

[0048] Other features and advantages of the present application will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings based on these drawings without creative efforts.

[0050] In order to more fully understand the present application and its beneficial effects, the following description will be made in conjunction with the drawings, where the same reference numerals represent the same parts in the following description.

[0051] Figure 1It is a flowchart of steps of a vehicle control method provided in an exemplary embodiment of the present disclosure;

[0052] Figure 2 It is a schematic flowchart of generating a target semantic instruction provided in an exemplary embodiment of the present disclosure;

[0053] Figure 3 It is a schematic flowchart of determining a unique fixation target provided in an exemplary embodiment of the present disclosure;

[0054] Figure 4 It is a schematic diagram of a target user's fixation target provided in an exemplary embodiment of the present disclosure;

[0055] Figure 5 It is a schematic diagram of a vehicle control device provided in an exemplary embodiment of the present disclosure;

[0056] Figure 6 It is a schematic architecture diagram of a vehicle provided in an exemplary embodiment of the present disclosure. Detailed implementation manners

[0057] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present application.

[0058] As people's attention to environmental protection and low carbon continues to increase, the development pace of new energy vehicles has also significantly accelerated. The automotive industry is accelerating the integration of technologies related to energy, transportation, information and communication, and electrification, networking, and intelligence have become the development trends of the automotive industry. New technologies for new energy vehicles have emerged like bamboo shoots after a spring rain. For example:

[0059] Application No.: CN202410658157.7, Publication No.: CN118238797B, Invention Title: New Energy Vehicle Energy Intelligent Management System, Control Method and Related Equipment;

[0060] Application No.: CN202410672579.X, Publication No.: CN118597091A, Invention Title: New Energy Vehicle Energy Intelligent Management Method, System and Related Equipment;

[0061] Application No.: 202410669449.0, Invention Title: New Energy Vehicle Energy Intelligent Management System, Control Method and Related Equipment;

[0062] All describe hybrid technologies mainly based on electricity, which have multiple advantages such as fast, economical, quiet, smooth, and green.

[0063] The application number is CN202211678720.4, the publication number is CN117382629B, and the invention title is Power Control Method, Device, Medium, Vehicle Controller and Vehicle for Vehicles;

[0064] The application number is CN202311164098.X, the publication number is CN116890770B, and the invention title is Vehicle Control System, Method and Vehicle;

[0065] The application number is CN202311170393.6, the publication number is CN117533292B, and the invention title is Vehicle Control System, Control Method, Controller and Vehicle;

[0066] All describe a new energy power system with four in-wheel motors independently driven as the core, which greatly improves the safety and power performance of new energy vehicles.

[0067] Based on the problems mentioned in the foregoing background technology, in the related art, the application of intelligent voice systems in the in-vehicle environment has been gradually popularized. Common intelligent voice systems have the function of continuous conversation, that is, after the user wakes up the voice system, the system will continuously listen to the user's voice and perform semantic analysis. If the content spoken by the user is recognized to conform to a certain preset command, the system will immediately respond and execute the command. This design can improve the user experience in many scenarios, especially during driving, where the user can control various functions of the in-vehicle system, such as navigation, music playback, and in-vehicle settings, without having to manually operate. Through continuous speech recognition, the user can maintain a smooth conversation with the voice system without having to manually wake up the voice system each time.

[0068] However, in the actual use process, the speech recognition system is prone to misjudging the user's speech commands. Especially in the in-vehicle environment, when the user is having a conversation with the co-pilot or rear passengers, or making a phone call, the speech recognition system may erroneously recognize these voices that are not for vehicle control as commands to the voice system and thus perform irrelevant operations. Such misjudgments not only affect the user's driving experience but may also pose safety hazards. Especially during driving, misoperations may lead to unnecessary interference or accidents. In addition, in some scenarios, such as voice navigation or autonomous driving and other control functions, it is difficult for users to accurately express their control commands and intentions for the vehicle. For example, when describing the surrounding environment and orientation, the traditional voice input method is difficult to meet the precise control requirements. Therefore, the accuracy of the existing voice control for vehicles is relatively low.

[0069] This application provides a vehicle control method. Please refer to Figure 1 , the vehicle control method provided by the embodiments of this application includes steps 101-step 104, which will be introduced in detail below.

[0070] Step 101: Obtain the voice information in the vehicle, and perform semantic recognition processing on the voice information to determine the initial semantic instruction.

[0071] In some embodiments, the voice information issued by the user in the vehicle can be collected through an in-vehicle voice device or component. For example, the voice information "Please switch to night mode" issued by the user to the central control screen can be captured.

[0072] In some embodiments, after obtaining the voice information, preprocessing operations such as noise reduction, speech-to-text conversion, and semantic recognition need to be performed on the voice information to facilitate the conversion of the voice information into a semantic instruction that can interact with the vehicle, and then the vehicle can be accurately controlled based on the semantic instruction.

[0073] It should also be noted that in the vehicle voice control scenario of the present application, one type of scenario is that the user directly issues a clear voice control to a certain in-vehicle interaction device. For example, if the user issues the voice "Switch the central control screen to night mode", the semantic instruction to switch the central control screen to night mode can be directly obtained.

[0074] Another type is the scenario of performing voice recognition and vehicle control only for users with vehicle voice control permissions. For example, if it is pre-set that only the driver can perform voice control on the vehicle, then when capturing the in-vehicle voice information and generating the subsequent semantic instruction, it is necessary to determine whether the generated semantic instruction is issued by the authorized user. And after determining that the semantic instruction is issued by the authority, if the semantic instruction does not clearly specify the control object, for example, the semantic instruction is only "Switch to night mode", it is also necessary to further analyze which in-vehicle interaction device the semantic instruction is issued to, to ensure the accuracy of the semantic instruction itself and the precision of vehicle voice control. Therefore, the vehicle control method of the present application can be applied to the above two scenarios.

[0075] In some embodiments, step 101 may include:

[0076] First, perform noise reduction processing on the voice information to obtain the first voice data after noise reduction;

[0077] Next, perform endpoint detection processing on the first voice data to determine the second voice data including the start time point and the end time point;

[0078] Then, input the second voice data into the first processing model to perform character recognition on the second voice data to obtain character information;

[0079] Finally, input the character information into the second processing model for semantic recognition to obtain the initial semantic instruction.

[0080] Such as Figure 2As shown, specifically, first, noise reduction processing is performed on the received voice signal to remove noise, such as background noise in the vehicle, wind noise, traffic noise, etc. By performing noise reduction processing, the first voice data after noise reduction can be obtained, making the subsequent voice recognition process clearer and more accurate.

[0081] In some embodiments, endpoint detection technology is used to process the first voice data to identify the start time point and end time point of the voice, and determine the effective voice segment. The role of endpoint detection is to identify the effective voice part in the voice signal and exclude the silent or meaningless parts. Through endpoint detection processing, the second voice data obtained contains the start and end time points of the user's actual voice, and the second voice data is used as a more accurate voice segment for subsequent semantic analysis.

[0082] In some embodiments, the second voice data is input into a first processing model, and the first processing model can be an ASR model, that is, a voice recognition model. The ASR model can analyze the audio features in the second voice data and perform text recognition processing on the audio features to convert them into corresponding text information.

[0083] In some embodiments, the extracted text information is input into a second processing model, and the second processing model can be an NLU model, that is, a natural language understanding model. Then, the second processing model is used to further process the text information to extract the intent and instructions in the text information and understand the user's semantics. The purpose is to identify the intent that the user wants to express from the pure text and obtain the initial semantic instruction, which includes the specific vehicle control operations or requests that the user hopes the system will execute.

[0084] Through the above methods, the present application can accurately identify and understand the voice information sent by the user, and preprocess the voice information to obtain the initial semantic instruction, providing accurate semantic support for the subsequent execution of voice commands.

[0085] Step 102: Perform lip movement detection on the target user in the vehicle to obtain a first detection result for determining the validity of the initial semantic instruction.

[0086] Among them, the target user refers to a user with vehicle voice control authority. For example, the target user can be the driver of the vehicle. It can be understood that in the scenario where it is necessary to determine whether the semantic instruction is issued by an authorized user mentioned above, after obtaining the initial semantic instruction, an instruction to perform lip movement detection on the target user is triggered to detect whether the initial semantic instruction comes from the voice information sent by the target user.

[0087] In some embodiments, step 102 may include:

[0088] First, image information including the target user is collected by an in-vehicle image acquisition device;

[0089] Next, perform object recognition on the image information to obtain an image of the lip area of the target user in the image information.

[0090] Finally, detect whether there is lip movement information in the image information of the target user based on the image of the lip area of the target user, and obtain a first detection result for determining the validity of the initial semantic instruction.

[0091] In some embodiments, the in-vehicle image acquisition device may include a camera or sensor installed inside the vehicle for capturing images of the in-vehicle environment in real time, with a focus on capturing the facial images of users. The in-vehicle image acquisition device continuously monitors the faces of each user, especially for the target user to collect image information. The collected image information may include the entire face or partial (such as the lip area) image data, providing a basis for subsequent lip reading recognition and motion analysis. By collecting clear facial image data through the in-vehicle image acquisition device, it can ensure that the quality of the images is sufficient to support lip movement detection and subsequent analysis.

[0092] In some embodiments, perform facial recognition and local feature recognition on the target user. For example, through a face detection algorithm, the system can identify the facial area of the target user from the entire image and further accurately extract the lip area.

[0093] Specifically, the face area can be located (such as using Haar cascade classifiers, deep learning algorithms, etc.), and then through specific facial key point detection methods, the specific position of the lips can be further identified, enabling efficient identification and tracking of the lip area of the target user in real-time images, especially in the case of occlusion or light changes in the target user's face, providing key data for subsequent lip movement detection.

[0094] During the detection process, the lip movement detection algorithm can analyze the dynamic changes in the image of the lip area of the target user, especially the movement of the lips. For example, the opening and closing, movement, or morphological changes of the lips are closely related to the pronunciation or vocalization process. By extracting the movement features of the lips, the algorithm can determine whether there is an obvious lip movement pattern, indicating that the target user is vocalizing or speaking, and further determine whether the initial semantic instruction issued by the target user is valid. Correspondingly, if lip movement information is detected, it indicates that the target user is vocalizing, and the system continues with speech recognition and semantic analysis; if no lip movement information is detected, it is considered that the initial semantic instruction is misrecognized and the invalid initial semantic instruction is discarded.

[0095] In some embodiments, it is possible to obtain images of the lip area in at least two pieces of image information and generate a first detection result based on the images of the lip area in the at least two pieces of image information.

[0096] Specifically, after obtaining the lip region information in multiple images, these images can be dynamically analyzed. Each image represents the lip state at a certain point in time. By comparing the lip changes in these images, it can be determined whether the user is speaking or issuing a command. For example, if there are obvious opening and closing changes in the lips in two images, the system may infer that the target user is probably speaking. On the contrary, if the lips remain static between two images, it may indicate that the target user is not making a sound.

[0097] By comparing the dynamic features of the lips in these images (such as the opening amplitude, lip shape changes, speed, etc.), a first detection result is generated, which is used to determine the validity of the voice command. For example, if the lips are shown to be active in multiple images, it is considered that the target user is speaking and the voice command may be valid; if the lip movement is not obvious, it is considered that the target user is not making a sound and the voice command may be invalid.

[0098] Through the lip movement changes in the time series, the actual speech activities of the user can be captured more accurately, reducing misjudgments caused by small lip movements or facial occlusions in a single image.

[0099] Step 103: When the first detection result determines that the initial semantic command is valid, perform a line-of-sight detection on the target user to obtain a second detection result, and based on the second detection result, correct the initial semantic command to obtain a target semantic command.

[0100] It can be understood that after detecting that the initial semantic command is issued by the target user, it is determined that the initial semantic command is valid. However, if the initial semantic command does not clearly specify the in-vehicle interaction device to be controlled, for example, the initial semantic command is only "switch to night mode", it is necessary to further correct the initial semantic command to obtain a more accurate corrected semantic command.

[0101] In some embodiments, step 103 may include:

[0102] Obtain the line-of-sight information of the target user within a preset time range before and after the starting time point;

[0103] Based on the line-of-sight information, determine the fixation target of the target user to obtain a second detection result.

[0104] As Figure 3 shown, at the starting time point when the voice command starts, for example, at the moment when the target user issues a voice message, trace back and obtain the line-of-sight information of the target user. The line-of-sight information refers to the target area that the target user's eyes are looking at within the preset time range before and after the voice command.

[0105] By way of example only, the preset time range can be 2 seconds before and after the starting time point, that is, within 2 seconds before and after the target user issues a voice command, the eye movement data of the target user can be collected. During this time window, the system can ensure that it captures the gaze changes of the target user's eyes and determines the gaze target of the target user when the command is issued.

[0106] In some embodiments, the line-of-sight information may include the gaze direction of the target user's eyes and the relationship between the eyes and the gazed object. For example, Figure 4 as shown, the gaze of the target user is concentrated on devices such as the center control screen, instrument panel, and HUD, as well as the duration of the gaze. This information is captured in real time by an in-vehicle eye movement tracking device (such as a DMS camera) to obtain a second detection result.

[0107] Through the above method, the judgment method of the present application based on line-of-sight information can significantly improve the accuracy of voice control of the vehicle by the in-vehicle voice system in a complex driving environment, enhancing the user experience and driving safety.

[0108] In some embodiments, the second detection result can be determined by the following method:

[0109] When the target user continuously gazes at the same in-vehicle interaction object for more than a preset duration within the preset time range, the target is determined to be a valid gaze target;

[0110] When there is only one valid gaze target, the valid gaze target is determined to be the only gaze target of the target user;

[0111] When all the in-vehicle interaction objects that the target user continuously gazes at within the preset time range do not exceed the preset duration, or there are multiple valid gaze targets, the in-vehicle interaction object or valid gaze target with the maximum continuous gaze duration of the target user is determined to be the only gaze target of the target user.

[0112] Among them, the in-vehicle interaction object can be an in-vehicle device that can execute vehicle voice control instructions, such as the center control screen, instrument panel, and HUD device. If the target user continuously gazes at the same in-vehicle interaction object for more than the preset duration within the preset time range (for example, the preset duration is 1 second), the system considers this gaze target to be valid and marks it as a valid gaze target. This valid gaze target is used for subsequent voice command or operation judgment.

[0113] In some cases, the target user may only gaze at one in-vehicle interaction object, and the system can directly identify the only valid gaze target. If the system only detects one in-vehicle interaction object and the gaze time of this object exceeds the preset duration (for example, the target user continuously gazes at the center control screen for more than 2 seconds), the system will consider this object to be the only gaze target of the target user.

[0114] In this case, the valid fixation target is directly determined as the only fixation target, and subsequent command execution will be carried out around this target. For example, if the target user gazes at the center console screen and issues a voice command of "switch to night mode", the system automatically takes the center console screen as the operation object to execute the command and switches the center console screen to night mode.

[0115] It can be understood that if the fixation time of the target user on a single target device does not exceed the preset duration within the preset time range, for example, the target user's fixation time is scattered on different in-vehicle devices and does not exceed the preset threshold, or there are multiple valid fixation targets, for example, the target user gazes at both the center console screen and the dashboard for more than the preset duration, the system will make a further judgment.

[0116] In this case, the system compares the continuous fixation durations of multiple fixation targets and selects the target with the longest fixation duration as the only fixation target. For example, if the target user gazes at the center console screen for 2 seconds and gazes at the dashboard for 1 second, the system can detect that the center console screen is the only fixation target of the target user.

[0117] If the fixation durations of multiple targets are similar or none of them significantly exceeds the others, the system can select the most suitable fixation target according to other logics (such as priority, device type, etc.). For example, the system can preferentially select a more important or frequently interacted device (such as the center console screen) as the fixation target.

[0118] In some embodiments, the method of the present application may further include:

[0119] When the target user continuously gazes at the same area without an in-vehicle interaction object within the preset time range for more than the preset duration, it is determined whether there is another user in the area;

[0120] If there is another user in the area, the second detection result is that the initial semantic instruction is determined to be invalid;

[0121] If there is no other user in the area, the in-vehicle interaction object with the maximum continuous fixation duration of the target user is determined as the only fixation target of the target user, and the second detection result is obtained.

[0122] In some embodiments, the line of sight of the target user within the preset time range can be monitored to confirm whether the target user gazes at an area in the vehicle without an in-vehicle interaction object. For example, this area can be the co-pilot seat, the rear seat area, etc. These areas usually do not involve direct interaction with in-vehicle devices. Therefore, if the target user continuously gazes in these areas, it may be to talk to other passengers or gaze at non-in-vehicle-related areas such as the rear row.

[0123] If it is detected that the target user continuously gazes at these non-interactive areas without in-vehicle interaction objects for longer than a preset duration (e.g., more than 2 seconds), it is necessary to further determine whether there are other users in this area. For example, other users can be passengers sitting in the co-pilot or the rear row of the vehicle.

[0124] Specifically, the seat sensing information of the vehicle can be obtained first. The seat sensing information of the cockpit can be obtained through sensors installed on the cockpit seats, such as pressure sensors, capacitive sensors, etc. When someone sits on the seat, the weight of the human body will exert pressure on the seat surface. The pressure sensor can sense this change in pressure and convert it into an electrical signal. For example, in the seat of the co-pilot of the vehicle, the pressure sensors can be distributed in the seat cushion and backrest parts. When the pressure value exceeds a certain threshold, it can be determined that there is someone on the co-pilot seat. The capacitive sensor, on the other hand, uses the change in capacitance between the human body and the sensor to detect whether there is someone. When the human body approaches the sensor, it will change the electric field distribution around the sensor, thereby causing a change in capacitance, and it can determine whether there is someone in the seat area.

[0125] Therefore, if it is determined that there are other users in this area by combining the seat sensing information of the cockpit, it indicates that the target user is probably talking to other users rather than interacting with the in-vehicle system, and it is necessary to judge whether to execute this voice command based on this information.

[0126] It can be understood that if it is detected that the area where the target user gazes is the co-pilot area or the rear row area, and there are other passengers in this area (such as co-pilot or rear row passengers), it is speculated that the target user may not be interacting with the in-vehicle system but communicating with other passengers. In this case, the voice command in this gazing area is determined to be invalid and discarded because the target user's attention is not focused on the in-vehicle interaction device but on the conversation with other passengers, avoiding mis-executing the command. For example, when the target user speaks to the co-pilot, the voice command will not be recognized as an order but regarded as an instruction or conversation to the co-pilot. Therefore, the second detection result of the system is to determine that the initial semantic command is invalid, avoiding misresponding to the voice from other passengers.

[0127] Correspondingly, if the system detects that the area where the target user gazes is the co-pilot or the rear row area, but there are no other passengers in this area (for example, the co-pilot seat is empty), it indicates that the gazing behavior of the target user is more likely to have nothing to do with the in-vehicle system interaction. However, the voice information sent by the target user may still be affected by other important factors in the command judgment. Then, the voice recognition judgment logic for the entire vehicle cockpit is further executed.

[0128] In this case, based on other in-vehicle interaction objects that the target user is gazing at, such as the central control screen, instrument panel, HUD, etc., their true intention will be judged. If the gaze duration of the target user when gazing at these interaction targets is relatively long, the system will determine that the gazing target is the only gazing target of the target user. By comparing the continuous gaze durations of all in-vehicle interaction objects, the in-vehicle interaction target with the longest gaze duration is selected, that is, the interaction device that the target user is most concerned about. For example, if the target user gazes at the central control screen for 2 seconds and gazes at the instrument panel for 1 second, it can be determined that the central control screen is the only gazing target and relevant commands are executed.

[0129] In this case, the second detection result selects the in-vehicle interaction object with the longest gaze time as the only gazing target, ensuring that the voice command is only executed when the target user clearly gazes at the relevant device, so as to achieve the purpose of semantic correction of the initial semantic command.

[0130] In some embodiments, the target semantic command is obtained in the following manner:

[0131] Obtain the only gazing target determined based on the second detection result;

[0132] Set the control object of the initial semantic command as the only gazing target to obtain a target semantic command for controlling the only gazing target.

[0133] Specifically, associate the control object of the initial semantic command with this gazing target. For example, if the target user says "Switch to night mode" and is gazing at the central control screen, the system will combine the "Switch to night mode" command with the central control screen to determine that the operation object is the central control screen.

[0134] Through this association, a more contextually meaningful target semantic command is generated. This command not only includes the original operation intention, such as "Switch to night mode", but also clearly defines the only gazing target to be controlled, such as the target being the central control screen. The target semantic command is the final control command, and the system will execute specific operations according to this command to obtain a final and more accurate target semantic command by correcting the initial semantic command.

[0135] Step 104: Execute the target semantic command to control the vehicle.

[0136] Specifically, after generating the target semantic command, the system transmits this target semantic command to the vehicle control system to execute corresponding operations. For example, if the target semantic command is "Switch the central control screen to night mode", then the vehicle control system controls the display mode of the central control screen to switch to night mode to achieve high-precision voice control of the vehicle based on the semantically accurate target semantic command.

[0137] In summary, in the embodiments of the present application, through the above technical solutions, lip movement detection is first used to determine whether a semantic instruction is issued by the target user, so as to avoid receiving and executing invalid semantic instructions issued by other users. After confirming that a valid semantic instruction is received, gaze detection is performed on the target user to accurately correct the actual execution target of the voice instruction. As a result, not only can the accuracy and integrity of the semantic instruction description be improved, but also the accuracy and safety of voice control of the vehicle can be enhanced based on the accurate semantic instruction.

[0138] Figure 5 It is a schematic structural diagram of a vehicle control device provided in an embodiment of the present application. Please refer to Figure 5 , the vehicle control device may include a semantic acquisition module 201, a lip movement detection module 202, a semantic correction module 203, and a vehicle control module 204. Among them, the semantic acquisition module 201 is configured to acquire voice information in the vehicle and perform semantic recognition processing on the voice information to determine an initial semantic instruction; the lip movement detection module 202 is configured to perform lip movement detection on the target user in the vehicle to obtain a first detection result for determining the validity of the initial semantic instruction; the semantic correction module 203 is configured to perform gaze detection on the target user to obtain a second detection result when the first detection result determines that the initial semantic instruction is valid, and correct the initial semantic instruction based on the second detection result to obtain a target semantic instruction; the vehicle control module 204 is configured to execute the target semantic instruction to control the vehicle.

[0139] Among them, the semantic acquisition module 201, the lip movement detection module 202, the semantic correction module 203, and the vehicle control module 204 can be respectively used to execute the steps 101-104 in the corresponding embodiments of the above vehicle control method. For the specific implementation manners of these modules and more detailed content, reference can be made to the corresponding method part, which will not be elaborated here one by one.

[0140] The embodiments of the present application also provide a computer-readable storage medium, on which instructions are stored, and when the instructions are executed by a processor, the processor is configured to execute the above vehicle control method.

[0141] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0142] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the specified functions in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the specified functions in one or more of the blocks.

[0143] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufacture including instruction means that implement the specified functions in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the specified functions in one or more of the blocks.

[0144] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the specified functions in one or more of the blocks.

[0145] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0146] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0147] Computer-readable media include both permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media do not include transitory media such as modulated communication signals and carrier waves.

[0148] As Figure 6 shown, it is a schematic diagram of the architecture of a vehicle provided in an embodiment mode of the present application. In this embodiment, the vehicle 400 may include a controller 300, and a computer program is stored on the controller 300. When the computer program is executed by a processor, the steps of the above vehicle control method are implemented. In this embodiment, the vehicle may be a fuel vehicle, a plug-in hybrid vehicle, a new energy vehicle, etc., and the present disclosure does not make specific limitations thereto.

[0149] In the description of the present application, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the present application, the meaning of "a plurality" is two or more, unless otherwise specifically defined.

[0150] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0151] Among the embodiments, embodiments, and related technical features of the present application, they can be combined and replaced with each other without conflict.

[0152] The above are only the preferred embodiments of the present application and do not impose any form of limitation on the present application. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present application without departing from the content of the technical solution of the present application still fall within the scope of the technical solution of the present application.

Claims

1. A vehicle control method, characterized in that, Including: Obtain voice information in the vehicle, and perform semantic recognition processing on the voice information to determine an initial semantic instruction; Perform lip movement detection on a target user in the vehicle to obtain a first detection result for determining the validity of the initial semantic instruction; When the first detection result determines that the initial semantic instruction is valid, perform line-of-sight detection on the target user to obtain a second detection result, and correct the initial semantic instruction based on the second detection result to obtain a target semantic instruction, including: when all in-vehicle interaction objects that the target user continuously gazes at within a preset time range do not exceed a preset duration, determine the in-vehicle interaction object with the maximum continuous gaze duration of the target user as the only gaze target of the target user to obtain the second detection result; or When the target user continuously gazes at the same area without in-vehicle interaction objects within a preset time range and exceeds the preset duration, determine whether there are other users in the area; if there are other users in the area, the second detection result is to determine that the initial semantic instruction is invalid; if there are no other users in the area, determine the in-vehicle interaction object with the maximum continuous gaze duration of the target user as the only gaze target of the target user to obtain the second detection result; Execute the target semantic instruction to control the vehicle.

2. The method according to claim 1, characterized in that, The performing semantic recognition processing on the voice information to determine an initial semantic instruction includes: Perform noise reduction processing on the voice information to obtain first voice data after noise reduction; Perform endpoint detection processing on the first voice data to determine second voice data including a start time point and an end time point; Input the second voice data into a first processing model to perform character recognition on the second voice data to obtain character information; Input the character information into a second processing model to perform semantic recognition to obtain the initial semantic instruction.

3. The method according to claim 2, characterized in that The performing line-of-sight detection on the target user to obtain a second detection result includes: Obtain the line-of-sight information of the target user within the preset time range before and after the start time point; Determine the gaze target of the target user based on the line-of-sight information to obtain the second detection result.

4. The method according to claim 3, characterized in that, The determining the gaze target of the target user based on the line-of-sight information to obtain the second detection result includes: When the target user continuously gazes at the same in-vehicle interaction object within the preset time range and exceeds the preset duration, determine that the target is a valid gaze target; When there is only one valid gaze target, determine the valid gaze target as the only gaze target of the target user; When there are multiple valid gaze targets, determine the in-vehicle interaction object or valid gaze target with the maximum continuous gaze duration of the target user as the only gaze target of the target user.

5. The method according to any one of claims 1-4, characterized in that, The correcting the initial semantic instruction based on the second detection result to obtain a target semantic instruction includes: Obtain the only gaze target determined based on the second detection result; Set the control object of the initial semantic instruction as the only fixation target to obtain the target semantic instruction for controlling the only fixation target.

6. The method according to claim 1, characterized in that, The lip movement detection of the target user in the vehicle to obtain the first detection result for determining the validity of the initial semantic instruction includes: Collect the image information including the target user through the vehicle-mounted image acquisition device; Perform target recognition on the image information to obtain the lip part image of the target user in the image information; Detect whether there is the lip movement information in the image information of the target user according to the lip part image of the target user to obtain the first detection result for determining the validity of the initial semantic instruction.

7. The method according to claim 6, characterized in that, The detecting whether there is the lip movement information in the image information of the target user according to the lip part image of the target user to obtain the first detection result for determining the validity of the initial semantic instruction further includes: Obtain at least two lip part images in the image information and generate the first detection result according to the at least two lip part images in the image information.

8. The method according to claim 6, wherein The lip movement detection of the target user in the vehicle to obtain the first detection result for determining the validity of the initial semantic instruction further includes: If there is the lip movement information in the image information of the target user, the first detection result determines that the initial semantic instruction is valid; If there is no lip movement information in the image information of the target user, the first detection result determines that the initial semantic instruction is invalid.

9. A vehicle control device, characterized in that, Includes: A semantic acquisition module for acquiring voice information in the vehicle and performing semantic recognition processing on the voice information to determine the initial semantic instruction; A lip movement detection module for performing lip movement detection on the target user in the vehicle to obtain the first detection result for determining the validity of the initial semantic instruction; A semantic correction module for, when the first detection result determines that the initial semantic instruction is valid, performing gaze detection on the target user to obtain the second detection result, and correcting the initial semantic instruction based on the second detection result to obtain the target semantic instruction, and further for, when all the in-vehicle interaction objects continuously fixated by the target user within the preset time range do not exceed the preset duration, determining the in-vehicle interaction object with the maximum continuous fixation duration of the target user as the only fixation target of the target user to obtain the second detection result; Or When the target user continuously fixates on the same area without in-vehicle interaction objects for more than the preset duration within the preset time range, determine whether there are other users in the area; If there are other users in the area, the second detection result determines that the initial semantic instruction is invalid; if there are no other users in the area, determine the in-vehicle interaction object with the maximum continuous fixation duration of the target user as the only fixation target of the target user to obtain the second detection result; A vehicle control module for executing the target semantic instruction to control the vehicle.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.

11. A controller, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.

12. A vehicle, characterized in that, It includes the controller according to claim 11.

13. A computer program product, characterized in that, It includes a computer program or instructions which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Vehicle control system, method and vehicle

    CN116890770B

  • Vehicle power control method and device, medium, vehicle controller and vehicle

    CN117382629A

  • Vehicle power control method, device, medium, vehicle controller and vehicle

    CN117382629B

  • Vehicle control system, control method, controller and vehicle

    CN117533292A

  • Vehicle control system, control method, controller and vehicle

    CN117533292B