Vehicle control methods, devices, vehicles and storage media

By acquiring voice information inside and outside the vehicle, using a voice processing model to identify the location and frequency of the sound source, and constructing a model to accurately identify the driver's voice, the problems of noise misidentification and illegal control in in-vehicle intelligent voice technology are solved, thereby improving vehicle safety and comfort.

CN116386630BActive Publication Date: 2025-11-14CHINA FAW CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310323585.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-29
Publication Date
2025-11-14
Estimated Expiration
2043-03-29

AI Technical Summary

Technical Problem

Existing in-vehicle intelligent voice technology cannot accurately recognize the driver's voice, easily misinterpreting noise as commands, and unauthorized personnel can use high-frequency sound waves to overstep their authority and control the vehicle, posing a security threat.

Method used

By acquiring the vehicle's voice information within the target range, the location and audio frequency information of the sound source are determined. A voice processing model is then used for recognition, and the voice processing model is constructed to filter out the target voice information and control the vehicle in response to the instructions of the legitimate passenger.

Benefits of technology

It achieves accurate recognition of the driver's voice, filters noise and illegal sound interference, and ensures the safety and comfort of the vehicle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116386630B_ABST
    Figure CN116386630B_ABST
Patent Text Reader

Abstract

This invention discloses a vehicle control method, device, vehicle, and storage medium. The method includes: acquiring voice information of the vehicle within a target range, wherein the voice information includes human voice information and noise information used to interfere with the human voice information; acquiring the location information and audio frequency information of the sound source corresponding to the voice information; calling a voice processing model to perform information recognition on the location information and audio frequency information to obtain target voice information in the voice information, wherein the target voice information is information from a passenger in the vehicle, and the voice processing model is trained using historical voice information of the passenger as training samples, the historical voice information including historical location information and historical audio frequency information; and controlling the vehicle in response to a target command corresponding to the target voice information. This invention solves the technical problem of inaccurate driver voice recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicles, and more specifically, to a vehicle control method, apparatus, vehicle, and storage medium. Background Technology

[0002] Currently, in-vehicle intelligent voice technology is becoming increasingly popular in vehicles. Based on this technology, drivers can use voice information to control some functions of the car, avoiding the need for drivers to take their hands off the steering wheel while driving, thus greatly ensuring the safety and comfort of the driver during the driving process.

[0003] In related technologies, the use of in-vehicle intelligent voice technology may misidentify noise in voice information as commands, or unauthorized personnel may use high-frequency sound waves in voice information that are inaudible to the human ear to overstep their authority and control the vehicle, thereby threatening the driver's life and property. Therefore, the above-mentioned problems involve technical issues such as the inability to accurately recognize the driver's voice.

[0004] There is currently no effective solution to the problem that existing technologies cannot accurately recognize driver voices. Summary of the Invention

[0005] This invention provides a vehicle control method, apparatus, vehicle, and storage medium to at least solve the technical problem of inaccurate driver voice recognition.

[0006] According to one aspect of the present invention, a vehicle control method is provided, comprising: acquiring voice information of the vehicle within a target range, wherein the voice information includes human voice information and noise information used to interfere with the human voice information; acquiring location information of the sound source corresponding to the voice information and audio frequency information of the sound source corresponding to the voice information; calling a voice processing model to perform information recognition on the location information and audio frequency information to obtain target voice information in the voice information, wherein the target voice information is information from a passenger in the vehicle, and the voice processing model is trained using historical voice information of the passenger as training samples, the historical voice information including historical location information and historical audio frequency information; and controlling the vehicle in response to a target command corresponding to the target voice information.

[0007] Optionally, a speech processing model is invoked to perform information recognition on the location information and audio frequency information to obtain the target speech information in the speech information, including: in response to the location information matching the target location information in the speech processing model, determining the matching result of the audio frequency information and the target audio frequency information in the speech processing model; in response to the matching result being used to characterize the matching of the audio frequency information and the target audio frequency information, determining the speech information as the target speech information.

[0008] Optionally, the method further includes: in response to the matching result indicating a mismatch between the audio frequency information and the target audio frequency information, prohibiting the target command corresponding to the voice information from controlling the vehicle, and the vehicle issuing a prompt signal.

[0009] Optionally, the method further includes: acquiring historical speech information of passengers in the vehicle within the target range; and constructing a speech processing model based on the historical location information of the sound source corresponding to the historical speech information and the historical audio frequency information of the sound source corresponding to the historical speech information.

[0010] Optionally, a speech processing model is constructed based on the historical location information of the sound source corresponding to the historical speech information and the historical audio frequency information of the sound source corresponding to the historical speech information. The model includes: determining the maximum and minimum values ​​in the historical location information; determining the target location information in the speech processing model based on the maximum and minimum values, and determining the historical audio frequency information of the passenger as the target audio frequency information; and storing the target audio frequency information and the target location information in the speech processing sub-model to obtain the speech processing model.

[0011] Optionally, the target location information in the speech processing model is determined based on the maximum and minimum values, including: standardizing the maximum and minimum values ​​to obtain the target location information.

[0012] Optionally, the method further includes updating the speech processing model according to a predetermined time period.

[0013] According to another aspect of the present invention, a vehicle control device is also provided, comprising: a first acquisition unit, configured to acquire voice information of the vehicle within a target range, wherein the voice information includes human voice information and noise information used to interfere with the human voice information; a second acquisition unit, configured to acquire location information of the sound source corresponding to the voice information and audio frequency information of the sound source corresponding to the voice information; a recognition unit, configured to invoke a voice processing model to perform information recognition on the location information and audio frequency information to obtain target voice information in the voice information, wherein the target voice information is information from a passenger in the vehicle, and the voice processing model is obtained by training with historical voice information of the passenger as training samples, the historical voice information including historical location information and historical audio frequency information; and a control unit, configured to control the vehicle in response to a target command corresponding to the target voice information.

[0014] According to another aspect of the present invention, a vehicle is also provided. This vehicle is used to execute the vehicle control method of the present invention.

[0015] According to another aspect of the present invention, a computer-readable storage medium is also provided. The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device where the computer-readable storage medium is located to perform the vehicle control method of the present invention.

[0016] In this embodiment of the invention, voice information of the vehicle within a target range is acquired, wherein the voice information includes human voice information and noise information used to interfere with the human voice information; the location information of the sound source corresponding to the voice information and the audio frequency information of the sound source corresponding to the voice information are acquired; a voice processing model is invoked to perform information recognition on the location information and audio frequency information to obtain target voice information in the voice information, wherein the target voice information is information from passengers in the vehicle, and the voice processing model is trained using the historical voice information of the passengers as training samples, the historical voice information including historical location information and historical audio frequency information; the vehicle is controlled in response to the target command corresponding to the target voice information. In other words, this invention, by acquiring voice information of the vehicle within a target range, determining the location information and audio frequency information of the sound source corresponding to the voice information, and using a voice processing model to recognize the determined location information and audio frequency information, obtains the target voice information of passengers in the vehicle, and further controls the vehicle based on the target command corresponding to the target voice information, thereby achieving the technical effect of accurately recognizing the driver's voice and solving the technical problem of inaccurate driver voice recognition. Attached Figure Description

[0017] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:

[0018] Figure 1 This is a flowchart of a vehicle control method according to an embodiment of the present invention;

[0019] Figure 2 This is a schematic diagram of a vehicle control method according to an embodiment of the present invention;

[0020] Figure 3 This is a flowchart illustrating the construction of a speech processing model according to an embodiment of the present invention;

[0021] Figure 4 This is a schematic diagram of a vehicle control device according to an embodiment of the present invention. Detailed Implementation

[0022] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0023] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0024] Example 1

[0025] According to an embodiment of the present invention, an embodiment of a vehicle control method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0026] Figure 1 This is a flowchart of a vehicle control method according to an embodiment of the present invention, such as... Figure 1 As shown, the method may include the following steps:

[0027] Step S102: Obtain the voice information of the vehicle within the target range, wherein the voice information includes human voice information and noise information used to interfere with the human voice information.

[0028] In the technical solution provided by step S102 of the present invention, voice information of a vehicle within a target range can be acquired. The target range can be a pre-defined range, including a range inside the vehicle and / or a pre-defined range outside the vehicle. The voice information can include human voice information and noise information used to interfere with the human voice information. The human voice information can be voice information emitted by a person, such as the sound emitted by a passenger in the vehicle or the sound emitted by a pedestrian outside the vehicle. It can include information that the human ear can recognize and / or high-frequency or low-frequency information that the human ear cannot recognize. The noise information can be used to interfere with the human voice information and can include environmental noise information and / or other noise information other than human voice information. It should be noted that no specific limitations are placed on the content included in the voice information here.

[0029] Optionally, a target range can be preset, and based on the preset target range, the voice information of the vehicle within the target range can be obtained.

[0030] For example, the target range can be pre-defined as a circle with a radius of 50 meters centered on the vehicle. This allows the acquisition of voice information of the vehicle within the set target range, which may include human voice information inside and outside the vehicle, as well as noise information used to interfere with the human voice information. This is just an example, and no specific restrictions are placed on the method and size of determining the target range.

[0031] For another example, the target range can be pre-set as a circle with a radius of 50 meters centered on the vehicle. The system can collect human voice information and noise information used to interfere with the human voice information within the set target range every ten seconds.

[0032] Step S104: Obtain the location information of the sound source corresponding to the speech information and the audio frequency information of the sound source corresponding to the speech information.

[0033] In the technical solution provided in step S104 of the present invention, voice information of the vehicle within the target range is acquired. Based on the acquired voice information, the location information of the sound source corresponding to the acquired voice information and the audio frequency information of the sound source corresponding to the acquired voice information can be determined. The sound source can be the source of sound, for example, the source of human voice information or the source of noise information. The location information of the sound source can be the distance or coordinate information of the sound source from the vehicle-mounted voice microphone; for example, the location information can be 1 meter or (0.5 meters, 0.5 meters). The audio frequency information of the sound source can be the sound wave frequency information of the acquired sound source signal; for example, it can be 2 Hz (abbreviated as Hz). This is only an example, and no specific limitations are made on the method and magnitude of determining the location information and audio frequency information.

[0034] Optionally, by acquiring the voice information of the vehicle within a pre-set target range, the location information of the sound source corresponding to the voice information and the audio frequency information of the sound source corresponding to the voice information can be determined.

[0035] For example, by periodically collecting voice information of the vehicle within a pre-set target range, the distance between the sound source corresponding to the acquired voice information and the vehicle's voice microphone can be determined, and the distance information is 1 meter. The audio frequency information of the sound source corresponding to the acquired voice information can also be determined, and the audio frequency information is 2 Hz.

[0036] Step S106: Call the speech processing model to perform information recognition on the location information and audio frequency information to obtain the target speech information in the speech information. The target speech information is the information from the passenger in the vehicle. The speech processing model is trained using the historical speech information of the passenger as training samples. The historical speech information includes historical location information and historical audio frequency information.

[0037] In the technical solution of step S106 of the present invention, the speech processing model is invoked to identify the location information of the sound source corresponding to the acquired speech information and the audio frequency information of the sound source corresponding to the speech information, thereby obtaining the target speech information in the speech information. The target speech information is information from passengers in the vehicle, which can be the driver's information or the information of a legitimate user of the vehicle. Speech information other than the target speech information is called abnormal speech information. The speech processing model is a model trained using the historical speech information of the passengers as training samples. The speech processing model can also be called a sound wave model. The historical speech information of the passengers can be the driver's stored historical speech information or the historical speech information of a legitimate user of the vehicle. A legitimate user of the vehicle can be a passenger. Historical speech information includes historical location information and historical audio frequency information.

[0038] Optionally, by acquiring the historical location information and historical audio frequency information of the passenger, the historical location information can be selected as a modeling feature, and combined with the historical audio frequency information, the historical speech information of the passenger can be used as training samples to train a speech processing model. Based on the constructed speech processing model, and the acquired location information and audio frequency information of the sound source corresponding to the speech information, the speech processing model can be invoked to perform information recognition on the location information and audio frequency information to obtain the target speech information in the speech information.

[0039] For example, by acquiring the historical location information and historical audio frequency information of a passenger, the historical location information can be selected as a modeling feature, and combined with the historical audio frequency information, the historical speech information of the passenger can be used as training samples to train a speech processing model. Based on the constructed speech processing model, and given that the location information of the sound source corresponding to the speech information is 1 meter and the audio frequency information of the sound source corresponding to the speech information is 2Hz, the speech processing model can be invoked to perform information recognition on the determined location information and audio frequency information. The determined location information can be used to mask speech information outside the vehicle that is beyond the range of the location information, and the determined audio frequency information can be used to recognize the speech information of the passenger inside the vehicle, thereby obtaining the target speech information of the passenger.

[0040] Step S108: In response to the target command corresponding to the target voice information, control the vehicle.

[0041] In step S108 of the present invention, by acquiring the target voice information from the voice information, the target command corresponding to the target voice information can be determined. In response to the target command corresponding to the target voice information, the vehicle can be controlled. The target command can be the instruction information issued by the target voice information, such as making a phone call, sending a text message, or turning on navigation. This is merely an example and does not impose specific limitations on the content of the target command. The target command can be used to control the corresponding functions of the vehicle based on the instruction information issued by the target voice information.

[0042] Optionally, by acquiring the target voice information in the voice information, the target command corresponding to the target voice information can be determined, and the vehicle can be controlled according to the target command corresponding to the target voice information.

[0043] For example, if the vehicle obtains the voice information of a passenger who wants to make a phone call, the vehicle will execute the call-making instruction in response to the passenger's voice information, thus completing the function of making a phone call.

[0044] For another example, if the voice information of a passenger in the vehicle is to open the navigation, the vehicle will execute the navigation command in response to the voice information of the passenger, thus completing the function of opening the navigation application for navigation.

[0045] In steps S102 to S108 of this application, the vehicle's voice information within a target range is acquired. This voice information includes human voice information and noise information used to interfere with the human voice. The location information and audio frequency information of the sound source corresponding to the voice information are also acquired. A voice processing model is invoked to identify the location and audio frequency information to obtain the target voice information. The target voice information is information from passengers in the vehicle. The voice processing model is trained using historical voice information of the passengers as training samples. This historical voice information includes historical location and audio frequency information. The vehicle is then controlled in response to the target command corresponding to the target voice information. In other words, this invention acquires the vehicle's voice information within a target range, determines the location and audio frequency information of the sound source corresponding to the voice information, and uses a voice processing model to identify the determined location and audio frequency information to obtain the target voice information of passengers in the vehicle. Furthermore, based on the target command corresponding to the target voice information, the vehicle is controlled, thereby achieving the technical effect of accurately recognizing the driver's voice and solving the technical problem of inaccurate driver voice recognition.

[0046] The method described in this embodiment will be further described below.

[0047] As an optional embodiment, step S106 involves calling a speech processing model to perform information recognition on location information and audio frequency information to obtain target speech information in the speech information, including: in response to the location information matching the target location information in the speech processing model, determining the matching result of audio frequency information and target audio frequency information in the speech processing model; and in response to the matching result being used to characterize the matching of audio frequency information and target audio frequency information, determining the speech information as target speech information.

[0048] In this embodiment, a speech processing model is invoked to identify the location information of the sound source corresponding to the acquired speech information and the audio frequency information of the sound source corresponding to the speech information. In response to the acquired location information matching the target location information in the speech processing model, the matching result between the acquired audio frequency information and the target audio frequency information in the speech processing model can be determined. The determined matching result is used to characterize the match between the acquired audio frequency information and the target audio frequency information, thus confirming that the speech information is the target speech information. The target location information can be the historical location information of passengers in the vehicle pre-stored in the speech processing model, or the distance or coordinate information of the sound source emitted by the passenger from the vehicle's microphone pre-stored. For example, the target location information can be 1 meter or (0.5 meters, 0.5 meters). The target audio frequency information can be the historical audio frequency information of passengers in the vehicle pre-stored in the speech processing model, or the audio frequency information corresponding to the sound source emitted by the passenger pre-stored. For example, the target audio frequency information can be 2Hz. This is merely an example, and no specific limitations are made on the determination method and size of the target location information and the target audio frequency information.

[0049] Optionally, after obtaining the location information of the sound source corresponding to the speech information and the audio frequency information of the sound source corresponding to the speech information, a speech processing model can be invoked to perform information recognition on the location information and audio frequency information. In response to the matching of the obtained location information with the target location information in the speech processing model, the speech information outside the vehicle is masked by the location information. Furthermore, the matching result of the obtained audio frequency information with the target audio frequency information in the speech processing model can be determined. In response to the determined matching result, the obtained audio frequency information is used to represent that it matches the target audio frequency information. The speech information inside the vehicle is identified by the audio frequency information to filter out noise information, thereby determining the target speech information in the speech information.

[0050] For example, if the location information of the sound source corresponding to the voice information is 1 meter and the audio frequency information of the sound source corresponding to the voice information is 2Hz, the voice processing model can be called to perform information recognition on the acquired location information and audio frequency information. In response to the acquisition of location information matching the target location information pre-stored in the voice processing model, which is less than 1.5 meters away, the acquired location information can be used to mask voice information outside the vehicle that exceeds the location information range. Furthermore, the matching result between the acquired audio frequency information and the target audio frequency information pre-stored in the voice processing model can be determined. In response to the determined matching result, it is used to represent that the acquired audio frequency information matches the target audio frequency information. The determined audio frequency information can be used to identify the voice information of the passenger inside the vehicle, thereby obtaining the target voice information of the passenger.

[0051] This embodiment uses a speech processing model to identify the location and audio frequency information of the sound source corresponding to the acquired speech information. In response to the acquisition of location and audio frequency information matching the target location and target audio frequency information in the speech processing model, it can filter out external speech information that exceeds the location information range based on the location information, and identify the speech information of passengers inside the vehicle based on the audio frequency information to filter out noise information. This allows the target speech information in the speech information to be determined, achieving the technical effect of accurately recognizing the driver's speech and solving the technical problem of not being able to accurately recognize the driver's speech.

[0052] As an optional embodiment, the method further includes: in response to the matching result indicating a mismatch between the audio frequency information and the target audio frequency information, prohibiting the target command corresponding to the voice information from controlling the vehicle, and the vehicle issuing a prompt signal.

[0053] In this embodiment, a speech processing model is invoked to identify the location information and audio frequency information of the sound source corresponding to the acquired speech information. In response to a match between the acquired location information and the target location information in the speech processing model, the matching result between the acquired audio frequency information and the target audio frequency information in the speech processing model is determined. This determined matching result indicates a mismatch between the acquired audio frequency information and the target audio frequency information, prohibiting the target command corresponding to the speech information from controlling the vehicle, and the vehicle issues a warning signal. The warning signal can be a text message, a voice message, etc. For example, it could be a text message sent to the vehicle's passengers to indicate that a mismatched voice message is attempting to control the vehicle, or it could be an alarm message issued to the vehicle. It should be noted that no specific restrictions are placed on the content of the warning signal here.

[0054] Optionally, after obtaining the location information of the sound source corresponding to the voice information and the audio frequency information of the sound source corresponding to the voice information, the voice processing model can be invoked to perform information recognition on the location information and audio frequency information. In response to the matching of the obtained location information with the target location information in the voice processing model, the matching result of the obtained audio frequency information with the target audio frequency information in the voice processing model can be further determined. In response to the determined matching result being used to characterize the mismatch between the obtained audio frequency information and the target audio frequency information, the target command corresponding to the voice information is prohibited from controlling the vehicle, and the vehicle issues a prompt signal.

[0055] For example, if the location information of the sound source corresponding to the voice information is 1 meter and the audio frequency information of the sound source corresponding to the voice information is 2Hz, the voice processing model can be called to perform information recognition on the acquired location information and audio frequency information. In response to the acquisition of location information matching the target location information pre-stored in the voice processing model as a distance of less than 1.5 meters, the matching result of the acquired audio frequency information and the target audio frequency information pre-stored in the voice processing model can be determined. In response to the determined matching result being used to indicate that the acquired audio frequency information does not match the target audio frequency information, the target command corresponding to the voice information is prohibited from controlling the vehicle, and a text message is sent to the passengers of the vehicle to indicate that there is a mismatched voice information attempting to control the vehicle.

[0056] As an optional embodiment, the method further includes: acquiring historical voice information of passengers in the vehicle within the target range; and constructing a voice processing model based on the historical location information of the sound source corresponding to the historical voice information and the historical audio frequency information of the sound source corresponding to the historical voice information.

[0057] In this embodiment, historical voice information of passengers in the vehicle within the target range can be obtained, along with historical location information and historical audio frequency information of the corresponding sound sources, which can be used to construct a voice processing model.

[0058] Optionally, historical voice information of passengers in the vehicle within the target range can be obtained. The obtained historical voice information can be filtered and processed, and abnormal and duplicate historical voice information can be deleted. Furthermore, the historical location information and historical audio frequency information of the sound source corresponding to the historical voice information can be obtained. The obtained historical location information can be selected as a modeling feature, and the obtained historical audio frequency information can be recorded to complete the construction of the voice processing model.

[0059] In related technologies, controlling voice assistants through machine learning suffers from problems such as misinterpreting noise as commands and unauthorized personnel using high-frequency sound waves inaudible to the human ear to circumvent vehicle control restrictions. This embodiment addresses these issues by constructing a voice processing model to identify the location and frequency information of the sound source corresponding to the acquired voice information. This allows for the identification of the target voice information within the voice data, achieving accurate driver voice recognition and resolving the technical problem of inaccurate driver voice recognition.

[0060] As an optional implementation method, a speech processing model is constructed based on the historical location information of the sound source corresponding to the historical speech information and the historical audio frequency information of the sound source corresponding to the historical speech information. The model includes: determining the maximum and minimum values ​​in the historical location information; determining the target location information in the speech processing model based on the maximum and minimum values, and determining the historical audio frequency information of the passenger as the target audio frequency information; and storing the target audio frequency information and the target location information in the speech processing sub-model to obtain the speech processing model.

[0061] In this embodiment, by obtaining the historical location information of the sound source corresponding to the historical speech information and the historical audio frequency information of the sound source corresponding to the historical speech information, the maximum and minimum values ​​in the historical location information can be determined. Based on the determined maximum and minimum values ​​in the historical location information, the target location information in the speech processing model can be determined, and the historical audio frequency information of the passenger is determined as the target audio frequency information. Furthermore, by storing the determined target audio frequency information and target location information in the speech processing sub-model, the speech processing model can be obtained.

[0062] Optionally, by obtaining the historical location information and historical audio frequency information of the sound source corresponding to the historical speech information, the maximum and minimum distances between the sound source emitted by the passenger and the vehicle-mounted voice microphone in the historical location information can be obtained. Based on the determined maximum and minimum values ​​in the historical location information, the target location information in the speech processing model can be determined, and the historical audio frequency information of the passenger can be determined as the target audio frequency information. Furthermore, by storing the determined target audio frequency information and target location information in the speech processing sub-model, the construction of the speech processing model can be completed.

[0063] As an optional implementation method, the target location information in the speech processing model is determined based on the maximum and minimum values, including: standardizing the maximum and minimum values ​​to obtain the target location information.

[0064] In this embodiment, the maximum and minimum values ​​in the historical location information are determined, and the target location information is obtained by standardizing these values. The maximum value can be the maximum distance between the sound source emitted by the passenger and the vehicle's microphone in the historical location information, denoted by Vmax. The minimum value can be the minimum distance between the sound source emitted by the passenger and the vehicle's microphone in the historical location information, denoted by Vmin. Standardization can be used to unify the distance data in the historical location information. For example, the maximum and minimum values ​​in the historical location information can be determined first, and the range can be calculated (the range being the difference between the maximum and minimum values). Then, the minimum value is subtracted from each distance value in the determined historical location information, and the result is divided by the range to complete the standardization process. This is merely an example and does not impose specific limitations on the method for determining the standardization process.

[0065] Optionally, the maximum and minimum distances between the sound source emitted by the passenger and the vehicle-mounted voice microphone can be obtained from the historical location information. The maximum and minimum distances in the determined historical location information can be standardized to obtain the target location information.

[0066] For example, the maximum and minimum distances (Vmax and Vmin) between the sound source emitted by a passenger and the vehicle's microphone can be obtained from historical location information. The maximum and minimum values ​​in the historical location information can be standardized using the following formula:

[0067] (V-Vmin) / (Vmax-Vmin)

[0068] Where V represents every distance value in the determined historical location information, excluding the maximum value Vmax and the minimum value Vmin. After standardizing the maximum and minimum distances between the sound source emitted by the passenger and the vehicle's voice microphone in the acquired historical location information, the target location information can be obtained.

[0069] As an optional embodiment, the method further includes updating the speech processing model according to a predetermined time period.

[0070] In this embodiment, a speech processing model is constructed and can be updated according to a predetermined time period. The predetermined time period can be a pre-set time interval for updating the speech processing model, such as a pre-set daily update interval. It should be noted that no specific limitation is placed on the determination of the predetermined time period here.

[0071] Optionally, the constructed speech processing model can be updated according to a predetermined time period, thereby ensuring the safety of the in-vehicle voice system.

[0072] For example, once a voice processing model is built, it can be updated daily according to a pre-set time cycle to ensure the safety of the in-vehicle voice system. This achieves the technical effect of accurately recognizing the driver's voice and solves the technical problem of not being able to accurately recognize the driver's voice.

[0073] This embodiment acquires voice information of a vehicle within a target range, including human voice information and noise information used to interfere with the human voice information; it acquires the location information and audio frequency information of the sound source corresponding to the voice information; it calls a voice processing model to recognize the location information and audio frequency information to obtain the target voice information in the voice information, wherein the target voice information is information from passengers in the vehicle, and the voice processing model is trained using the historical voice information of the passengers as training samples, the historical voice information including historical location information and historical audio frequency information; and it controls the vehicle in response to the target command corresponding to the target voice information. In other words, this invention acquires voice information of a vehicle within a target range, determines the location information and audio frequency information of the sound source corresponding to the voice information, and uses a voice processing model to recognize the determined location information and audio frequency information to obtain the target voice information of passengers in the vehicle. Furthermore, based on the target command corresponding to the target voice information, it controls the vehicle, thereby achieving the technical effect of accurately recognizing the driver's voice and solving the technical problem of inaccurate driver voice recognition.

[0074] Example 2

[0075] The technical solutions of the embodiments of the present invention will be illustrated below with reference to preferred embodiments.

[0076] With the rapid development of the automotive industry, in-vehicle intelligent voice technology is becoming increasingly common in vehicles. Based on this technology, drivers can control some of the car's functions using their voice, preventing them from taking their hands off the steering wheel and greatly enhancing driver safety and comfort. However, the use of in-vehicle intelligent voice technology also presents certain safety risks. For example, it can misinterpret noise as commands, and unauthorized personnel can use high-frequency sound waves, which are inaudible to the human ear, to illegally control the vehicle, threatening the driver's life and property.

[0077] To address the aforementioned issues, a machine learning-based method for defending against silent command control of voice assistants has been proposed in related technologies. This method uses machine learning to control the voice assistant, but it cannot solve the problems of misidentifying noise as commands and unauthorized personnel using high-frequency sound waves inaudible to the human ear to circumvent vehicle control. A domain-adaptive silent voice attack detection method has also been proposed, primarily targeting adversarial attack detection methods. However, this method cannot control in-vehicle intelligent voice technology, thus failing to effectively guarantee driver safety. A variational inference silent attack detection method enhanced by memory networks has also been proposed. This method uses memory storage to differentiate audio based on the type of voice input, preventing unauthorized speech attacks. However, this method also cannot solve the problems of misidentifying noise as commands and unauthorized personnel using high-frequency sound waves inaudible to the human ear to circumvent vehicle control.

[0078] Furthermore, addressing the lack of systematic models and memory storage for voice signals from different directions and frequencies in existing voice assistant anti-theft technologies, this invention provides a vehicle control method. This method acquires voice information from within a target range, determines the location and frequency information of the corresponding sound source, and uses a voice processing model to identify the determined location and frequency information to obtain the target voice information of the occupants in the vehicle. Based on the target commands corresponding to the target voice information, the vehicle is controlled. This embodiment effectively filters out noise and masks high-frequency or low-frequency human voice information outside the vehicle that is inaudible to the human ear, ensuring that only the target voice information of the occupants inside the vehicle is collected, such as the driver and other legitimate vehicle users. It also ensures that regardless of whether the voice information is inside or outside the vehicle, whether it is a high-frequency or low-frequency signal, the in-vehicle voice system masks all voice information except the target voice information and only identifies the target commands corresponding to the target voice information of the occupants. Furthermore, the voice processing model is updated according to a predetermined time period, greatly ensuring the security of the in-vehicle voice system.

[0079] Figure 2 This is a schematic diagram of a vehicle control method according to an embodiment of the present invention, as shown below. Figure 2As shown, the voice information existing inside and outside the vehicle can include: legal voice, illegal voice, and environmental noise. By collecting voice information inside and outside the vehicle within the target range every ten seconds, the first five seconds collect noise information that interferes with human voice information, and the last five seconds collect all human voice information, thus obtaining the location information and audio frequency information of the sound source corresponding to the voice information. Further obtaining the historical location information and historical audio frequency information of the passengers allows the historical location information to be selected as a modeling feature, combined with the historical audio frequency information, to complete the construction of a sound wave model. The constructed sound wave model can identify the location information and audio frequency information of the sound source corresponding to the acquired voice information. It can use the location information to mask external voice information and use the audio frequency information to identify internal voice information, filtering out noise information. This allows the identification of the target voice information within the voice information, and further, based on the target command corresponding to the target voice information, the vehicle can be controlled.

[0080] Figure 3 This is a flowchart of constructing a speech processing model according to an embodiment of the present invention, such as... Figure 3 As shown, the process of constructing a speech processing model may include the following steps:

[0081] Step S301: Process abnormal voice information.

[0082] In step S301 above, the historical voice information of the passenger is obtained, and abnormal historical voice information in the obtained historical voice information can be processed and filtered.

[0083] Step S302: Select location information as modeling features.

[0084] In step S302 above, the historical location information and historical audio frequency information of the passenger's historical voice information are obtained, and the historical location information can be selected as the modeling feature.

[0085] Step S303: Determine the maximum and minimum values ​​in the historical location information.

[0086] In step S303 above, the historical location information of the passenger is obtained, and the maximum value Vmax and minimum value Vmin of the distance between the sound source emitted by the passenger and the vehicle-mounted voice microphone can be determined.

[0087] Step S304: Standardize the maximum and minimum values.

[0088] In step S304 above, the maximum value Vmax and minimum value Vmin of the distance between the sound source emitted by the passenger and the vehicle-mounted voice microphone in the historical location information are determined. The maximum and minimum values ​​in the determined historical location information can be standardized using the following formula:

[0089] (V-Vmin) / (Vmax-Vmin)

[0090] Where V represents each distance value in the determined historical location information, excluding the maximum value Vmax and the minimum value Vmin.

[0091] Step S305: Construct a speech processing model.

[0092] In step S305 above, the speech processing model is further constructed based on the historical audio frequency information of the passengers.

[0093] In this embodiment, when modeling the speech processing model, the repeated and abnormal historical speech information is first filtered and deleted. Then, the historical location information is used as a modeling feature, and the maximum and minimum values ​​in the historical location information are standardized. The standardized historical location information and historical audio frequency information are added to the speech processing sub-model to complete the construction of the speech processing model.

[0094] In this embodiment, by periodically collecting voice information from inside and outside the vehicle within the target range, the location information and audio frequency information of the sound source corresponding to the voice information can be obtained. Further, based on the historical location information and historical audio frequency information of the passenger, the historical location information is selected as a modeling feature, and combined with the historical audio frequency information, a voice processing model is constructed. The constructed voice processing model is then used to identify the location information and audio frequency information of the sound source corresponding to the acquired voice information, achieving dual shielding against illegal target commands. The location information can be used for the first layer of shielding against external voice information, and the audio frequency information can be used to identify internal voice information, filtering out noise information and achieving a second layer of shielding. This allows for the identification of the target voice information within the voice information and accurate recognition of the driver's voice information.

[0095] This embodiment acquires voice information of a vehicle within a target range, including human voice information and noise information used to interfere with the human voice information; it acquires the location information and audio frequency information of the sound source corresponding to the voice information; it calls a voice processing model to recognize the location information and audio frequency information to obtain the target voice information in the voice information, wherein the target voice information is information from a passenger in the vehicle, and the voice processing model is trained using the passenger's historical voice information as training samples, the historical voice information including historical location information and historical audio frequency information; and it controls the vehicle in response to the target command corresponding to the target voice information. In other words, this invention acquires voice information of a vehicle within a target range, determines the location information and audio frequency information of the sound source corresponding to the voice information, and uses a voice processing model to recognize the determined location information and audio frequency information to obtain the target voice information emitted by the passenger. Furthermore, based on the target command corresponding to the target voice information, it controls the vehicle, thereby achieving the technical effect of accurately recognizing the driver's voice and solving the technical problem of inaccurate driver voice recognition.

[0096] Example 3

[0097] According to an embodiment of the present invention, a vehicle control device is also provided. It should be noted that this vehicle control device can be used to execute the vehicle control method of Embodiment 1.

[0098] Figure 4 This is a schematic diagram of a vehicle control device according to an embodiment of the present invention. Figure 4 As shown, the vehicle control device 400 may include: a first acquisition unit 402, a second acquisition unit 404, an identification unit 406, and a control unit 408.

[0099] The first acquisition unit 402 is used to acquire the voice information of the vehicle within the target range, wherein the voice information includes human voice information and noise information used to interfere with the human voice information.

[0100] The second acquisition unit 404 is used to acquire the location information of the sound source corresponding to the speech information and the audio frequency information of the sound source corresponding to the speech information.

[0101] The recognition unit 406 is used to call the speech processing model to recognize the location information and audio frequency information to obtain the target speech information in the speech information. The target speech information is the information from the passenger in the vehicle. The speech processing model is trained using the historical speech information of the passenger as training samples. The historical speech information includes historical location information and historical audio frequency information.

[0102] The control unit 408 is used to control the vehicle in response to the target command corresponding to the target voice information.

[0103] Optionally, the recognition unit 406 includes: a first determining module, configured to determine the matching result of audio frequency information and target audio frequency information in the speech processing model in response to the matching of location information and target location information in the speech processing model; and a second determining module, configured to determine the speech information as target speech information in response to the matching result representing the matching of audio frequency information and target audio frequency information.

[0104] Optionally, the device further includes: a prompting unit, configured to, in response to a matching result indicating a mismatch between the audio frequency information and the target audio frequency information, prohibit the target command corresponding to the voice information from controlling the vehicle, and the vehicle issues a prompting signal.

[0105] Optionally, the device further includes: a third acquisition unit for acquiring historical voice information of passengers in the vehicle within the target range; and a construction unit for constructing a voice processing model based on the historical location information of the sound source corresponding to the historical voice information and the historical audio frequency information of the sound source corresponding to the historical voice information.

[0106] Optionally, the construction unit includes: a first determining module for determining the maximum and minimum values ​​in the historical location information; a second determining module for determining the target location information in the speech processing model based on the maximum and minimum values, and determining the historical audio frequency information of the passenger as the target audio frequency information; and a storage module for storing the target audio frequency information and the target location information in the speech processing sub-model to obtain the speech processing model.

[0107] Optionally, the second determining module includes a processing submodule for standardizing the maximum and minimum values ​​to obtain target location information.

[0108] Optionally, the device further includes an update unit for updating the speech processing model according to a predetermined time period.

[0109] In this embodiment of the invention, a first acquisition unit acquires voice information of the vehicle within a target range, wherein the voice information includes human voice information and noise information used to interfere with the human voice information; a second acquisition unit acquires the location information of the sound source corresponding to the voice information and the audio frequency information of the sound source corresponding to the voice information; a recognition unit calls a voice processing model to recognize the location information and audio frequency information to obtain the target voice information in the voice information, wherein the target voice information is information from a passenger in the vehicle, and the voice processing model is trained using the passenger's historical voice information as training samples, the historical voice information including historical location information and historical audio frequency information; a control unit controls the vehicle in response to the target command corresponding to the target voice information. In other words, this invention acquires voice information of the vehicle within a target range, determines the location information and audio frequency information of the sound source corresponding to the voice information, and uses a voice processing model to recognize the determined location information and audio frequency information to obtain the target voice information emitted by the passenger. Furthermore, based on the target command corresponding to the target voice information, the vehicle is controlled, thereby achieving the technical effect of accurately recognizing the driver's voice and solving the technical problem of inaccurate driver voice recognition.

[0110] Example 4

[0111] According to an embodiment of the present invention, a vehicle is also provided for performing any of the vehicle control methods in Embodiment 1.

[0112] Example 5

[0113] According to an embodiment of the present invention, a computer-readable storage medium is also provided, the storage medium including a stored program, wherein the program executes any of the vehicle control methods in Embodiment 1.

[0114] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0115] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0116] In the several embodiments provided by this invention, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed can be through some interfaces; the indirect coupling or communication connection of units or modules can be electrical or other forms.

[0117] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0118] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0119] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0120] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for controlling a vehicle, characterized in that, include: Acquire voice information of the vehicle within the target range, wherein the voice information includes human voice information and noise information used to interfere with the human voice information; Obtain the location information of the sound source corresponding to the speech information and the audio frequency information of the sound source corresponding to the speech information; The speech processing model is invoked to perform information recognition on the location information and the audio frequency information to obtain the target speech information in the speech information. The target speech information is information from the passenger in the vehicle. The speech processing model is trained using the historical speech information of the passenger as training samples. The historical speech information includes historical location information and historical audio frequency information. The vehicle is controlled in response to the target command corresponding to the target voice information.

2. The method according to claim 1, characterized in that, The speech processing model is invoked to perform information recognition on the location information and the audio frequency information to obtain the target speech information in the speech information, including: In response to the location information matching the target location information in the speech processing model, the matching result of the audio frequency information and the target audio frequency information in the speech processing model is determined; In response to the matching result indicating that the audio frequency information matches the target audio frequency information, the speech information is determined to be the target speech information.

3. The method according to claim 2, characterized in that, The method further includes: In response to the matching result indicating that the audio frequency information does not match the target audio frequency information, the target command corresponding to the voice information is prohibited from controlling the vehicle, and the vehicle issues a prompt signal.

4. The method according to claim 1, characterized in that, The method further includes: Obtain the historical voice information of the passenger in the vehicle within the target range; The speech processing model is constructed based on the historical location information of the sound source corresponding to the historical speech information and the historical audio frequency information of the sound source corresponding to the historical speech information.

5. The method according to claim 4, characterized in that, Based on the historical location information of the sound source corresponding to the historical speech information and the historical audio frequency information of the sound source corresponding to the historical speech information, the speech processing model is constructed, including: Determine the maximum and minimum values ​​in the historical location information; Based on the maximum and minimum values, the target location information in the speech processing model is determined, and the historical audio frequency information of the passenger is determined as the target audio frequency information. The target audio frequency information and the target location information are stored in the speech processing sub-model to obtain the speech processing model.

6. The method according to claim 5, characterized in that, Based on the maximum and minimum values, the target location information in the speech processing model is determined, including: The target location information is obtained by standardizing the maximum and minimum values.

7. The method according to any one of claims 4-6, characterized in that, The method further includes: The speech processing model is updated according to a predetermined time period.

8. A vehicle control device, characterized in that, include: The first acquisition unit is used to acquire voice information of the vehicle within the target range, wherein the voice information includes human voice information and noise information used to interfere with the human voice information; The second acquisition unit is used to acquire the location information of the sound source corresponding to the speech information and the audio frequency information of the sound source corresponding to the speech information. The recognition unit is used to call the speech processing model to perform information recognition on the location information and the audio frequency information to obtain the target speech information in the speech information. The target speech information is information from the passenger in the vehicle. The speech processing model is trained using the historical speech information of the passenger as training samples. The historical speech information includes historical location information and historical audio frequency information. The control unit is used to control the vehicle in response to the target command corresponding to the target voice information.

9. A vehicle, characterized in that, Used to perform the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device on which the computer-readable storage medium is located to perform the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Supplementary driving device based on GPS location and RF radio frequency function

    CN207441035U

  • Method for executing instruction, relevant apparatus and computer program product

    US20220301564A1