Vehicle, vehicle voice recognition method, device and medium
By identifying the vehicle environment and driving data to determine the target scenario, adjusting the priority of command types, and using the speech recognition model to prioritize the invocation of relevant command types, the accuracy problem of the in-vehicle speech recognition system when the environment changes is solved, achieving higher recognition accuracy and response speed.
Patent Information
- Application Number
- CN202411577035.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-06
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2044-11-06
AI Technical Summary
Existing in-vehicle voice recognition systems have low accuracy when environmental conditions change, which may lead to control commands not matching the user's intentions.
By identifying the vehicle's internal environment data, external environment data, and driving data, the target scenario is determined, and the first command type corresponding to the scenario is retrieved from the database. Its priority is adjusted, and the speech recognition model prioritizes calling the command type for speech recognition.
It improves the accuracy and response speed of voice recognition, ensuring that control commands are consistent with user intentions and enhancing the convenience of vehicle operation.
Smart Images

Figure CN119601004B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of speech recognition technology, and in particular relates to a vehicle, a vehicle speech recognition method, a device and a medium. Background Technology
[0002] With the development of in-vehicle voice recognition technology, there are now voice recognition systems that can automatically recognize user voice commands, identify user intentions, and predict vehicle control commands, thereby improving the convenience of users using vehicles.
[0003] In speech recognition systems of related technologies, the prediction of control commands may be affected by the results of speech recognition. For example, when the accuracy of speech recognition is low, the system may predict or recommend control commands that do not match the user's intentions, and the accuracy of speech recognition may deteriorate based on the environmental conditions of the vehicle. Summary of the Invention
[0004] The embodiments of this application provide a vehicle, a vehicle voice recognition method, an apparatus, and a medium, which can at least to some extent improve the accuracy of voice recognition.
[0005] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.
[0006] The first aspect of this application provides a vehicle voice recognition method, including:
[0007] Based on one or more of the vehicle’s internal environment data, external environment data, and driving data, identify the target scenario in which the vehicle is located.
[0008] According to the target scenario, a first instruction type corresponding to the target scenario is obtained from the database, wherein the database includes multiple scenarios of the vehicle and instruction types corresponding to each scenario;
[0009] Adjust the priority of the first instruction type to be greater than the priority of other instruction types in the database;
[0010] The user's first voice information is obtained, and the first instruction type is preferentially called based on the speech recognition model to recognize the first voice information in order to determine the first instruction corresponding to the first voice information.
[0011] Optionally, retrieving the first instruction type corresponding to the target scenario from the database based on the target scenario includes:
[0012] Obtain a first target identifier, wherein the first target identifier is used to characterize the target scenario;
[0013] Obtain a second target identifier corresponding to the first target identifier from the database, wherein the second target identifier is used to characterize the first instruction type, the database includes a plurality of first identifiers and a second identifier corresponding to each first identifier, the first identifier is used to characterize the scenario, and the second identifier is used to characterize the instruction type.
[0014] Optionally, adjusting the priority of the first instruction type to be greater than the priority of other instruction types in the database includes:
[0015] The target value of the first instruction type is adjusted to a first value, and the target values of the other instruction types are adjusted to a second value, wherein the first value is used to represent a first priority, the second value is used to represent a second priority, and the first priority is greater than the second priority.
[0016] Optionally, obtaining the user's first voice information, and prioritizing the invocation of the first instruction type to recognize the first voice information based on the speech recognition model, includes:
[0017] Obtain the first voice information, wherein the first voice information includes a first task to be executed and a second task to be executed;
[0018] The first instruction type is preferentially invoked based on the speech recognition model to identify the first task to be executed;
[0019] The speech recognition model is used to call other instruction types to identify other second tasks to be executed.
[0020] Optionally, when there are multiple first tasks to be executed, the step of prioritizing the invocation of the first instruction type based on the speech recognition model to identify the first execution includes:
[0021] Determine the timestamp of each of the first tasks to be executed;
[0022] The first instruction type is invoked first based on the speech recognition model, and each of the first tasks to be executed is identified according to the order of the timestamps.
[0023] Optionally, the method further includes:
[0024] Acquire multiple second voice information of the user in the target scenario over a period of time;
[0025] Based on the speech recognition model, multiple pieces of the second speech information are recognized to determine multiple second commands;
[0026] If multiple second instructions belong to the same instruction type, and multiple second instructions do not belong to the first instruction type, then the priority of the second instruction type to which the multiple second instructions belong is increased.
[0027] Optionally, the vehicle includes a voice acquisition device, and when acquiring the user's first voice information, the method further includes:
[0028] Obtain ambient noise intensity and / or vehicle speed;
[0029] If the ambient noise intensity is greater than or equal to a set noise intensity and / or the vehicle speed is greater than or equal to a set vehicle speed, then the noise reduction capability of the voice acquisition device is enhanced.
[0030] A second aspect of this application provides a vehicle voice recognition device, comprising:
[0031] The first identification unit is used to identify the target scenario in which the vehicle is located based on one or more of the vehicle's internal environment data, external environment data, and driving data.
[0032] The acquisition unit is configured to acquire a first instruction type corresponding to the target scenario from a database according to the target scenario, wherein the database includes multiple scenarios of the vehicle and instruction types corresponding to each scenario;
[0033] An adjustment unit is used to adjust the priority of the first instruction type to be greater than the priority of other instruction types in the database;
[0034] The second recognition unit is used to acquire the user's first voice information, and based on the voice recognition model, preferentially call the first instruction type to recognize the first voice information in order to determine the first instruction corresponding to the first voice information.
[0035] A third aspect of this application provides a computer-readable storage medium storing at least one computer program instruction, which is loaded and executed by a processor to perform the operations described in any of the methods described in the first aspect.
[0036] A fourth aspect of this application provides a vehicle including one or more processors and one or more memories, wherein at least one piece of program code is stored in the one or more memories, and the at least one piece of program code is loaded and executed by the one or more processors to perform the operations performed as described in any of the methods in the first aspect.
[0037] The one or more technical solutions provided in the embodiments of the present invention achieve at least the following technical effects or advantages:
[0038] The vehicle voice recognition method provided in this application identifies the target scenario of the vehicle based on one or more of the vehicle's internal environment data, external environment data, and driving data. Based on the target scenario, it retrieves a first instruction type corresponding to the target scenario from a database, wherein the database includes multiple vehicle scenarios and instruction types corresponding to each scenario. The priority of the first instruction type is adjusted to be higher than the priority of other instruction types in the database. The user's first voice information is obtained, and the first instruction type is preferentially called to recognize the first voice information based on a voice recognition model to determine the first instruction corresponding to the first voice information. Therefore, this application embodiment, by pre-constructing the correspondence between different vehicle scenarios and various vehicle instruction types, when the vehicle is identified as being in a target scenario, increases the priority of the target instruction type corresponding to that target scenario based on the preset correspondence. Thus, when user voice is received, the first instruction type related to the target scenario can be used to recognize the user voice, improving the accuracy and response speed of user voice recognition.
[0039] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0040] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:
[0041] Figure 1 A flowchart of a vehicle voice recognition method according to an embodiment of this application is shown;
[0042] Figure 2 A structural diagram of a vehicle voice recognition device according to an embodiment of this application is shown;
[0043] Figure 3 A schematic diagram of the structure of a computer system suitable for implementing embodiments of the present application is shown. Detailed Implementation
[0044] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0045] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.
[0046] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0047] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0048] It should also be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such uses of these terms can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described.
[0049] The vehicle voice recognition method of this application embodiment will be described in detail below with reference to the accompanying drawings.
[0050] See Figure 1 The flowchart of the vehicle voice recognition method according to an embodiment of this application is shown.
[0051] like Figure 1 As shown, the first aspect of this application provides a vehicle voice recognition method, including but not limited to:
[0052] Step S10. Identify the target scenario in which the vehicle is located based on one or more of the vehicle's internal environment data, external environment data, and driving data;
[0053] For example, vehicle interior environmental data includes, but is not limited to: interior temperature, interior humidity, interior light intensity, and interior noise. For instance, an onboard temperature sensor can detect the interior temperature, an onboard humidity sensor can detect the interior humidity, an onboard light intensity sensor can detect the light intensity, and an onboard noise sensor can detect the interior noise.
[0054] For example, vehicle external environment data includes, but is not limited to, weather data (such as sunny, rainy, cloudy, snowy, temperature, precipitation, wind speed, etc.), road condition data (such as traffic conditions, road type data, etc.), and external light intensity data. For instance, weather data can be obtained from an external weather server through the vehicle's communication terminal, road condition data can be obtained from an external traffic server through the vehicle's communication terminal, and external light intensity data can be obtained from the vehicle's light intensity sensor.
[0055] For example, driving data includes, but is not limited to, vehicle speed data, braking data, acceleration data, and steering data. For instance, vehicle speed data can be detected by wheel-end sensors, braking data can be detected by pressure sensors installed on the accelerator pedal, steering data can be detected by onboard angle sensors, and acceleration data can be detected by accelerometers, etc.
[0056] Based on the above, it is possible to identify the target scenario in which the vehicle is located. It is understandable that a target scenario may include multiple sub-scenarios under multiple different scenario categories. For example, in a certain target scenario, the vehicle may be in rainy weather, traffic congestion, low light intensity, low interior temperature, and slow speed.
[0057] Step S20. Based on the target scenario, obtain the first instruction type corresponding to the target scenario from the database, wherein the database includes multiple scenarios of the vehicle and instruction types corresponding to each scenario;
[0058] In this embodiment of the application, through big data analysis, different vehicle scenarios and different vehicle command types can be pre-classified, and a correspondence between scenarios and command types can be established. The scenario classification and command type classification of vehicles can have different classification criteria, which are not limited here. For ease of understanding, the scenario classification and command type classification of vehicles are illustrated below.
[0059] In some embodiments, vehicles are classified into the following scenarios:
[0060] 1) Classify by driving scenario: such as low-speed driving in the city, driving on the highway, traffic jams, waiting in the car, etc.
[0061] 2) Classify by weather conditions: such as sunny, heavy rain, fog, blizzard, extreme temperature, etc.
[0062] 3) Classify by in-vehicle environment: such as in-vehicle noise level, excessively high or low temperature, etc.
[0063] Therefore, we can construct multiple scenario classification sets, each of which includes multiple sub-scenarios.
[0064] In some embodiments, based on the user's voice input, the vehicle can be categorized into the following command types:
[0065] 1) Navigation-related: such as "navigate to a certain place", "next exit", "select route", etc.
[0066] 2) Safety driving-related: such as "slow down", "turn on the windshield wipers", "adjust the seat position", etc.
[0067] 3) Vehicle control functions: such as "turn on the air conditioner", "adjust the windows", "adjust the seat", etc.
[0068] 4) Entertainment control functions: such as "play music", "adjust volume", "switch radio stations", etc.
[0069] 5) Information query category: such as "check the weather" and "road conditions".
[0070] Therefore, multiple instruction type classification sets can be constructed, and each classification set includes multiple sub-scenarios.
[0071] After constructing classification sets for multiple scenarios and multiple instruction types, big data analysis can be used to determine the corresponding instruction type for each scenario and establish a correspondence between scenarios and instruction types. For ease of understanding, the instruction type matching process for some scenarios is explained below.
[0072] 1) Driving scenario matching.
[0073] For example, when driving at high speeds, the command type fields for navigation and safe driving are matched with the high-speed driving scenario to ensure that safety-related operations can be accurately identified and responded to quickly, such as "acceleration command" or "acceleration command".
[0074] For example, when driving at low speeds or on city roads, entertainment control commands (such as "play music" or "change channel") are matched with the low-speed or city road driving conditions.
[0075] For example, when the vehicle is parked or waiting, the response requirements for safety commands can be reduced, and entertainment, air conditioning, and other functions can be matched with the parking or waiting status.
[0076] 2) Weather conditions match.
[0077] For example, in severe weather conditions (such as heavy rain, fog, or snow), safety-related instruction types are matched with the severe weather conditions.
[0078] For example, when the weather is sunny or the in-car environment is stable, the types of entertainment control and information query commands can be matched with the sunny weather or in-car environment.
[0079] 3) Matching the in-vehicle environment.
[0080] For example, when the noise level inside the vehicle is too high, the volume adjustment command type will be matched with that noise level.
[0081] For example, when the temperature inside the vehicle is too high or too low, the type of air conditioning adjustment command will be matched with the scenario of the temperature being too high or too low inside the vehicle.
[0082] 4) User behavior matching.
[0083] For example, when it is detected that the user is focused on driving (such as at high speed, operating the steering wheel, etc.), safety-related instructions are matched with the scenario in which the user is focused on driving.
[0084] In some embodiments, when performing command type matching, scenario data from multiple different categories can be fused before matching the corresponding command type. For example, in-vehicle temperature sensor, humidity sensor, and outside weather data can jointly determine the user's air conditioning needs, while noise sensor and vehicle speed data can help the system determine whether enhanced filtering and processing of voice signals is required. Data fusion helps improve the accuracy of context perception, thereby improving the accuracy of user intent recognition.
[0085] Based on the above, the embodiments of this application can pre-construct the correspondence between different scenarios and different instruction types according to big data analysis, so that when the vehicle is identified as being in a certain target scenario, the instruction type matching the target scenario can be quickly located based on the preset correspondence, thereby improving work efficiency.
[0086] In step S20, obtaining the first instruction type corresponding to the target scenario from the database according to the target scenario includes:
[0087] Step S21. Obtain a first target identifier, wherein the first target identifier is used to characterize the target scenario;
[0088] Understandably, after classifying various scenarios based on big data analysis, each scenario can be assigned a first identifier. Since the first identifiers of each scenario are different, the scenario can be characterized by the first identifier, thereby reducing the amount of data processing. Similarly, after classifying various instruction types based on big data analysis, each instruction type can be assigned a second identifier. Since the second identifiers of each instruction type are different, the instruction type can be characterized by the second identifier, thereby reducing the amount of data processing.
[0089] Step S22. Obtain a second target identifier corresponding to the first target identifier from the database, wherein the second target identifier is used to characterize the first instruction type, the database includes a plurality of first identifiers and a second identifier corresponding to each first identifier, the first identifier is used to characterize the scenario, and the second identifier is used to characterize the instruction type.
[0090] Step S30. Adjust the priority of the first instruction type to be greater than the priority of other instruction types in the database;
[0091] Understandably, when a vehicle is identified as being in a target scenario, the priority of the first instruction type corresponding to that target scenario is adjusted to be higher than the priority of other instruction types. This is so that when the user's voice content is subsequently acquired, the first instruction type can be called first to recognize the user's voice content. By combining contextual awareness of the vehicle's internal and external environment and driving data, the accuracy of voice content recognition is higher and the response speed is faster.
[0092] It should be noted that the target scenario in which the vehicle is located may be dynamically changing during vehicle operation. Therefore, when the target scenario changes, the corresponding first command type may also change, and consequently, the priority between different command types may also change. For example, when driving at high speeds, commands related to driving safety, such as "accelerate," "decelerate," and "navigate to the next exit," will have lower priority than commands related to the entertainment system. Thus, the priority of the first command type can be dynamically adjusted according to the different target scenarios in which the vehicle is located. This ensures that subsequent recognition of the user's voice content incorporates the perception of the vehicle's current context, thereby improving the accuracy and response speed of speech recognition.
[0093] In some embodiments, adjusting the priority of the first instruction type to be greater than the priority of other instruction types in the database includes:
[0094] The target value of the first instruction type is adjusted to a first value, and the target values of the other instruction types are adjusted to a second value, wherein the first value is used to represent a first priority, the second value is used to represent a second priority, and the first priority is greater than the second priority.
[0095] It is understandable that after classifying various instruction types based on big data analysis, and when matching the first instruction type according to the target scenario, the target value of each instruction type can be dynamically adjusted. The target value is used to characterize the priority of the instruction type.
[0096] For example, the target value of the first instruction type can be adjusted to 0, and the target values of other instruction types can be adjusted to 1, where 0 has a higher priority than 1. Therefore, when acquiring a user's voice information, the higher-priority first instruction type can be prioritized for voice recognition, improving the accuracy and response speed of user voice information recognition.
[0097] Step S40. Obtain the user's first voice information, and based on the voice recognition model, preferentially call the first instruction type to recognize the first voice information to determine the first instruction corresponding to the first voice information.
[0098] Understandably, when acquiring a user's initial voice information, multimodal input data (such as gestures and touches) can be combined to improve the contextual awareness of the language recognition model. For example, when recognizing the voice command "turn on the air conditioner," if the model detects that the user is touching the air conditioner control area or making a related gesture, it can respond to the command faster and more accurately.
[0099] In some embodiments, obtaining the user's first voice information, and recognizing the first voice information by preferentially calling the first instruction type based on a speech recognition model, includes:
[0100] Step S41. Obtain the first voice information, wherein the first voice information includes a first task to be executed and a second task to be executed;
[0101] It is understandable that the user's initial voice input may include multiple concurrent tasks to be executed. For example, the user's voice input may include "turn on the windshield wipers, turn up the air conditioning, turn on the interior lights, etc." Therefore, for the voice recognition model, it is necessary to recognize multiple concurrent tasks to be executed and then output the target command corresponding to each task.
[0102] Step S42. Based on the speech recognition model, prioritize calling the first instruction type to recognize the first task to be executed;
[0103] When multiple concurrent tasks are pending, the speech recognition model prioritizes the identification of the first instruction model for the first task to be executed. This ensures that the most relevant instruction to the current vehicle scenario is accurately and quickly output, improving the efficiency of the speech recognition model. For example, if the current scenario is highway driving in the rain, then "turn on the windshield wipers" is the most relevant instruction type in the above-mentioned "turn on the windshield wipers, turn up the air conditioning, turn on the interior lights, etc.", and this task will be identified and responded to first, thus ensuring safe driving.
[0104] Step S43. Based on the speech recognition model, call the other instruction types to recognize other second tasks to be executed.
[0105] Therefore, when multiple tasks are running concurrently, the first instruction type can be prioritized to identify the first task to be executed based on the priority of the preset first instruction type. This ensures that the speech recognition model can prioritize outputting the target instruction that corresponds to both the target scenario and the user's voice, thereby improving the recognition accuracy and response speed of the speech recognition model.
[0106] In some embodiments, when there are multiple first tasks to be executed, the step of prioritizing the invocation of the first instruction type based on the speech recognition model to identify the first execution includes:
[0107] Step S421. Determine the timestamp of each of the first tasks to be executed;
[0108] Step S422. Based on the speech recognition model, the first instruction type is invoked first, and each of the first tasks to be executed is identified according to the order of the timestamps.
[0109] For example, in the target scenario of driving in low visibility in rainy weather, the user's first voice output includes multiple first tasks to be executed, such as "turn on the windshield wipers, refresh the navigation, turn on the fog lights," etc. It can be understood that the instruction types corresponding to each of the above first tasks to be executed are all instruction types that match the target scenario. Therefore, each task to be executed can be responded to in sequence according to the time of receipt of each first task to be executed, so as to meet the user's needs.
[0110] In some embodiments, the method further includes:
[0111] Step S50. Acquire multiple second voice information of the user in the target scenario within a certain period of time;
[0112] It is understandable that when a user repeatedly issues a second voice message for a specific target scenario, it indicates that the user has a personalized control need for that target scenario. Therefore, this application embodiment collects multiple second voice messages from the user in the target scenario over a period of time.
[0113] Step S51. Based on the speech recognition model, identify multiple second speech information to determine multiple second commands;
[0114] Step S52. If multiple second instructions are of the same instruction type and multiple second instructions do not belong to the first instruction type, then increase the priority of the second instruction type to which the multiple second instructions belong.
[0115] Understandably, if multiple second commands do not belong to the first command type, it indicates that the user has repeatedly issued commands with different priorities than the preset priority in the target scenario. For example, some users like to adjust the music when driving at low speeds, while others may focus more on air conditioning adjustments. Through long-term learning, the personalized priority configuration for each user can be gradually optimized. For another example, in urban road conditions, the preset first command type matching urban road conditions is entertainment-related. However, if the user repeatedly issues voice commands such as "turn off the music" in urban road conditions, it indicates that multiple recognized second commands do not match the preset command type. Therefore, to meet the user's personalized preference needs, this application embodiment adjusts the command types matching the target scenario and increases the priority of command types to adapt to the user's preferences and needs.
[0116] In some embodiments, the vehicle includes a voice acquisition device, and the method further includes, during the acquisition of the user's first voice information:
[0117] Obtain ambient noise intensity and / or vehicle speed;
[0118] If the ambient noise intensity is greater than or equal to a set noise intensity and / or the vehicle speed is greater than or equal to a set vehicle speed, then the noise reduction capability of the voice acquisition device is enhanced.
[0119] For example, the voice acquisition device includes an in-vehicle microphone for acquiring the user's voice information. When the ambient noise is high and / or the vehicle speed is high, the impact of wind noise and / or ambient noise on voice acquisition can be reduced by improving the microphone's sensitivity and noise reduction capability, thereby reducing voice recognition errors caused by noise, wind noise and other interference.
[0120] See Figure 2 The diagram shows a structural diagram of a vehicle voice recognition device according to an embodiment of this application.
[0121] like Figure 2 As shown, a second aspect of this application provides a vehicle voice recognition device 200, comprising:
[0122] The first identification unit 201 is used to identify the target scenario in which the vehicle is located based on one or more of the vehicle's internal environment data, external environment data, and driving data.
[0123] The acquisition unit 202 is configured to acquire a first instruction type corresponding to the target scenario from a database according to the target scenario, wherein the database includes multiple scenarios of the vehicle and instruction types corresponding to each scenario;
[0124] Adjustment unit 203 is used to adjust the priority of the first instruction type to be greater than the priority of other instruction types in the database;
[0125] The second recognition unit 204 is used to acquire the user's first voice information, and to recognize the first voice information by preferentially calling the first instruction type based on the voice recognition model, so as to determine the first instruction corresponding to the first voice information.
[0126] According to a third aspect of the embodiments of this application, a computer-readable storage medium is provided, the computer-readable storage medium storing at least one computer program instruction, the at least one computer program instruction being loaded and executed by a processor to perform the operation as described in any of the methods in the first aspect.
[0127] Computer-readable storage media may be portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the computer-readable storage medium of this application is not limited thereto. In this application, the readable storage medium may be any tangible medium that contains or stores a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.
[0128] A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.
[0129] Program code for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0130] See Figure 3 This is a schematic diagram of the structure of a computer system suitable for implementing the embodiments of this application in a vehicle.
[0131] According to a fourth aspect of the embodiments of this application, a vehicle is provided, including one or more processors and one or more memories, wherein at least one piece of program code is stored in the one or more memories, and the at least one piece of program code is loaded and executed by the one or more processors to perform the operation as performed by any of the methods in the first aspect.
[0132] like Figure 3 As shown, vehicle 400 is represented in the form of a general-purpose computing device. The components of vehicle 400 may include, but are not limited to: at least one processing unit 410, at least one storage unit 420, and a bus 430 connecting different system components (including storage unit 420 and processing unit 410).
[0133] The storage unit stores program code, which can be executed by the processing unit 410, causing the processing unit 410 to perform the steps described in the "Embodiment Method" section above according to various exemplary embodiments of this application.
[0134] Storage unit 420 may include readable media in the form of volatile storage units, such as random access memory (RAM) 421 and / or cache 422, and may further include read-only memory (ROM) 423.
[0135] Storage unit 420 may also include a program / utility 424 having a set (at least one) of program modules 425, such program modules 425 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.
[0136] Bus 430 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.
[0137] Vehicle 400 can also communicate with one or more external devices 500 (e.g., keyboard, pointing device, Bluetooth device, etc.), one or more devices that enable users to interact with vehicle 400, and / or any device that enables vehicle 400 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed through I / O (input / output) interface 450, which can also be connected to display unit 440 to display the communication content. Furthermore, vehicle 400 can communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 460. As shown, network adapter 460 communicates with other modules of vehicle 400 via bus 430. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with vehicle 400, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0138] The functions described herein can be implemented in hardware, software executed by a processor, firmware, or any combination thereof. If implemented in software executed by a processor, the functions can be stored as one or more instructions or codes on or transmitted via a computer-readable medium. Other examples and embodiments are within the scope and spirit of this invention and the appended claims. For example, due to the nature of software, the functions described above can be implemented using software executed by a processor, hardware, firmware, hardwired, or any combination thereof. Furthermore, the functional units can be integrated into a single processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit.
[0139] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between units or modules may be electrical or other forms.
[0140] The units described as separate components may or may not be physically separate. Similarly, the components of the control device may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0141] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0142] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A vehicle voice recognition method, characterized in that, include: Based on one or more of the vehicle’s internal environment data, external environment data, and driving data, identify the target scenario in which the vehicle is located. According to the target scenario, a first instruction type corresponding to the target scenario is obtained from the database, wherein the database includes multiple scenarios of the vehicle and instruction types corresponding to each scenario; Adjust the priority of the first instruction type to be greater than the priority of other instruction types in the database; The user's first voice information is obtained, and the first instruction type is preferentially called to recognize the first voice information based on the voice recognition model in order to determine the first instruction corresponding to the first voice information. The step of obtaining the user's first voice information, and prioritizing the invocation of the first instruction type to recognize the first voice information based on the speech recognition model, includes: Obtain the first voice information, wherein the first voice information includes a first task to be executed and a second task to be executed; The first instruction type is preferentially invoked based on the speech recognition model to identify the first task to be executed; The second task to be executed is identified by calling other instruction types based on the speech recognition model.
2. The method according to claim 1, characterized in that, The step of retrieving the first instruction type corresponding to the target scenario from the database according to the target scenario includes: Obtain a first target identifier, wherein the first target identifier is used to characterize the target scenario; Obtain a second target identifier corresponding to the first target identifier from the database, wherein the second target identifier is used to characterize the first instruction type, the database includes a plurality of first identifiers and a second identifier corresponding to each first identifier, the first identifier is used to characterize the scenario, and the second identifier is used to characterize the instruction type.
3. The method according to claim 1 or 2, characterized in that, Adjusting the priority of the first instruction type to be greater than the priority of other instruction types in the database includes: The target value of the first instruction type is adjusted to a first value, and the target values of the other instruction types are adjusted to a second value, wherein the first value is used to represent a first priority, the second value is used to represent a second priority, and the first priority is greater than the second priority.
4. The method according to claim 1, characterized in that, When there are multiple first tasks to be executed, the step of prioritizing the invocation of the first instruction type to identify the first tasks based on the speech recognition model includes: Determine the timestamp of each of the first tasks to be executed; The first instruction type is invoked first based on the speech recognition model, and each of the first tasks to be executed is identified according to the order of the timestamps.
5. The method according to claim 1, characterized in that, The method further includes: Acquire multiple second voice information of the user in the target scenario over a period of time; Based on the speech recognition model, multiple pieces of the second speech information are recognized to determine multiple second commands; If multiple second instructions belong to the same instruction type, and multiple second instructions do not belong to the first instruction type, then the priority of the second instruction type to which the multiple second instructions belong is increased.
6. The method according to claim 1, characterized in that, The vehicle includes a voice acquisition device, and the method further includes, when acquiring the user's first voice information: Obtain ambient noise intensity and / or vehicle speed; If the ambient noise intensity is greater than or equal to a set noise intensity and / or the vehicle speed is greater than or equal to a set vehicle speed, then the noise reduction capability of the voice acquisition device is enhanced.
7. A vehicle voice recognition device, characterized in that, include: The first identification unit is used to identify the target scenario in which the vehicle is located based on one or more of the vehicle's internal environment data, external environment data, and driving data. The acquisition unit is configured to acquire a first instruction type corresponding to the target scenario from a database according to the target scenario, wherein the database includes multiple scenarios of the vehicle and instruction types corresponding to each scenario; An adjustment unit is used to adjust the priority of the first instruction type to be greater than the priority of other instruction types in the database; The second recognition unit is used to acquire the user's first voice information, and to recognize the first voice information by preferentially calling the first instruction type based on the voice recognition model, so as to determine the first instruction corresponding to the first voice information. The step of obtaining the user's first voice information, and prioritizing the invocation of the first instruction type to recognize the first voice information based on the speech recognition model, includes: Obtain the first voice information, wherein the first voice information includes a first task to be executed and a second task to be executed; The first instruction type is preferentially invoked based on the speech recognition model to identify the first task to be executed; The second task to be executed is identified by calling other instruction types based on the speech recognition model.
8. A computer-readable storage medium storing at least one computer program instruction, the at least one computer program instruction being loaded and executed by a processor to perform the operation performed by the method as described in any one of claims 1-6.
9. A vehicle, characterized in that, It includes one or more processors and one or more memories, wherein at least one piece of program code is stored in the one or more memories, and the at least one piece of program code is loaded and executed by the one or more processors to perform the operation performed by the method as described in any one of claims 1-6.
Citation Information
Patent Citations
Vehicle-mounted voice recognition data processing method and system
CN109920429A
Decoding method in speech recognition scene and related device
CN117292685A