Method and device for issuing vehicle-mounted voice recognition model and vehicle
By deploying speech recognition models and test models in the vehicle system and comparing and filtering vehicle operating parameters, the accuracy problem of the vehicle speech recognition system in complex environments was solved, and the reliability and effectiveness of the model in practical applications were improved.
Patent Information
- Application Number
- CN202511691643.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-18
- Publication Date
- 2026-02-10
AI Technical Summary
Existing in-vehicle voice recognition systems have low accuracy in complex driving environments, and the reliability and effectiveness of the models are difficult to fully verify before large-scale application.
Deploy in-vehicle speech recognition models and speech recognition test models. By comparing the model output results with vehicle operating parameters, filter and update the speech recognition test models to ensure their reliability and effectiveness in practical applications.
It improves the reliability and effectiveness of speech recognition test models in practical applications, reduces network bandwidth consumption, reduces security risks, and lowers the verification cost in edge scenarios.
Smart Images

Figure CN121506136A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of automobiles, and in particular to a method and device for publishing a vehicle-mounted voice recognition model and a vehicle. BACKGROUND
[0002] With the development of automobile intelligence, the in-vehicle voice recognition system has become an important part of improving driving experience and operational convenience. In particular, with the rapid development of the visual-language-action model (VLA) combined with intelligent driving and voice recognition, the requirement for voice recognition accuracy will be higher.
[0003] However, the existing voice recognition system still has the problem of low recognition accuracy when facing complex driving environments, such as noise interference, different accents, or ambiguous instructions. Once the voice recognition is incorrect, it may lead to incorrect operation execution, affecting driving safety and user experience.
[0004] In the process of implementing the present application, the inventors have found that at least the following problems exist in the prior art: In order to continuously improve the accuracy of voice recognition, the model needs to be continuously optimized, but the reliability and effectiveness of the model in actual application are difficult to fully verify before large-scale application. SUMMARY
[0005] Therefore, the embodiments of the present application provide a method and device for publishing a vehicle-mounted voice recognition model and a vehicle, which can test the reliability and effectiveness of the voice recognition test model in actual application before large-scale application.
[0006] A method for publishing a vehicle-mounted voice recognition model, comprising:
[0007] deploying a vehicle-mounted voice recognition model published by a mass-produced vehicle and a voice recognition test model, and respectively receiving a voice instruction sent by a user in the vehicle;
[0008] comparing a first voice recognition result of the voice instruction output by the vehicle-mounted voice recognition model and a second voice recognition result of the voice instruction output by the voice recognition test model to obtain model comparison data;
[0009] detecting the second voice recognition result to obtain running detection data using a running parameter of the vehicle, and providing the model comparison data and the running detection data to a background, and in a case where the model comparison data and the running detection data both satisfy a publishing condition, publishing the voice recognition test model as a vehicle-mounted voice recognition model.
[0010] The comparison of the first voice recognition result of the voice instruction output by the vehicle-mounted voice recognition model and the second voice recognition result of the voice instruction output by the voice recognition test model comprises:
[0011] comparing the speech recognition text in the first speech recognition result with the test recognition text in the second speech recognition result to obtain text comparison data;
[0012] comparing the vehicle operation instruction in the speech recognition text with the test operation instruction in the test recognition text to obtain operation comparison data.
[0013] The vehicle operation instruction comprises a vehicle operation action, a vehicle component to which the vehicle operation action is directed, and a vehicle operation parameter for the vehicle component;
[0014] The test operation instruction comprises an operation test action, a vehicle component to which the operation test action is directed, and an operation test parameter for the vehicle component;
[0015] The comparison of the vehicle operation instruction in the speech recognition text with the test operation instruction in the test recognition text comprises:
[0016] The comparison of the vehicle operation action with the operation test action, the comparison of the vehicle component to which the vehicle operation action is directed with the vehicle component to which the operation test action is directed, and the comparison of the vehicle operation parameter for the vehicle component with the operation test parameter for the vehicle component.
[0017] The detection of the second speech recognition result based on the running parameter of the vehicle to obtain running detection data comprises:
[0018] Based on the driving parameter and / or the equipment parameter in the running parameter of the vehicle, the logical reasonableness of the operation test action in the second speech recognition result is detected, and the running detection data comprises a detection result for the logical reasonableness, the driving parameter and / or the equipment parameter.
[0019] Further comprising:
[0020] In the case where the file comparison data indicates that the speech recognition text is similar to the test recognition text, and the operation comparison data indicates that the vehicle operation instruction and the test operation instruction are the same, the step of detecting the second speech recognition result based on the running parameter of the vehicle to obtain running detection data is executed.
[0021] Further comprising:
[0022] In the case where the model comparison data indicates that the first speech recognition result is inconsistent with the second speech recognition result, and / or the running detection data indicates that the second speech recognition result is logically unreasonable, the in-vehicle speech data of a preset time period before and after the time point corresponding to the speech instruction and the time point corresponding to the speech instruction are obtained and provided to the background.
[0023] Further comprising:
[0024] In a case where the model comparison data indicates that the first speech recognition result is consistent with the second speech recognition result, and the running detection data indicates that the second speech recognition result is logically reasonable, obtaining and providing in-vehicle speech data at a time point corresponding to the speech instruction to a background.
[0025] Further comprising:
[0026] In a case where the background determines that one of the model comparison data and the running detection data does not satisfy a release condition, updating the speech recognition test model by using the model comparison data and the running detection data, and deploying the updated speech recognition test model in a mass-produced vehicle;
[0027] And / or,
[0028] The speech recognition test model is a new model independent of the vehicle-mounted speech recognition model, or the speech recognition test model is an upgraded version of the vehicle-mounted speech recognition model.
[0029] According to a second aspect of an embodiment of the present application, there is provided a device for releasing a vehicle-mounted speech recognition model, comprising:
[0030] A speech module is configured to deploy a vehicle-mounted speech recognition model released by a mass-produced vehicle and a speech recognition test model, and receive a speech instruction sent by a user in a vehicle respectively;
[0031] A comparison module is configured to compare a first speech recognition result output by the vehicle-mounted speech recognition model for the speech instruction and a second speech recognition result output by the speech recognition test model for the speech instruction, to obtain model comparison data;
[0032] A detection module is configured to detect running detection data of the second speech recognition result by using a running parameter of the vehicle, and provide the model comparison data and the running detection data to a background. In a case where the background determines that the model comparison data and the running detection data both satisfy a release condition, the speech recognition test model is released as a vehicle-mounted speech recognition model.
[0033] According to a third aspect of an embodiment of the present application, there is provided a vehicle comprising the device for releasing a vehicle-mounted speech recognition model.
[0034] According to a fourth aspect of an embodiment of the present application, there is provided an electronic device for releasing a vehicle-mounted speech recognition model, comprising:
[0035] One or more processors;
[0036] a storage device storing one or more programs,
[0037] When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described above.
[0038] According to a fifth aspect of an embodiment of the present application, a computer readable medium is provided, having stored thereon a computer program which, when executed by a processor, implements the method as described above.
[0039] An embodiment of the above-mentioned application has the following advantages or beneficial effects: The vehicle-mounted speech recognition model deployed in the mass-produced vehicle and the speech recognition test model are respectively received by the user in the vehicle; the vehicle-mounted speech recognition model is compared with the speech recognition test model; the logic of the second speech recognition result is detected by the running parameters of the vehicle. Whether to release the speech recognition test model is determined by the comparison data of the model and the running detection data. Before the speech recognition model is released, the vehicle-mounted speech recognition model deployed in the mass-produced vehicle and the running parameters of the vehicle are compared, so as to improve the reliability and effectiveness of the speech recognition test model in actual application.
[0040] Further effects of the above-mentioned non-conventional optional mode will be described in the following combined with the specific embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0041] The accompanying drawings are used to better understand the present application, and do not constitute an improper limitation on the present application. Among them:
[0042] Figure 1 is the main flowchart of the method for releasing the vehicle-mounted speech recognition model according to an embodiment of the present application;
[0043] Figure 2 is the flowchart of comparing the vehicle-mounted speech recognition model and the speech recognition test model according to an embodiment of the present application;
[0044] Figure 3 is the main structure diagram of the device for releasing the vehicle-mounted speech recognition model according to an embodiment of the present application;
[0045] Figure 4 is an exemplary system architecture diagram to which the embodiment of the present application can be applied;
[0046] Figure 5 is a structure diagram of a computer system of a terminal device or a server suitable for implementing the embodiment of the present application. DETAILED DESCRIPTION
[0047] Exemplary embodiments of the present application are described herein with reference to the accompanying drawings, which are meant to be exemplary in nature, and include various details intended to facilitate understanding of the application. Accordingly, one skilled in the art should realize that various changes and modifications can be made to the embodiments described herein, without departing from the scope and spirit of the application. Similarly, it will be appreciated that, in the interest of clarity and briefness, the following description has omitted description of well-known functions and constructions.
[0048] To improve the reliability and effectiveness of the speech recognition test model in practical application before large-scale application, the technical solutions in the following embodiments of the present application can be used.
[0049] Referring to Figure 1 , Figure 1 is a main flowchart of a method for publishing a vehicle-mounted speech recognition model according to an embodiment of the present application. The vehicle-mounted speech recognition model and the speech recognition test model are compared, and the second speech recognition result of the speech recognition test model is detected by using the running parameters of the vehicle. As shown in Figure 1 Specifically, the method comprises the following steps:
[0050] S101, the published vehicle-mounted speech recognition model and the speech recognition test model are deployed in the mass-produced vehicle to receive the voice instructions sent by the user in the vehicle.
[0051] In the embodiments of the present application, the published vehicle-mounted speech recognition model and the speech recognition test model are deployed in the mass-produced vehicle. The vehicle-mounted speech recognition model is used to receive and recognize the voice instructions sent by the user in the vehicle. The speech recognition test model is a model for testing.
[0052] The speech recognition test model is a new model independent of the vehicle-mounted speech recognition model. Alternatively, the speech recognition test model is an upgraded version of the vehicle-mounted speech recognition model. As an example, the version of the vehicle-mounted speech recognition model is 1.0, and the version of the speech recognition test model is 2.0.
[0053] The published vehicle-mounted speech recognition model and the speech recognition test model are deployed in the mass-produced vehicle to compare the published model with the unpublished model. Relying on various actual use scenarios of the mass-produced vehicle, a large amount of edge scene voice data can be continuously collected. For example, extreme noise, rare accents, and ambiguous instructions. The above voice data comes from real driving environment, and the scale of the voice data naturally grows with the use of the mass-produced vehicle, breaking through the limitations of fixed test cases, providing rich voice data for the training of the speech recognition test model, and thus being able to optimize the speech recognition capability in extreme scenarios.
[0054] The user sends a voice instruction in the vehicle, and the voice instruction includes not only the user voice but also environmental noise. Specifically, the user sends a voice instruction in the vehicle, and a voice collection device in the vehicle, such as a microphone array, captures a voice signal and converts the voice signal into an electrical signal. The electrical signal is synchronously distributed to a vehicle-mounted voice recognition model and a voice recognition test model through a vehicle-mounted high-speed bus.
[0055] The vehicle-mounted voice recognition model and the voice recognition test model share a hardware device but independently occupy a computing thread to implement voice recognition. For example, the hardware device includes a voice collection device and a memory. The voice recognition test model does not intervene in vehicle control during voice recognition and only stores a voice recognition result in a preset storage area.
[0056] S102, comparing a first voice recognition result for the voice instruction output by the vehicle-mounted voice recognition model and a second voice recognition result for the voice instruction output by the voice recognition test model to obtain model comparison data.
[0057] The vehicle-mounted voice recognition model and the voice recognition test model perform voice recognition respectively, compare the first voice recognition result of the vehicle-mounted voice recognition model and the second voice recognition result output by the voice recognition test model for the voice instruction, and take the comparison result as model comparison data.
[0058] As an example, the model comparison data includes a comparison result, a version identifier of the vehicle-mounted voice recognition model, and a version identifier of the voice recognition test model. In this way, the server side can update the voice recognition model according to the version identifiers and the comparison result.
[0059] S103, detecting the second voice recognition result to obtain running detection data with a running parameter of the vehicle, and providing the model comparison data and the running detection data to a background. If the model comparison data and the running detection data both satisfy a release condition, the voice recognition test model is released as the vehicle-mounted voice recognition model.
[0060] In the embodiments of the present application, the purpose of sending the voice instruction is to control the vehicle, and the second voice recognition result can be detected by the running parameter of the vehicle to detect the logical consistency of the second voice recognition result and the running parameter of the vehicle.
[0061] Specifically, the logical rationality of the operation action in the second voice recognition result is detected based on a driving parameter and / or a device parameter in the running parameter of the vehicle, and the running detection data includes the logical rationality, the driving parameter, and / or the device parameter.
[0062] The operation parameters of the vehicle refer to driving parameters and device parameters. As an example, the driving parameters include one or more of speed, gear and driving mode. The device parameters include device name and regulation parameters.
[0063] For example, the driving parameter includes the parking gear, and the second voice recognition result is to open the cruise control. The cruise control function is only started when the vehicle is running and the vehicle speed reaches the minimum speed of the cruise control. If the second voice recognition result is not logically consistent with the current driving parameter, the logical rationality includes logical inconsistency.
[0064] For example, the device parameter includes the device name: air conditioner; and the regulation parameter: off. The second voice recognition result is to increase the vehicle speed. If the second voice recognition result is not logically consistent with the current device parameter, the logical rationality includes logical inconsistency.
[0065] For example, the driving parameter includes the speed: 5 km / h, and the second voice recognition result is to increase the speed to 40 km / h. If the second voice recognition result is logically consistent with the current driving parameter, the logical rationality includes logical consistency.
[0066] In the above embodiment of the present application, the vehicle-mounted voice recognition model deployed in the mass-produced vehicle and the voice recognition test model are compared, and the voice recognition results of the vehicle-mounted voice recognition model and the voice recognition test model are compared. The logicality of the second voice recognition result is detected by the operation parameters of the vehicle. Whether the voice recognition test model is released is determined by comparing the model data and the operation detection data. Before the voice recognition model is released, the vehicle-mounted voice recognition model deployed in the mass-produced vehicle and the operation parameters of the vehicle are compared, and the reliability and effectiveness of the voice recognition test model in actual application are further determined.
[0067] Referring to Figure 2 That is 200, Figure 2 is a flowchart of comparing the vehicle-mounted voice recognition model and the voice recognition test model according to the embodiment of the present application. Specifically, the following steps are included:
[0068] S201, compare the voice recognition text in the first voice recognition result with the test recognition text in the second voice recognition result to obtain text comparison data.
[0069] The vehicle-mounted voice recognition model outputs the voice recognition text in the first voice recognition result. As an example, the first voice recognition result includes the voice recognition text and the sending time of the voice instruction. The voice recognition test model outputs the test recognition text in the second voice recognition result. The vehicle-mounted voice recognition model has higher stability than the voice recognition test model, and can compare the voice recognition text and the test recognition text.
[0070] In one embodiment of the present application, the text similarity between the speech recognition text and the test recognition text is determined using edit distance. Edit distance refers to the minimum number of single-character editing operations required to convert one text into another, including one or more of the following: insertion, deletion, and replacement.
[0071] If the text similarity between the speech recognition text and the test recognition text is greater than or equal to the similarity threshold, it is determined that the speech recognition text is similar to the test recognition text. The text comparison data in the model comparison data includes the test recognition text, and the speech recognition text is similar to the test recognition text.
[0072] If the text similarity between the speech recognition text and the test recognition text is less than the similarity threshold, it is determined that the speech recognition text is different from the test recognition text. The text comparison data in the model comparison data includes the speech recognition text, the test recognition text, and the speech recognition text is different from the test recognition text.
[0073] As an example, the similarity threshold is set to 95%. The speech recognition text: Turn on the air conditioner to 23 degrees. The test recognition text: Turn on the air conditioner to 24 degrees. The text length is 7, the edit distance is 1, and the similarity is (7-1) / 7≈85.7%<95%, so it is determined that the speech recognition text is different from the test recognition text.
[0074] S202, compare the vehicle operation instruction in the speech recognition text with the test operation instruction in the test recognition text to obtain operation comparison data.
[0075] In the case of controlling the vehicle using voice instructions, the vehicle operation instruction in the speech recognition text is usually used to control the vehicle. The vehicle operation instruction in the speech recognition text and the test operation instruction in the test recognition text can be compared to obtain operation comparison data.
[0076] As an example, the vehicle operation instruction of the speech recognition text is obtained through key information extraction and intent recognition, and the test operation instruction of the test recognition text is obtained.
[0077] In one embodiment of the present application, the vehicle operation instruction includes: a vehicle operation action, a vehicle component to which the vehicle operation action is directed, and a vehicle operation parameter for the vehicle component; and the test operation instruction includes: an operation test action, a vehicle component to which the operation test action is directed, and an operation test parameter for the vehicle component.
[0078] The vehicle operation action corresponds to the operation test action, the vehicle component to which the vehicle operation action is directed corresponds to the vehicle component to which the operation test action is directed, the operation test parameter for the vehicle component corresponds to the vehicle operation parameter for the vehicle component, and the parameters are compared. That is, the vehicle operation action is compared with the operation test action, the vehicle component to which the vehicle operation action is directed is compared with the vehicle component to which the operation test action is directed, and the vehicle operation parameter for the vehicle component is compared with the operation test parameter for the vehicle component.
[0079] Specifically, the operation action comparison includes comparing whether the vehicle operation action and the operation test action are completely identical. For example, the vehicle operation action is "turn on", and the operation test action is "turn on". The vehicle operation action and the operation test action are identical. The vehicle operation action is "turn on", and the operation test action is "turn off". The vehicle operation action and the operation test action are not identical.
[0080] The vehicle component comparison includes comparing whether the vehicle component to which the vehicle operation action is directed and the vehicle component to which the operation test action is directed are identical. For example, the vehicle component to which the vehicle operation action is directed is "air conditioner", and the vehicle component to which the operation test action is directed is "air conditioner". They are considered identical. The vehicle component to which the vehicle operation action is directed is "air conditioner", and the vehicle component to which the operation test action is directed is "navigation". They are considered different.
[0081] The test parameter comparison includes comparing whether the vehicle operation parameter for the vehicle component and the operation test parameter for the vehicle component are identical.
[0082] For specific parameters such as temperature and volume, an allowable error range of the parameters is set. For example, the allowable error of the temperature parameter is ±1℃, and the allowable error of the volume parameter is ±1. When the parameter is within the allowable error range, the vehicle operation parameter for the vehicle component and the operation test parameter for the vehicle component are identical. When the parameter is outside the allowable error range, the vehicle operation parameter for the vehicle component and the operation test parameter for the vehicle component are different.
[0083] As an example, the vehicle operation action and the operation test action are identical, the vehicle component to which the vehicle operation action is directed and the vehicle component to which the operation test action is directed are identical, and the vehicle operation parameter for the vehicle component and the operation test parameter for the vehicle component are identical. Therefore, the vehicle operation instruction and the test operation instruction are identical.
[0084] As another example, one parameter in the vehicle operation instruction and the corresponding parameter in the test operation instruction are different. Therefore, the vehicle operation instruction and the test operation instruction are different.
[0085] The operation comparison data in the model comparison data includes the vehicle operation instruction, the test operation instruction, and the comparison result of the vehicle operation instruction and the test operation instruction.
[0086] In one embodiment of the present invention, the prerequisite for obtaining operation detection data by executing the second speech recognition result of vehicle operation parameter detection is that the speech recognition text is similar to the test recognition text, and the vehicle operation command and the test operation command are the same. By fulfilling this prerequisite, the vehicle's computational load is reduced, and the amount of data sent to the backend is decreased. Therefore, if the speech recognition text and the test recognition text are not similar, or if the vehicle operation command and the test operation command are different, there is no need to send data to the backend.
[0087] In the above embodiments, the speech recognition text and test recognition text are compared, as are the vehicle operation instructions and test operation instructions, to eliminate invalid text and semantic discrepancies. Vehicle operating parameters are used for scene verification to eliminate logically contradictory data. After filtering, invalid data is significantly reduced, thus lowering network bandwidth consumption and providing the server with more valid data to optimize the speech recognition model.
[0088] In one embodiment of the present invention, in order to update the speech recognition test model based on data, model comparison data and running test data can be uploaded to the background to update the speech recognition test model in the background.
[0089] In addition, to improve the relevance of the background update of the voice recognition test model, in-vehicle voice data can also be uploaded.
[0090] If the first and second speech recognition results are inconsistent, and / or the logic of the second speech recognition result is unreasonable, it indicates that the speech recognition test model has a high error rate in recognizing voice commands. Therefore, it is necessary to send in-vehicle voice data within a preset time period to the backend. The preset time period is the period before and after the time point corresponding to the voice command. The in-vehicle voice data within the preset time period involves the application scenario of the voice command, and the speech recognition test model can be updated according to the application scenario using the aforementioned in-vehicle voice data. As an example, the preset time period is 5 seconds.
[0091] The first speech recognition result is consistent with the second speech recognition result, and the second speech recognition result is logically sound. This means that the speech recognition test model can correctly recognize voice commands, and it only needs to send the in-vehicle voice data corresponding to the time point of the voice command to the backend. The backend only updates the speech recognition test model based on the in-vehicle voice data at that time point.
[0092] In the above embodiments, in-vehicle voice data is filtered by model comparison data and operational detection data, thereby improving the relevance and effectiveness of the transmitted in-vehicle voice data.
[0093] In an embodiment of the present invention, the backend can determine whether the release conditions are met based on model comparison data and operation detection data. If the release conditions are met, the speech recognition test model can be released as an in-vehicle speech recognition model.
[0094] As an example, the release conditions include: the accuracy of the model comparison data and the running detection data are both above 80%, and there is no safety level error, then the speech recognition test model is gradually used as the vehicle-mounted speech recognition model in all mass-produced vehicles. The safety level error is an instruction recognition error related to the safety control of the vehicle.
[0095] In the case where the model comparison data or the running detection data does not meet the release conditions, the backend adjusts and updates the speech recognition test model using the model comparison data, the running detection data, and the voice instructions. The updated speech recognition test model is deployed in mass-produced vehicles for continued testing.
[0096] As an example, the speech recognition test model is deployed in 20% of the test vehicles of different vehicle models, different regions, and different driving habits, and the backend continues to receive the model comparison data and the running detection data sent by the test vehicles to determine whether the release conditions are met.
[0097] As an example, when the comparison results of the comparison data are inconsistent, the 5-second voice segments before and after the instruction (including 1-second ambient sound before and after the instruction) can be automatically intercepted, and the current software log (including the model version, recognition confidence), the "to be verified" voice recognition module log, and the vehicle state snapshot (including vehicle speed, engine speed, microphone array signal-to-noise ratio, etc. Parameters) are encapsulated and returned to the backend.
[0098] Further, different stages are used to perform algorithm gray release and verification:
[0099] Stage one: based on the difference data returned by the background, the "to be verified" algorithm is optimized and adjusted to generate a new "to be verified" algorithm version.
[0100] Stage two: the optimized "to be verified" algorithm is first deployed on 20% of the test vehicles, which cover different vehicle models, user groups in different use areas, and driving habits. During the test period, the algorithm runs in parallel with the current algorithm as a "to be verified" algorithm, only performs recognition result comparison, does not participate in vehicle control, and performs small-scale test deployment.
[0101] Stage three: A / B test comparison and analysis: continuously monitor the recognition difference rate, misrecognition rate, and other key indicators of the optimized "to be verified" algorithm and the current algorithm in the test vehicles. If the accuracy of the "to be verified" algorithm in the returned data is above 80% and there is no safety level error (such as instruction recognition error related to vehicle safety control), it enters the next stage; otherwise, return to stage one for re-optimization.
[0102] Phase four: expand the deployment range: when the conditions of phase three are met, the optimized "to be verified" algorithm is deployed to cover 20% of the vehicles, and monitoring and comparative analysis are continued. If the performance is stable, it will be gradually expanded to cover all vehicles.
[0103] In the above embodiment, the model comparison data and the running detection data are used as the basis for judging whether the release conditions are met through different stages, which eliminates the safety risks brought by directly replacing the vehicle-mounted voice recognition model with the voice recognition test model, and uses the actual scene voice recognition test model of the vehicle. While ensuring the continuous iteration of the voice recognition test model, the driving safety line is also strengthened.
[0104] The performance of the voice edge scene is affected by many factors such as vehicle speed, such as high-speed wind noise, vehicle hardware, such as microphone array sensitivity difference, and environmental noise, etc. A large number of test vehicles are needed to cover different scenes for voice recognition test model testing. In the embodiment of the present application, based on the large-scale mass-produced vehicles, the edge scene is continuously discovered under the premise of high efficiency and low cost, which continuously improves the accuracy and reliability of the voice recognition of the verified recognition model. Without additional investment in test resources, and can cover the difference characteristics of different vehicle models, speeds and environments at the same time, which greatly reduces the cost of edge scene verification. It not only solves the technical pain points of insufficient coverage of edge scenes in research and development, but also realizes the efficient evolution and reliable landing of the voice recognition test model through fine data uploading, providing a full-link solution for the continuous optimization of the vehicle-mounted voice recognition model.
[0105] Reference Figure 3 , Figure 3 is the main structure diagram of the device for releasing the vehicle-mounted voice recognition model according to the embodiment of the present application. The device for releasing the vehicle-mounted voice recognition model can realize the method for releasing the vehicle-mounted voice recognition model, such as Figure 3 as shown in 300, the device for releasing the vehicle-mounted voice recognition model specifically includes:
[0106] The voice module 301 is used to deploy the vehicle-mounted voice recognition model and the voice recognition test model released by the mass-produced vehicles, and respectively receives the voice instructions sent by the user in the vehicle;
[0107] The comparison module 302 is used to compare the first voice recognition result of the voice instruction output by the vehicle-mounted voice recognition model and the second voice recognition result of the voice instruction output by the voice recognition test model, and obtains the model comparison data;
[0108] The detection module 303 is configured to detect operation detection data from the second speech recognition result based on an operation parameter of the vehicle, and provide the model comparison data and the operation detection data to a background. In a case where the model comparison data and the operation detection data both satisfy a release condition, the speech recognition test model is released as a vehicle-mounted speech recognition model.
[0109] In an embodiment of the present application, the comparison module 302 is configured to compare speech recognition text in the first speech recognition result with test recognition text in the second speech recognition result to obtain text comparison data.
[0110] The vehicle operation instruction in the speech recognition text is compared with the test operation instruction in the test recognition text to obtain operation comparison data.
[0111] In an embodiment of the present application, the vehicle operation instruction includes a vehicle operation action, a vehicle component to which the vehicle operation action is directed, and a vehicle operation parameter for the vehicle component.
[0112] The test operation instruction includes an operation test action, a vehicle component to which the operation test action is directed, and an operation test parameter for the vehicle component.
[0113] The comparison module 302 is configured to compare the vehicle operation action with the operation test action, compare the vehicle component to which the vehicle operation action is directed with the vehicle component to which the operation test action is directed, and compare the vehicle operation parameter for the vehicle component with the operation test parameter for the vehicle component.
[0114] In an embodiment of the present application, the detection module 303 is configured to detect logical reasonableness of the operation test action in the second speech recognition result based on a driving parameter and / or a device parameter in the operation parameter of the vehicle, and the operation detection data includes a detection result for the logical reasonableness, the driving parameter, and / or the device parameter.
[0115] In an embodiment of the present application, the detection module 303 is configured to perform the step of detecting operation detection data from the second speech recognition result based on an operation parameter of the vehicle in a case where the file comparison data indicates that the speech recognition text is similar to the test recognition text, and the operation comparison data indicates that the vehicle operation instruction and the test operation instruction are the same.
[0116] In an embodiment of the present application, the detection module 303 is configured to, when the model comparison data indicates that the first speech recognition result is inconsistent with the second speech recognition result and / or the running detection data indicates that the second speech recognition result is logically unreasonable, obtain and provide the in-vehicle voice data corresponding to the time point before and after the voice instruction to the background.
[0117] In an embodiment of the present application, the detection module 303 is configured to, when the model comparison data indicates that the first speech recognition result is consistent with the second speech recognition result and the running detection data indicates that the second speech recognition result is logically reasonable, obtain and provide the in-vehicle voice data corresponding to the time point of the voice instruction to the background.
[0118] In an embodiment of the present application, the detection module 303 is configured to, when the background determines that one of the model comparison data and the running detection data does not satisfy the release condition, update the speech recognition test model using the model comparison data and the running detection data, and deploy the updated speech recognition test model in the mass-produced vehicle.
[0119] In an embodiment of the present application, the speech recognition test model is a new model independent of the vehicle-mounted speech recognition model or the speech recognition test model is an upgraded version of the vehicle-mounted speech recognition model.
[0120] The device for releasing the vehicle-mounted speech recognition model in the embodiments of the present application can be applied to a vehicle.
[0121] Figure 4 An exemplary system architecture 400 to which the method for releasing the vehicle-mounted speech recognition model or the device for releasing the vehicle-mounted speech recognition model in the embodiments of the present application can be applied is shown.
[0122] As shown in Figure 4 The vehicle system architecture 400 can include various systems, such as a driving control system 401, a power system 402, a sensor system 403, a control system 404, an auxiliary lane changing system 405, one or more peripheral devices 406, a power supply 407, a computer system 408, and a user interface 409. The method for releasing the vehicle-mounted speech recognition model provided in the embodiments of the present application can be implemented by interacting with each of the above systems, or can be implemented by an external device controlling the above systems or a robot driving a vehicle operating the above systems. Alternatively, the vehicle system architecture 400 can include more or fewer systems, and each system can include multiple elements. In addition, each system and element of the vehicle system architecture 400 can be interconnected by wire or wirelessly.
[0123] The vehicle system architecture 400 includes a driving control system 401, which can be in a full or partial autonomous driving mode. For example, the driving control system 401 can automatically control the vehicle to travel according to a control signal or a control instruction without human interaction or through interaction with an external device or a robot that drives the vehicle.
[0124] The power system 402 can include components that provide power motion for the vehicle. For example, the power system 402 can include an engine, an energy source, a transmission, wheels, tires, etc. The engine can be a combustion engine, an electric motor, an air compression engine, or a combination of other types of engines, such as a hybrid engine composed of a gasoline engine and an electric motor, a hybrid engine composed of a combustion engine and an air compression engine. The engine converts the energy source into mechanical energy to provide the transmission. Examples of the energy source can include gasoline, diesel, other petroleum-based fuels, propane, other compressed gas-based fuels, ethanol, solar panels, batteries, and other sources of electricity. The energy source can also provide energy for other systems of the vehicle. In addition, the transmission can include a gearbox, a differential, a drive shaft, a clutch, etc.
[0125] The sensor system 403 can include sensors that sense the environment around the vehicle, such as sensors that sense whether there are obstacles around, etc., and pressure sensors that sense whether there are passengers on the seats, etc. For example, a positioning system (which can be a global positioning system (GPS) system, a Beidou system, or other positioning systems), a radar, a laser rangefinder, an inertial measurement unit (IMU), a camera, etc. The positioning system can be used to locate the geographical position of the vehicle. The IMU is used to sense the changes in the position and orientation of the vehicle based on inertial acceleration. In one embodiment, the IMU can be a combination of an accelerometer and a gyroscope. The radar can use radio signals to sense objects within the environment around the vehicle. In some embodiments, in addition to sensing objects, the radar can also be used to sense the speed and / or direction of travel of the objects, etc.
[0126] To detect environmental information, objects, etc. outside the vehicle, the camera, etc. can be configured at appropriate positions outside the vehicle. For example, to obtain environmental images of the side of the vehicle, the camera can be on the side mirror of the vehicle. The camera can be a still or video camera.
[0127] The control system 404 can include software systems that implement vehicle driving control, such as systems that perform vehicle surroundings analysis, systems that perform seatbelt pretensioning, systems that perform route planning, systems that avoid obstacles, vision systems that perform image analysis, and the like. The control system 404 can also include hardware systems such as throttle, steering wheel systems, seatbelt systems, airbag systems, peripheral devices (e.g., projection devices, displays, etc.), and the like. Additionally, the control system 404 can include components in addition to, or instead of, those shown and described. Or some of the components shown above can be reduced.
[0128] Additionally, the control system 404 can interact with external sensors, other autonomous driving devices, other computer systems, or users through peripheral devices 406. The peripheral devices 406 can include wireless communication systems, onboard computers, microphones and / or speakers, cameras, and projectors, among others.
[0129] In some embodiments, the peripheral devices 406 provide a means for a user of the control system 404 to interact with a user interface. For example, an onboard computer can provide information to a user of the vehicle. The user interface can also operate the onboard computer to receive input from the user. The onboard computer can be operated through a touchscreen. In other cases, the peripheral devices can provide a means for communicating with other devices located within the vehicle. For example, a microphone can receive audio (e.g., voice commands or other audio input) from a user of the control system. Similarly, a speaker can output audio to a user of the control system.
[0130] The wireless communication system can wirelessly communicate with one or more devices directly or via a communication network. For example, the wireless communication system can communicate with devices using cellular networks, WiFi and wireless local area networks (WLAN) networks, among others, and can also communicate directly with devices using infrared links, Bluetooth, or ZigBee. Other wireless protocols, such as various autonomous driving communication systems, among others, can also be used.
[0131] The power source 407 can provide power to various components of the vehicle. The power source 407 can be a rechargeable lithium battery or a lead-acid battery.
[0132] Some or all of the functionality to implement vehicle control for ramp entry scenarios is controlled by the computer system 408. The computer system 408 can include at least one processor that executes instructions stored in a non-transitory computer-readable medium, such as a memory. The computer system 408 provides the control system described above with the executable code to implement vehicle control for ramp entry scenarios.
[0133] The processor can be any conventional processor, such as a commercially available central processing unit (CPU). Alternatively, the processor can be a dedicated device such as an application specific integrated circuit (ASIC) or other hardware-based processor. Those skilled in the art will appreciate that the processor, computer, or memory can actually comprise a plurality of processors, computers, or memories, which can or can not be stored in the same physical housing. For example, the memory can be a hard drive or other storage medium located in a housing different from that of the computer. Accordingly, reference to a processor or computer will be understood to encompass reference to a collection of processors or computers or memories, which can or can not operate in parallel. Rather than using a single processor to perform the steps described herein, such as steering assembly and deceleration assembly, some components can each have their own processor that only performs determinations related to the functionality specific to the component.
[0134] A user interface 409 for providing information to or receiving information from a user of the vehicle. Optionally, the user interface 409 can include one or more input / output devices in the set of peripheral devices 406, such as a wireless communication system, an on-board computer, a microphone, and a speaker.
[0135] It should be understood that the above components are only an example, in actual application, components in each module or system described above can be added or deleted according to actual needs, Figure 4 It should not be understood as a limitation to the embodiments of the present application.
[0136] The following refers to Figure 5 which shows a structural schematic diagram of a computer system 500 of a terminal device suitable for implementing the embodiments of the present application. Figure 5 The terminal device shown is only an example and should not bring any limitation to the functions and use range of the embodiments of the present application.
[0137] As Figure 5 shown, the computer system 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 502 or programs loaded from a storage portion 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the system 500 are also stored. The CPU 501, ROM 502, and RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0138] The following components are connected to the I / O interface 505: an input part 506 including a keyboard, a mouse, etc.; an output part 507 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage part 508 including a hard disk, etc.; and a communication part 509 including a network interface card such as a LAN card, a modem, etc. The communication part 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as necessary. A removable medium 511 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is attached to the drive 510 as necessary, so that a computer program read out therefrom is installed in the storage part 508 as necessary.
[0139] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network by the communication part 509, and / or installed from the removable medium 511. When the computer program is executed by the central processing unit (CPU) 501, the above-described functions defined in the system of the present disclosure are executed.
[0140] It should be noted that the computer-readable medium shown in the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or component. In the present application, the computer-readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to an electromagnetic signal, an optical signal or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or component. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0141] The flowcharts and block diagrams in the drawings illustrate the possible implementation architectures, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each block in the flowcharts or block diagrams can represent a module, a program segment or a part of code containing one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in different order than that shown in the drawings. For example, two blocks that are shown in succession can actually be executed substantially in parallel, and sometimes in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams or flowcharts, and the combination of blocks in the block diagrams or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0142] The modules described in the embodiments of the present invention can be implemented in software or hardware. The described modules can also be housed in a processor; for example, a processor can be described as including a voice module, a comparison module, and a detection module. The names of these modules do not necessarily limit the module itself; for example, a voice module can also be described as "used to deploy in mass-produced vehicles a released in-vehicle voice recognition model and a voice recognition test model, respectively receiving voice commands sent by the user inside the vehicle."
[0143] In another aspect, the present invention also provides a computer-readable medium, which may be included in the device described in the above embodiments; or it may exist independently and not assembled into the device. The computer-readable medium carries one or more programs, which, when executed by the device, cause the device to include:
[0144] The in-vehicle voice recognition model and voice recognition test model deployed in mass-produced vehicles receive voice commands sent by users inside the vehicle.
[0145] The first speech recognition result for the voice command output by the vehicle speech recognition model and the second speech recognition result for the voice command output by the speech recognition test model are compared to obtain model comparison data.
[0146] The second speech recognition result is detected using the vehicle's operating parameters to obtain operation detection data. The model comparison data and the operation detection data are then provided to the backend. If the backend determines that both the model comparison data and the operation detection data meet the release conditions, the speech recognition test model is released as an in-vehicle speech recognition model.
[0147] According to the technical solution of this invention, an in-vehicle voice recognition model and a voice recognition test model deployed in mass-produced vehicles respectively receive voice commands sent by users inside the vehicle; the in-vehicle voice recognition model is compared with the voice recognition test model; and the logicality of the second voice recognition result is detected using vehicle operating parameters. The decision to release the voice recognition test model is made based on model comparison data and operational detection data. Before releasing the voice recognition model, comparing the already released in-vehicle voice recognition model deployed in mass-produced vehicles with the vehicle's operating parameters improves the reliability and effectiveness of the voice recognition test model in practical applications.
[0148] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention. It should be noted that the acquisition, storage, and application of user personal information involved in the technical solutions of this disclosure comply with relevant laws and regulations and do not violate public order and good morals.
Claims
1. A method for publishing an in-vehicle voice recognition model, characterized in that, include: The in-vehicle voice recognition model and voice recognition test model deployed in mass-produced vehicles receive voice commands sent by users inside the vehicle. The first speech recognition result for the voice command output by the vehicle speech recognition model and the second speech recognition result for the voice command output by the speech recognition test model are compared to obtain model comparison data. The second speech recognition result is detected using the vehicle's operating parameters to obtain operation detection data. The model comparison data and the operation detection data are then provided to the backend. If the backend determines that both the model comparison data and the operation detection data meet the release conditions, the speech recognition test model is released as an in-vehicle speech recognition model.
2. The method for publishing an in-vehicle voice recognition model according to claim 1, characterized in that, The step of comparing the first speech recognition result output by the vehicle-mounted speech recognition model for the speech command and the second speech recognition result output by the speech recognition test model for the speech command includes: The speech recognition text in the first speech recognition result is compared with the test recognition text in the second speech recognition result to obtain text comparison data; The vehicle operation instructions in the speech recognition text are compared with the test operation instructions in the test recognition text to obtain operation comparison data.
3. The method for publishing an in-vehicle speech recognition model according to claim 2, characterized in that, The vehicle operation instructions include: vehicle operation actions, vehicle components targeted by the vehicle operation actions, and vehicle operation parameters for the vehicle components. The test operation instructions include: operation test actions, vehicle components targeted by the operation test actions, and operation test parameters for the vehicle components. The step of comparing the vehicle operation instructions in the speech recognition text with the test operation instructions in the test recognition text includes: The vehicle operation actions are compared with the operation test actions, the vehicle components targeted by the vehicle operation actions are compared with the vehicle components targeted by the operation test actions, and the vehicle operation parameters for the vehicle components are compared with the operation test parameters for the vehicle components.
4. The method for publishing an in-vehicle speech recognition model according to claim 1, characterized in that, The step of obtaining operational detection data by detecting the second speech recognition result using the vehicle's operational parameters includes: Based on the driving parameters and / or equipment parameters in the vehicle's operating parameters, the logical rationality of the operation test action in the second voice recognition result is detected. The operation detection data includes the detection result for logical rationality, the driving parameters and / or the equipment parameters.
5. The method for publishing an in-vehicle speech recognition model according to claim 2, characterized in that, Also includes: If the file comparison data indicates that the speech recognition text is similar to the test recognition text, and the operation comparison data indicates that the vehicle operation command and the test operation command are the same, then the step of obtaining operation detection data by detecting the second speech recognition result using the vehicle's operating parameters is executed.
6. The method for publishing an in-vehicle speech recognition model according to claim 1, characterized in that, Also includes: If the model comparison data indicates that the first speech recognition result is inconsistent with the second speech recognition result, and / or the operation detection data indicates that the logic of the second speech recognition result is unreasonable, then the in-vehicle voice data within a preset time period before and after the time point corresponding to the voice command and the time point corresponding to the voice command are acquired and provided to the backend.
7. The method for publishing an in-vehicle speech recognition model according to claim 1, characterized in that, Also includes: If the model comparison data indicates that the first speech recognition result is consistent with the second speech recognition result, and the operation detection data indicates that the second speech recognition result is logically reasonable, then the in-vehicle voice data at the time point corresponding to the voice command is acquired and provided to the backend.
8. The method for publishing an in-vehicle speech recognition model according to claim 1, characterized in that, Also includes: If the background determines that either the model comparison data or the operation detection data does not meet the release conditions, the speech recognition test model is updated using the model comparison data and the operation detection data, and the updated speech recognition test model is deployed in the mass-produced vehicle. And / or, The speech recognition test model is a new model independent of the in-vehicle speech recognition model, or the speech recognition test model is an upgraded version of the in-vehicle speech recognition model.
9. A device for publishing an in-vehicle voice recognition model, characterized in that, include: The voice module is used to deploy in the released in-vehicle voice recognition model and voice recognition test model in mass-produced vehicles, respectively, to receive voice commands sent by users in the vehicle; The comparison module is used to compare the first speech recognition result of the vehicle speech recognition model for the speech command with the second speech recognition result of the speech recognition test model for the speech command to obtain model comparison data. The detection module is used to detect the second speech recognition result using the vehicle's operating parameters to obtain operating detection data, and to provide the model comparison data and the operating detection data to the backend. If the backend determines that both the model comparison data and the operating detection data meet the release conditions, it releases the speech recognition test model as an in-vehicle speech recognition model.
10. A vehicle, characterized in that, Includes the apparatus for publishing an in-vehicle voice recognition model as described in claim 9.