Voice sample acquisition method and device, vehicle and computer program product
By capturing the wrongly identified data samples corrected by users, generating positive and negative samples and saving them, the difficulty of collecting and annotating entity recognition samples in vehicle voice navigation is solved, efficient and low-cost data acquisition and model optimization are achieved, and the performance and user experience of the voice recognition model are improved.
Patent Information
- Application Number
- CN202510423452.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-06-27
AI Technical Summary
In the field of vehicle-mounted voice navigation technology, it is difficult to obtain sufficient labeling data, which affects the training effect and accuracy of speech recognition models. Especially when dealing with place names and entity recognition of diversity and complexity, it is difficult for the existing technology to effectively collect and label entity recognition samples, and it cannot meet the entity update requirements.
By obtaining the navigation voice commands and correction operation instructions input by the user, the first entity text information and the first annotation information, as well as the second entity text information and the second annotation information modified by the user, the similarity between the two is calculated. When the similarity is greater than the preset value, a pair of positive and negative samples are generated and saved for training or optimizing the speech recognition model.
This method can efficiently and at low cost to obtain a large number of real and valuable speech recognition data samples, effectively collect and label entity recognition samples, meet entity update needs, improve the recognition accuracy and generalization capabilities of the speech recognition model, and improve the interactive experience between users and the speech recognition model.
Smart Images

Figure CN120220658A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving technology, and in particular to a voice sample collection method and device, a vehicle, and a computer program product. Background Art
[0002] Speech recognition technology is constantly developing and improving, and its applications can be divided into entity recognition such as place names and speech instruction recognition. Entity recognition focuses on accurately identifying specific entities in text, such as place names; while speech instruction recognition focuses on understanding user intentions and instructions, such as voice control, search, etc. Speech entity recognition is widely used in the field of in-vehicle voice navigation.
[0003] Due to the diversity and complexity of specific entities (such as place names), it is very difficult to obtain sufficient annotated data, which affects the training effect and accuracy of speech recognition models. The proprietary pronunciation of place names and the diversity and professionalism of POI data make the industry usually adopt data purchase to augment data, which requires a lot of manpower and time. Especially for some uncommon or specific regions, the scarcity of annotated data is more prominent. In addition, the performance of speech recognition models depends to a large extent on their ability to handle various accents, speaking speeds, background noise and other complex situations. The finished annotated data is usually single in nature and poor in diversity. Moreover, with the changes of the times, new entities such as place names continue to emerge, while old entities may gradually fade out of people's field of vision. Batch correction data usually has a lag.
[0004] In summary, how to effectively collect and label entity recognition samples and meet entity update requirements is a technical problem that needs to be urgently solved in the field of in-vehicle voice navigation technology. Summary of the invention
[0005] The purpose of the present invention is to propose a voice sample collection method and device, a vehicle, and a computer program product to achieve effective entity recognition sample collection and labeling and meet entity update requirements.
[0006] To achieve the above object, according to a first aspect of the present invention, a method for collecting speech samples is provided, comprising:
[0007] Acquire a navigation voice command input by a user, recognize the navigation voice command, obtain first entity text information and first annotation information, and output and display the first entity text information and the first annotation information;
[0008] Acquire a correction operation instruction input by a user, and obtain second entity text information and second annotation information according to the correction operation instruction;
[0009] Obtain the similarity between the first annotation information and the second annotation information. When the similarity is greater than a preset value, obtain and save a pair of positive and negative samples for training or optimizing the speech recognition model; wherein, the positive sample includes the second entity text information and the second annotation information, and the negative sample includes the first entity text information and the first annotation information.
[0010] According to a second aspect of the present invention, there is provided a speech sample acquisition device, including:
[0011] A voice command recognition module, configured to obtain a navigation voice command input by a user, recognize the navigation voice command, obtain first entity text information and first annotation information, and output and display the first entity text information and the first annotation information;
[0012] A correction information acquisition module, configured to obtain a correction operation command input by a user, and obtain second entity text information and second annotation information according to the correction operation command;
[0013] An identification sample acquisition module, configured to obtain the similarity between the first annotation information and the second annotation information. When the similarity is greater than a preset value, obtain and save a pair of positive and negative samples for training or optimizing the speech recognition model; wherein, the positive sample includes the second entity text information and the second annotation information, and the negative sample includes the first entity text information and the first annotation information.
[0014] According to a third aspect of the present invention, there is provided a vehicle, including:
[0015] A communication interface, configured to communicate with other electronic devices;
[0016] A memory, configured to store computer program instructions;
[0017] A processor, configured to execute the computer program instructions to support the device to implement the method as described in the first aspect.
[0018] According to a fourth aspect of the present invention, there is provided a computer program product, including computer program instructions, and the computer program instructions direct a computer device to perform operations corresponding to the method as described in the first aspect.
[0019] The above-mentioned speech sample acquisition method and device, vehicle, and computer program product have the following beneficial effects:
[0020] The above method is implemented during the user's use of in-vehicle voice navigation. Compared with the traditional methods of manually annotating or simulating and generating data samples, the above method captures the misrecognized data samples corrected by the user, generates a pair of positive and negative samples and saves them, and can obtain a large number of real and valuable speech recognition data samples more efficiently and at low cost, effectively collecting and annotating entity recognition samples, meeting the entity update requirements, and enhancing the interaction experience between the user and the speech recognition model; the captured positive and negative samples can be used to train or optimize the speech recognition model, support the continuous learning and iteration of the speech recognition model, enable it to continuously adapt to new environments and user needs, and can significantly improve the recognition accuracy and generalization ability of the speech recognition model. In addition, the above method can be designed as a device and a computer program product, which is easy to implement and integrate in the existing in-vehicle voice navigation system. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required in the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0022] Figure 1 It is a flowchart of a speech sample collection method in an embodiment of the present invention.
[0023] Figure 2 It is a framework structure diagram of a speech sample collection device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] The detailed description of the drawings is intended to be illustrative of the current embodiments of the present invention, rather than representing the only form in which the present invention can be implemented. It should be understood that the same or equivalent functions can be accomplished by different embodiments intended to be included within the spirit and scope of the present invention.
[0025] Refer to Figure 1 , an embodiment of the present invention provides a speech sample collection method, and the method includes the following steps:
[0026] Step S10, obtain the navigation voice command input by the user, recognize the navigation voice command, obtain the first entity text information and the first annotation information, and output and display the first entity text information and the first annotation information;
[0027] Specifically, when the user uses in-vehicle voice navigation, the user inputs navigation voice commands through the microphone of the in-vehicle voice system. For example, when requesting navigation to a certain location, the speech recognition model performs speech recognition, converts the user's navigation voice commands into first entity text information, and at the same time, generates first annotation information, where the first annotation information is the speech feature of the first entity text information.
[0028] Step S20: Obtain the correction operation command input by the user, and obtain second entity text information and second annotation information according to the correction operation command;
[0029] Specifically, the user checks whether the first entity text information and the first annotation information displayed on the in-vehicle display screen are correct. When they are incorrect, the user can make corrections. The user can input correction operation commands in various ways, such as through voice commands, touch screen input, physical buttons, etc. The correction operation commands may include the correct voice commands re-said by the user, selecting the correct options, editing text information, etc.; by recording the correction operation commands input by the user, the second entity text information and the second annotation information are obtained. The second voice information is an accurate expression of the user's intention and represents the navigation command that the user hopes the system will understand and execute. The second annotation information is the speech feature of the first entity text information.
[0030] Step S30: Obtain the similarity between the first annotation information and the second annotation information. When the similarity is greater than a preset value, obtain and save a pair of positive and negative samples for training or optimizing the speech recognition model; where the positive sample includes the second entity text information and the second annotation information, and the negative sample includes the first entity text information and the first annotation information.
[0031] Specifically, use a preset similarity measurement method to compare the similarity between the first annotation information and the second annotation information. The similarity measurement can be a string-based similarity algorithm (such as edit distance, cosine similarity, etc.), or a more complex machine learning model to evaluate the similarity degree of the two annotation information, with the aim of determining the difference degree between the first annotation information initially recognized by the speech recognition model and the second annotation information after the user's correction.
[0032] The method of this embodiment presets a similarity threshold, which is used to determine when a pair of annotation information is considered to be different enough to be a valuable training sample. If the similarity is greater than the similarity threshold, it means that the first annotation information and the second annotation information can form a pair of valid positive and negative samples. If the similarity is less than or equal to the similarity threshold, it means that the first annotation information and the second annotation information do not form a pair of valid positive and negative samples. The positive samples include the second entity text information and the second annotation information after user correction, which are the correct entity information that the speech recognition model should learn to recognize. The negative samples include the first entity text information and the first annotation information initially misrecognized by the speech recognition model, which are the errors that the speech recognition model needs to learn to avoid. This pair of positive and negative samples is then saved to the database or training dataset for subsequent speech recognition model training or optimization.
[0033] Through the method of this embodiment, the acquisition of in-vehicle voice navigation entity recognition samples is realized. The speech recognition model can effectively learn from the user's feedback and improve the performance of the speech recognition model through continuous training. This method ensures the quality and relevance of the training data, and helps the speech recognition model better adapt to the changes and diversities in the actual usage scenarios.
[0034] In some embodiments, the first annotation information is the phoneme of the first entity text information, and the second annotation information is the phoneme of the second entity text information.
[0035] Specifically, the first speech recognition information is "Jinyan Jiayuan", and the first annotation information is jin / yan / jia / yuan; the second speech recognition information is "Jinyan Jiayuan", and the second annotation information is jin / yan / jia / yuan.
[0036] In some embodiments, the first speech text is the navigation destination name recognized by the speech recognition model for the navigation voice command; the second speech text is the navigation destination name expected by the user.
[0037] Specifically, the method of this embodiment focuses on improving the accuracy of destination name recognition in the voice navigation system. When the user inputs a navigation voice command, the speech recognition model will try to parse the user's speech and convert it into text. The result of this process is the first speech text, which represents the destination name that the speech recognition model thinks the user wants to navigate to. For example, when the user says "Navigate to Jinyan Jiayuan", the speech recognition model may recognize it as "Jinyan Jiayuan". If the destination name recognized by the speech recognition model is inaccurate or not what the user expects, the user will make a correction. The user provides the correct destination name by inputting a correction operation command, which becomes the second speech text. Continuing the above example, if the user actually wants to navigate to "Jinyan Jiayuan", they will correct it to this name and its annotation information.
[0038] In some embodiments, the correction operation instruction includes that the user manually deletes the displayed first entity text information and manually inputs the second entity text information and the second annotation information on the display interface, and instructs to re-search for the navigation destination.
[0039] Specifically, in step S10, the first entity text information and the first annotation information are displayed in the address bar of the navigation system interface. When the user finds that the first entity text information and the first annotation information are incorrect, the user can manually delete the displayed incorrect information through a touch screen or other input means. After deleting the incorrect information, the user manually inputs the correct destination name, that is, the second entity text information. At the same time, the user also needs to provide the second annotation information. After completing the above operations, the user instructs the navigation system to re-search for the navigation destination according to the new information, which can be completed by clicking a button such as "Search", "OK" or the like.
[0040] Corresponding to the above embodiment, refer to Figure 2 , another embodiment of the present invention provides a voice sample collection device, including:
[0041] A voice command recognition module 1, configured to obtain a navigation voice command input by a user, recognize the navigation voice command, obtain first entity text information and first annotation information, and output and display the first entity text information and the first annotation information;
[0042] A correction information acquisition module 2, configured to obtain a correction operation instruction input by a user, and obtain second entity text information and second annotation information according to the correction operation instruction;
[0043] An identification sample collection module 3, configured to obtain the similarity between the first annotation information and the second annotation information. When the similarity is greater than a preset value, obtain and save a pair of positive and negative samples for training or optimizing a voice recognition model; wherein, the positive sample includes the second entity text information and the second annotation information, and the negative sample includes the first entity text information and the first annotation information.
[0044] In some embodiments, the first annotation information is the phoneme of the first entity text information, and the second annotation information is the phoneme of the second entity text information.
[0045] In some embodiments, the first voice text is the navigation destination name recognized by the voice recognition model for the navigation voice command; the second voice text is the navigation destination name expected by the user.
[0046] In some embodiments, the correction operation instruction includes that the user manually deletes the displayed first entity text information on the display interface, manually inputs the second entity text information and the second annotation information, and instructs to re-search for a navigation destination.
[0047] It should be noted that the device provided in this embodiment can be used to execute the method described in the above embodiment.
[0048] Therefore, the content not detailed in this embodiment can be obtained by referring to the content of the method in the above embodiment, so it will not be elaborated here.
[0049] Another embodiment of the present invention provides a vehicle, including:
[0050] A communication interface for communicating with other electronic devices;
[0051] A memory for storing computer program instructions;
[0052] A processor for executing the computer program instructions to support the vehicle to implement the method described in the above embodiment.
[0053] In this embodiment, the memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store operating devices, application programs required for at least one function, etc., and the data storage area can store relevant data, etc. In addition, the memory can be a high-speed random access memory, or a non-volatile memory, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a FlashCard, etc., or the memory can also be other volatile solid-state storage devices.
[0054] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or the processor can also be any conventional processor. The processor is the control center of the vehicle, and uses various interfaces and lines to connect various parts of the vehicle.
[0055] Another embodiment of the present invention provides a computer program product, including computer program instructions that direct a computer device to perform operations corresponding to the method described in the above embodiment.
[0056] Specifically, the computer program product includes a series of computer program instructions, which are codes written in the computer program. They define how to perform specific operations. These computer program instructions are designed to be loaded onto a computer device and guide the device to perform specific operations, which refer to the respective steps in the voice sample collection method described in the above embodiment. In this way, the computer program product of this embodiment provides a complete software solution, which can run on various computer devices and implement the voice sample collection method of the above embodiment.
[0057] The embodiments of the present invention have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to technologies in the market, or to enable other ordinary skill in the art in the technical field to understand the disclosed embodiments.
Claims
1. A method for collecting speech samples, characterized in that: include: Acquire a navigation voice command input by a user, recognize the navigation voice command, obtain first entity text information and first annotation information, and output and display the first entity text information and the first annotation information; Acquire a correction operation instruction input by a user, and obtain second entity text information and second annotation information according to the correction operation instruction; Obtain the similarity between the first annotation information and the second annotation information. When the similarity is greater than a preset value, obtain and save a pair of positive and negative samples for training or optimizing a speech recognition model; wherein the positive sample includes the second entity text information and the second annotation information, and the negative sample includes the first entity text information and the first annotation information.
2. The method according to claim 1, characterized in that The first annotation information is the phonemes of the first entity text information, and the second annotation information is the phonemes of the second entity text information.
3. The method according to claim 2, characterized in that The first voice text is the name of the navigation destination obtained by the voice recognition model recognizing the navigation voice instruction; the second voice text is the name of the navigation destination desired by the user.
4. The method according to claim 3, characterized in that The correction operation instruction includes the user manually deleting the first entity text information displayed on the display interface and manually inputting the second entity text information and the second annotation information, and instructing to search the navigation destination again.
5. A speech sample collection device, characterized in that: include: A voice command recognition module, used to obtain a navigation voice command input by a user, recognize the navigation voice command, obtain first entity text information and first annotation information, and output and display the first entity text information and the first annotation information; A correction information acquisition module, used to acquire a correction operation instruction input by a user, and obtain the second entity text information and the second annotation information according to the correction operation instruction; An identification sample acquisition module is used to obtain the similarity between the first annotation information and the second annotation information. When the similarity is greater than a preset value, a pair of positive and negative samples for training or optimizing a speech recognition model is obtained and saved; wherein the positive sample includes the second entity text information and the second annotation information, and the negative sample includes the first entity text information and the first annotation information.
6. The device according to claim 5, characterized in that The first annotation information is the phonemes of the first entity text information, and the second annotation information is the phonemes of the second entity text information.
7. The device according to claim 6, characterized in that The first voice text is the name of the navigation destination obtained by the voice recognition model recognizing the navigation voice instruction; the second voice text is the name of the navigation destination desired by the user.
8. The device according to claim 7, characterized in that The correction operation instruction includes the user manually deleting the first entity text information displayed on the display interface and manually inputting the second entity text information and the second annotation information, and instructing to search the navigation destination again.
9. A vehicle, characterized in that: include: A communication interface, used to communicate with other electronic devices; a memory for storing computer program instructions; A processor, configured to execute the computer program instructions to enable the apparatus to implement the method according to any one of claims 1 to 4.
10. A computer program product, characterized in that The method comprises computer program instructions, wherein the computer program instructions instruct a computer device to execute operations corresponding to the method according to any one of claims 1 to 4.
Citation Information
Cited By
Training method and training device for labeling model
CN120748381A