Methods, devices, electronic equipment, and storage media for generating training data

By acquiring popular locations in the target city, constructing multiple incorrect names, and generating two rounds of navigation correction dialogue, the problem of navigation path errors in the in-vehicle intelligent voice interaction system was solved, improving the error correction capability of the navigation model and the user experience.

CN122090823APending Publication Date: 2026-05-26VOYAH AUTOMOBILE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
VOYAH AUTOMOBILE TECH CO LTD
Filing Date
2026-01-07
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In in-vehicle intelligent voice interaction systems, the lack of high-quality training data for multi-round error correction scenarios leads to frequent navigation path errors, requiring users to interact with the vehicle system through multi-round dialogue for error correction.

Method used

By acquiring popular locations in the target city, constructing multiple incorrect names, and generating multiple sets of two-round navigation correction dialogues to be applied to the training data of the navigation model, the error patterns in real speech recognition are simulated, and structured two-round error correction dialogue data is automatically generated.

Benefits of technology

It improves the navigation model's ability to understand users' error correction intentions and the accuracy of correcting incorrect place names, thereby enhancing navigation interaction and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122090823A_ABST
    Figure CN122090823A_ABST
Patent Text Reader

Abstract

This application provides a method, apparatus, electronic device, and storage medium for generating training data. The method involves acquiring a preset number of popular locations related to a target city; for each popular location, constructing multiple incorrect names based on the target Chinese characters within the location and a preset list of pronunciation errors; and generating multiple sets of two-round navigation correction dialogues based on the multiple incorrect names and the popular locations themselves, which are then applied to the training data of a navigation model. Applying this technical solution can avoid the technical problem of inaccurate navigation applications in vehicle systems due to missing relevant training data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle technology, and more particularly to a method, apparatus, electronic device, and storage medium for generating training data. Background Technology

[0002] In in-vehicle intelligent voice interaction systems, users often use voice commands to perform operations such as setting navigation destinations and planning routes with the vehicle's infotainment system.

[0003] Currently, due to factors such as environmental noise, differences in user pronunciation, and limitations in the performance of voice signal acquisition equipment, the vehicle's infotainment system may produce incorrect destination name recognition results after receiving user voice commands. This error will directly lead to incorrect navigation routes, requiring users to interact with the vehicle's infotainment system through multiple rounds of dialogue to correct the error.

[0004] However, the existing training data lacks high-quality data for such multi-round error correction scenarios, resulting in inaccurate navigation applications in vehicle systems. Summary of the Invention

[0005] This application provides a method, apparatus, electronic device, and storage medium for generating training data, in order to avoid inaccurate navigation applications in vehicle systems due to the lack of relevant training data.

[0006] In a first aspect, embodiments of this application provide a method for generating training data, including:

[0007] Retrieve a preset number of popular locations related to the target city;

[0008] For each popular location, multiple incorrect names are constructed based on the target Chinese characters in the popular location and a preset list of pronunciation errors;

[0009] Based on the multiple incorrect names of the popular locations and the popular locations themselves, multiple sets of two-round navigation correction dialogues are generated to be applied to the training data of the navigation model.

[0010] In one or more embodiments, constructing multiple incorrect names for the popular locations based on the target Chinese characters in the popular locations and a preset list of pronunciation errors includes:

[0011] Based on the target Chinese characters in the popular locations, the intelligent agent generates a list of homophones and near-homophones of the target Chinese characters.

[0012] If the first letter of a homophone in the homophone list belongs to a letter in the pronunciation error list, the target Chinese character and the homophone list are sent to the Chinese character checking agent.

[0013] After the Chinese character checking agent verifies the correctness of the target Chinese character and the list of homophones, the dialogue data construction agent generates multiple incorrect names for the popular location based on the target Chinese character and the list of homophones.

[0014] In one or more embodiments, the step of constructing an intelligent agent based on dialogue data to generate multiple incorrect names for the popular locations according to the target Chinese character and the list of homophones and near-homophones includes:

[0015] The intelligent agent constructed using the dialogue data replaces the target Chinese characters with characters from a list of homophones and near-homophones, generating multiple incorrect names for the popular locations.

[0016] In one or more embodiments, before constructing an agent using dialogue data to generate multiple incorrect names for the popular locations based on the target Chinese character and the list of homophones and near-homophones, the method further includes:

[0017] The intelligent agent for checking Chinese characters performs correctness checks on homophones / near-homophones and / or checks on the correctness of preset pronunciation errors based on the target Chinese character and the list of homophones / near-homophones.

[0018] If the correctness check passes, the target Chinese character and the list of homophones and near-homophones will be sent to the dialogue data construction agent;

[0019] If the correctness check fails, instruct the Chinese character extension generation agent to regenerate the list of homophones and near-homophones of the target Chinese character.

[0020] In one or more embodiments, generating multiple sets of two-round navigation correction dialogues for use as training data for a navigation model, based on multiple incorrect names of the popular locations and the popular locations themselves, includes:

[0021] For each incorrect name, an intelligent agent is constructed using dialogue data to build a set of two-round navigation correction dialogues based on the incorrect name and the target Chinese characters in the popular locations;

[0022] The dialogue verification agent performs an integrity check on the two-round navigation correction dialogue. If the check is successful, the two-round navigation correction dialogue is added to the training data.

[0023] In one or more embodiments, the method further includes:

[0024] If the two rounds of navigation correction dialogue are not satisfactory, the two rounds of navigation correction dialogue are fed back to the dialogue data construction agent for reconstruction until the integrity check is passed or the number of feedbacks reaches a preset number. Then, the two most recent rounds of navigation correction dialogue are added to the training data.

[0025] In one or more embodiments, obtaining a preset number of popular locations related to the target city includes:

[0026] S1, obtain the preset number of multiple popular locations related to the target city through a location-generating agent;

[0027] S2, the location detection agent performs deduplication and fake location removal based on the multiple popular locations and the preset location database;

[0028] S3, if the first number of processed popular locations is less than the preset number, repeat steps S1-S3 until the first number reaches the preset number.

[0029] Secondly, embodiments of this application provide a training data generation apparatus, comprising:

[0030] The acquisition module is used to acquire a preset number of popular locations related to the target city;

[0031] A construction module is used to construct multiple incorrect names for each popular location based on the target Chinese characters in the popular location and a preset list of pronunciation errors;

[0032] The generation module is used to generate multiple sets of two-round navigation correction dialogues based on the multiple incorrect names of the popular locations and the popular locations themselves, for use as training data for the navigation model.

[0033] In one or more embodiments, the construction module constructs multiple incorrect names for the popular locations based on the target Chinese characters in the popular locations and a preset list of pronunciation errors, specifically for:

[0034] Based on the target Chinese characters in the popular locations, the intelligent agent generates a list of homophones and near-homophones of the target Chinese characters.

[0035] If the first letter of a homophone in the homophone list belongs to a letter in the pronunciation error list, the target Chinese character and the homophone list are sent to the Chinese character checking agent.

[0036] After the Chinese character checking agent verifies the correctness of the target Chinese character and the list of homophones, the dialogue data construction agent generates multiple incorrect names for the popular location based on the target Chinese character and the list of homophones.

[0037] In one or more embodiments, the construction module constructs an agent based on dialogue data to generate multiple incorrect names for the popular locations according to the target Chinese characters and the list of homophones and near-homophones, specifically for:

[0038] The intelligent agent constructed using the dialogue data replaces the target Chinese characters with characters from a list of homophones and near-homophones, generating multiple incorrect names for the popular locations.

[0039] In one or more embodiments, before the construction module generates multiple incorrect names for the popular locations based on the target Chinese characters and the list of homophones and near-homophones using dialogue data, the construction module is further configured to:

[0040] The intelligent agent for checking Chinese characters performs correctness checks on homophones / near-homophones and / or checks on the correctness of preset pronunciation errors based on the target Chinese character and the list of homophones / near-homophones.

[0041] If the correctness check passes, the target Chinese character and the list of homophones and near-homophones will be sent to the dialogue data construction agent;

[0042] If the correctness check fails, instruct the Chinese character extension generation agent to regenerate the list of homophones and near-homophones of the target Chinese character.

[0043] In one or more embodiments, the generation module generates multiple sets of two-round navigation correction dialogues based on multiple incorrect names of the popular locations and the popular locations themselves, for use in training data for the navigation model, specifically for:

[0044] For each incorrect name, an intelligent agent is constructed using dialogue data to build a set of two-round navigation correction dialogues based on the incorrect name and the target Chinese characters in the popular locations;

[0045] The dialogue verification agent performs an integrity check on the two-round navigation correction dialogue. If the check is successful, the two-round navigation correction dialogue is added to the training data.

[0046] In one or more embodiments, the generation module is further configured to:

[0047] If the two rounds of navigation correction dialogue are not satisfactory, the two rounds of navigation correction dialogue are fed back to the dialogue data construction agent for reconstruction until the integrity check is passed or the number of feedbacks reaches a preset number. Then, the two most recent rounds of navigation correction dialogue are added to the training data.

[0048] In one or more embodiments, the acquisition module acquires a preset number of popular locations related to the target city, specifically for:

[0049] S1, obtain the preset number of multiple popular locations related to the target city through a location-generating agent;

[0050] S2, the location detection agent performs deduplication and fake location removal based on the multiple popular locations and the preset location database;

[0051] S3, if the first number of processed popular locations is less than the preset number, repeat steps S1-S3 until the first number reaches the preset number.

[0052] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor;

[0053] The memory stores computer-executed instructions;

[0054] The processor executes computer execution instructions stored in the memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.

[0055] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and / or various possible implementations of the first aspect.

[0056] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the first aspect and / or various possible implementations of the first aspect.

[0057] The training data generation method, apparatus, electronic device, and storage medium provided in this application embodiment acquire a preset number of popular locations related to the target city; for each popular location, construct multiple incorrect names for the popular location based on the target Chinese characters in the popular location and a preset list of pronunciation errors; and generate multiple sets of two-round navigation correction dialogues based on the multiple incorrect names and popular locations for use as training data for the navigation model. In this technical solution, by acquiring a preset number of real popular locations related to the target city, a core vocabulary covering typical navigation destinations is constructed. For each location, based on its Chinese character composition, multiple incorrect names are generated using homophones, near-homophones, and specific pronunciation confusion rules, simulating typical error patterns caused by noise, accents, and other factors in real speech recognition. By pairing these incorrect names with the correct locations, structured two-round error correction dialogue data is automatically generated. This process can generate high-quality, diverse multi-round error correction training samples on a large scale, directly addressing the need for speech recognition error correction in navigation scenarios. It provides rich error correction interaction paradigms for training the navigation model, helping to improve the model's understanding of user error correction intentions and the accuracy of correcting incorrect place names, ultimately enhancing navigation interaction and user experience. Attached Figure Description

[0058] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0059] Figure 1 A flowchart illustrating the training data generation method provided in the embodiments of this application. Figure 1 ;

[0060] Figure 2 A flowchart illustrating the training data generation method provided in the embodiments of this application. Figure 2 ;

[0061] Figure 3 A flowchart illustrating the training data generation method provided in the embodiments of this application. Figure 3 ;

[0062] Figure 4 A schematic diagram of agent interaction for the training data generation method provided in the embodiments of this application;

[0063] Figure 5 A schematic diagram of the structure of the training data generation device provided in the embodiments of this application;

[0064] Figure 6 A schematic diagram of the structure of the electronic device provided in this application.

[0065] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0066] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0067] In in-vehicle intelligent voice interaction systems, users often use voice commands to perform operations such as setting navigation destinations and planning routes. Due to the influence of environmental noise (e.g., road noise, passenger conversations), differences in user pronunciation (e.g., dialects, accents), and the performance limitations of voice signal acquisition equipment, the in-vehicle system may produce typos in the destination name after receiving the user's voice command.

[0068] The aforementioned errors will directly lead to incorrect navigation routes, requiring users to engage in multiple rounds of dialogue with the vehicle's infotainment system to correct the errors.

[0069] Existing technologies primarily generate multi-turn dialogue data through manual annotation or rule engines, but these have the following limitations:

[0070] 1) Lack of error injection mechanism: The data is only processed based on the original dialogue data, without systematic error injection for potential typos in the first round of dialogue (such as words with similar pronunciations or similar shapes), resulting in insufficient matching between the generated error correction data and real user error correction scenarios;

[0071] 2) Uniform error correction content generation: It usually relies on predefined error correction templates (such as "Did you mean XX?") and does not combine the semantic features of user error correction behavior (such as the contextual association of idioms / words) to generate dynamic error correction content, which limits the naturalness and accuracy of error correction interaction;

[0072] 3) Insufficient data quality and diversity: Error correction samples are not dynamically generated based on real geographic data (e.g., a database of popular urban locations), and there is a lack of closed-loop verification of the error correction logic (e.g., whether the first round of errors triggers effective second-round error correction), resulting in a large number of invalid or redundant samples in the training data.

[0073] 4) Poor model adaptability: The data generation strategy was not optimized for the characteristics of the small parameter model on the vehicle terminal (e.g., low resource and low latency requirements), resulting in the generated data being out of touch with the model training objectives and affecting the model's performance in actual deployment.

[0074] Accordingly, in response to the technical problems existing in the prior art, the inventors of this application have the following concept: Observing that speech recognition errors, most misrecognition results are caused by homophones, near-homophones, and confusion in the pronunciation of specific dialects, the technical problem can be transformed into a computable generation task. By using a preset list of pronunciation errors, possible error patterns of popular locations can be constructed, and by using the correct location-error name pairing relationship, a standardized two-round error correction dialogue template can be reverse-engineered, thereby generating training data that can improve the accuracy of the navigation model.

[0075] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0076] It should be understood that this method is applied to electronic devices, specifically cloud servers or in-vehicle systems.

[0077] Figure 1 A flowchart illustrating the training data generation method provided in the embodiments of this application. Figure 1 ,like Figure 1 As shown, the method includes:

[0078] Step 11: Obtain a preset number of popular locations related to the target city;

[0079] In this step, a preset number of popular locations related to the target city are obtained. Each popular location may include: well-known landmarks, business districts, transportation hubs, etc.

[0080] In one possible implementation, the preset quantity could be 100.

[0081] Optionally, one possible implementation of step 11 could be:

[0082] S1, obtains a preset number of popular locations related to the target city through a location-generating agent;

[0083] In this implementation, the location generation agent can be a program or model with specific task instructions. After receiving the name of the target city, it automatically crawls, organizes and outputs multiple popular locations according to predefined strategies (such as obtaining popular tourist attractions, major business districts, and important transportation hubs) by integrating or calling tools such as search engines, map APIs, and tourism website data interfaces.

[0084] S2, through the location detection agent, performs deduplication and fake location removal based on multiple popular locations and a preset location database;

[0085] In this implementation, during the fake location removal process, a search tool can be used to perform a reverse query on the popular locations generated by S1. By cross-validating the authority and consistency of the information, it can be determined whether the popular location actually exists.

[0086] In the deduplication process, locations confirmed as real are compared with a pre-built or continuously maintained preset location library. If the location already exists in the preset location library, it is marked as a duplicate; otherwise, it is marked as a new location to be added to the library.

[0087] S3. If the first number of popular locations after processing is less than the preset number, repeat steps S1-S3 until the first number reaches the preset number.

[0088] In this implementation, after executing S1 and S2, the number of locations marked as new locations (i.e., the first number) is counted and compared with the preset number.

[0089] If the initial number is less than the preset number, it means that more new locations need to be added. At this time, the detailed information of the "fake locations" and "duplicate locations" identified in S2 is fed back to the location generation agent as feedback. Based on this feedback (e.g., "a fictional building does not exist, please avoid generating similar names"), the location generation agent adjusts its subsequent generation strategy to generate new locations that are more likely to pass the verification.

[0090] Step 12: For each popular location, construct multiple incorrect names for the popular location based on the target Chinese characters in the popular location and the preset list of pronunciation errors;

[0091] In this step, a series of possible incorrect representations are constructed for each popular location to simulate typical misrecognition situations in speech recognition. A single Chinese character can be extracted from the obtained popular locations as the target Chinese character.

[0092] Then, by combining this target Chinese character with a pre-defined list of pronunciation errors (which defines the correspondence of specific easily confused initials, such as nl, hf, etc.), multiple incorrect names for the popular location are constructed.

[0093] Step 13: Based on multiple incorrect names of popular locations and popular locations, generate multiple sets of two-round navigation correction dialogues to be applied to the training data of the navigation model.

[0094] In this step, multiple sets of two-round navigation correction dialogues are generated based on the correct names of popular locations (containing the target Chinese characters) and their corresponding multiple incorrect names, which are then used as training data for the navigation model.

[0095] In each two-round navigation correction dialogue, the process could be as follows: in the first round, the user requests an incorrect location name, and in the second round, the user corrects the location name and confirms the navigation.

[0096] The training data generation method provided in this application involves acquiring a preset number of popular locations related to a target city; for each popular location, constructing multiple incorrect names based on the target Chinese characters in the popular location and a preset list of pronunciation errors; and generating multiple sets of two-round navigation correction dialogues based on the multiple incorrect names and the popular locations themselves, which are then used as training data for the navigation model. In this technical solution, a core vocabulary covering typical navigation destinations is constructed by acquiring a preset number of real popular locations related to the target city. For each location, multiple incorrect names are generated based on its Chinese character composition, using homophones, near-homophones, and specific pronunciation confusion rules. This simulates typical error patterns caused by noise, accents, and other factors in real speech recognition. By pairing these incorrect names with the correct locations, structured two-round error correction dialogue data is automatically generated. This process can generate high-quality, diverse multi-round error correction training samples on a large scale, directly addressing the need to correct speech recognition errors in navigation scenarios. It provides rich error correction interaction paradigms for training the navigation model, helping to improve the model's understanding of user error correction intentions and the accuracy of correcting incorrect place names, ultimately enhancing navigation interaction and user experience.

[0097] Based on the above embodiments, Figure 2 A flowchart illustrating the training data generation method provided in the embodiments of this application. Figure 2 ,like Figure 2 As shown, step 12 may include:

[0098] Step 21: Generate an intelligent agent by expanding Chinese characters. Based on the target Chinese characters in popular locations, generate a list of homophones and near-homophones of the target Chinese characters.

[0099] In this step, the Chinese character extension generation agent receives the target Chinese character as input, and internally integrates and calls the Chinese character pinyin and phonology database.

[0100] Then, according to the preset rules (e.g., the same initial consonant and the same final vowel are defined as homophones; the same initial consonant and similar final vowel or vice versa are defined as near-homophones), all homophones and near-homophones of the Chinese character are queried and output to form a list of homophones and near-homophones.

[0101] Step 22: If the first letter of a near-homophone in the near-homophone list belongs to a letter in the pronunciation error list, send the target Chinese character and the near-homophone list to the Chinese character checking AI.

[0102] In this step, the electronic device has a preset list of pronunciation errors, which is essentially a mapping rule table for easily confused initials (e.g., n->l, h->f, etc.). After generating a list of homophones and near-homophones, the first letter of the pinyin for each homophone and near-homophone in the list is determined.

[0103] If the initial letter of the pinyin of a character can form a pair of confusing relationships with the initial letter of the pinyin of the target Chinese character in the pronunciation error list (for example, the initial letter h of a homophone "huī" of the target character "fēi" triggers the h / f confusion rule), then the entire character list will be marked as involving a specific pronunciation error, and thus enter the next step, that is, sending the target Chinese character and the list of homophonic and near-homophonic characters to the Chinese character checking intelligent agent.

[0104] Step 23, after the Chinese character checking intelligent agent verifies the correctness of the target Chinese character and the list of homophonic and near-homophonic characters, the intelligent agent for constructing dialogue data generates multiple incorrect names of popular locations based on the target Chinese character and the list of homophonic and near-homophonic characters.

[0105] In this step, after the Chinese character checking intelligent agent verifies the correctness of the target Chinese character and the list of homophonic and near-homophonic characters, the intelligent agent for constructing dialogue data replaces the target Chinese character according to the list of homophonic and near-homophonic characters to generate multiple incorrect names of popular locations.

[0106] Optionally, a possible implementation of the intelligent agent for constructing dialogue data generating multiple incorrect names of popular locations based on the target Chinese character and the list of homophonic and near-homophonic characters in step 23 is: the intelligent agent for constructing dialogue data replaces the target Chinese character with the characters in the list of homophonic and near-homophonic characters respectively to generate multiple incorrect names of popular locations.

[0107] In this implementation, the position of the target Chinese character in the string of the popular location is identified, and then each character in the list of homophonic and near-homophonic characters (which can include: the target Chinese character itself, used to generate the "correct" comparison sample, which can be filtered out later) is traversed and filled into this position in turn, so as to generate a series of new strings, and these strings are the multiple incorrect names of the popular location.

[0108] For example, for "Huangzhuang" and the target Chinese character "Huang" and its list of homophonic and near-homophonic characters "Huang, Wang", "Huangzhuang" and "Wangzhuang" will be generated.

[0109] Optionally, the implementation of the correctness verification before step 23 can be:

[0110] Step 1, the Chinese character checking intelligent agent conducts the correctness verification of homophony / near-homophony and / or the correctness verification of preset pronunciation errors according to the target Chinese character, the list of homophonic and near-homophonic characters;

[0111] In this implementation, the target Chinese character and the list of homophonic and near-homophonic characters from step 22 are received. Its verification process includes:

[0112] 1. Homophony / near-homophony correctness verification: Confirm that each character in the list of homophonic and near-homophonic characters is indeed in line with the predefined homophony or near-homophony standard with the target Chinese character, and exclude the characters that are mis-incorporated.

[0113] 2. Preset pronunciation error verification: For the list of homophones and near-homophones submitted for inspection due to step 22, additional verification is performed to confirm whether they actually conform to the preset pronunciation confusion rules.

[0114] Step 2: If the correctness check passes, send the target Chinese character and the list of homophones and near-homophones to the dialogue data construction agent;

[0115] In this implementation, if all checks pass, the list of homophones and near-homophones and the target Chinese character are sent to the dialogue data construction agent.

[0116] Step 3: If the correctness check fails, instruct the Chinese character extension generation agent to regenerate the list of homophones and near-homophones of the target Chinese character.

[0117] In this implementation, if the check fails (e.g., an unreasonable character is found), an error message and correction instruction are sent to the Chinese character extension generation agent, requesting it to regenerate the list of homophones and near-homophones.

[0118] The training data generation method provided in this application embodiment generates a list of homophones and near-homophones of the target Chinese character based on the target Chinese character in the popular location; if the first letter of the homophone in the list of homophones and near-homophones belongs to the letter in the list of pronunciation errors, the target Chinese character and the list of homophones and near-homophones are sent to the Chinese character checking agent; after the Chinese character checking agent checks the correctness of the target Chinese character and the list of homophones and near-homophones, the dialogue data construction agent generates multiple incorrect names of the popular location based on the target Chinese character and the list of homophones and near-homophones. This scheme automatically constructs a list of homophones and near-homophones of the target Chinese character through a Chinese character extension agent. Based on specific pronunciation confusion rules, it selects a subset of the list that conforms to typical speech recognition error patterns. After the Chinese character checking agent verifies its correctness, the dialogue data constructs an agent that systematically generates multiple incorrect names for popular locations using the verified error list. This enables the large-scale production of structured, correct location-typical error name pairing data, thus providing an accurate input foundation for generating high-quality multi-turn navigation error-correction dialogues. This ensures that the training data closely matches the error distribution in actual voice interaction, effectively improving the navigation model's ability to correct typical speech recognition errors and the accuracy of multi-turn dialogue understanding.

[0119] Based on the above embodiments, Figure 3 A flowchart illustrating the training data generation method provided in the embodiments of this application. Figure 3 ,like Figure 3 As shown, step 13 may include:

[0120] Step 31: For each incorrect name, construct an intelligent agent based on the dialogue data, and build a set of two-round navigation correction dialogues according to the incorrect name and the target Chinese characters in popular locations;

[0121] In this step, the dialogue data constructs an agent that receives two key inputs: an incorrect name and the target Chinese character in the corresponding popular location.

[0122] The dialogue data constructs an intelligent agent with several pre-stored dialogue templates, such as "Navigate to [wrong location]" as the first round; "No, it is the [correct Chinese character] of [idiom / word]" as the second round.

[0123] The agent fills the corresponding position in the template with the incorrect name and a pre-associated typical word or idiom (such as "yellow") related to the target Chinese character, thereby generating a set of structured two-round dialogue text.

[0124] For example, in the first round: Navigate to Sanyuan Bridge; in the second round: No, it's the Yuan Dynasty.

[0125] Step 32: Verify the integrity of the two-round navigation correction dialogue through the dialogue verification agent. If it passes, add the two-round navigation correction dialogue to the training data.

[0126] In this step, after receiving the two-round navigation correction dialogue generated in step 31, the dialogue verification agent performs automatic verification according to preset, computable rules, which may include:

[0127] 1) String comparison confirmed that the location names mentioned in the first round did indeed contain incorrect Chinese characters;

[0128] 2) Keyword matching or semantic analysis to confirm that the second round contains specific words that point to the correct Chinese characters;

[0129] 3) Simple logical coherence judgment.

[0130] If all validation rules pass, the group of dialogues is deemed qualified and is directly appended to or used as the final dataset for training the navigation model.

[0131] Step 33: If it fails, feed back the two rounds of navigation correction dialogue to the dialogue data construction agent for reconstruction until the integrity check is passed or the number of feedback reaches the preset number, then add the two most recent rounds of navigation correction dialogue to the training data.

[0132] In this step, if the verification in step 32 fails, the two-round navigation correction dialogue will not be discarded directly.

[0133] At this point, the two-round navigation correction dialogue and the specific rule items that failed the verification (e.g., error code: 102, the second round lacks referential words) are fed back to the dialogue data to construct the agent.

[0134] Then, the dialogue data constructs an agent that, based on the error code or description, attempts to modify its generation strategy (e.g., select another related word for the correct Chinese character and regenerate the second round of dialogue) and produces a new set of two-round navigation correction dialogues.

[0135] Finally, the new two rounds of navigation correction dialogue will be sent to step 32 for verification again until a dialogue that passes the integrity verification is generated, or a maximum number of loops is reached, i.e., a preset number of loops (e.g., 4 times). Reaching the preset number of loops means that automatic correction may fail. At this time, a compromise strategy is adopted, and the most recently generated dialogue version (i.e. the best version that can be obtained at present) is stored in the training data.

[0136] The training data generation method provided in this application involves constructing a set of two-round navigation correction dialogues for each incorrect name using a dialogue data construction agent based on the incorrect name and the target Chinese characters in popular locations. The dialogue verification agent performs integrity checks on the two-round navigation correction dialogues. If the checks are successful, the two-round navigation correction dialogues are added to the training data. If they fail, the two-round navigation correction dialogues are fed back to the dialogue data construction agent for reconstruction. This process continues until the integrity check passes or the number of feedback attempts reaches a preset number, at which point the most recent two-round navigation correction dialogue is added to the training data. This scheme constructs an intelligent agent based on the correspondence between incorrect names and correct Chinese characters using dialogue data. It then builds a structured two-round navigation correction dialogue. The dialogue verification agent verifies the completeness of the dialogue logic and the accuracy of the corrected terms, ensuring that the final training data meets the completeness and rationality requirements of real error correction interactions. This allows for the batch production of high-quality, highly targeted multi-round error correction dialogue samples, effectively solving the problems of high cost and incomplete coverage of manually constructed data. At the same time, iterative correction avoids low-quality or logically chaotic data from polluting the training set, providing accurate and diverse error correction dialogue training materials for the navigation model and significantly improving the model's understanding and correction capabilities in complex multi-round interaction scenarios.

[0137] Based on the above embodiments, Figure 4 A schematic diagram of agent interaction for the training data generation method provided in the embodiments of this application, as shown below. Figure 4 As shown, the agents executing this process may include: a Chinese character expansion generation agent, a Chinese character checking agent, a dialogue data construction agent, and a dialogue verification agent.

[0138] One possible implementation process is as follows:

[0139] Location generation agent: Give the city name to the location generation agent, and the location generation agent will use a search tool to generate 100 popular locations in the current city.

[0140] Location Detection Agent: This agent uses search tools to check whether the locations generated by the location generation agent actually exist. For locations that actually exist, it uses data query tools to check whether they exist in the database. Data that does not exist is stored in the database, and data already in the database is marked as duplicate; data that does not actually exist is marked as non-existent.

[0141] Interaction Process: The location generation agent provides 100 location data records to the location detection agent. The location detection agent feeds back the unqualified locations to the location generation agent, marks the reasons, and orders the location generation agent to supplement the corresponding number of location data records. The maximum number of interaction rounds between the two agents is set to 6 rounds.

[0142] Chinese Character Expansion Generation Agent: Randomly select a Chinese character from the locations generated in step 1, generate a list of homophones and near-homophones of this character, and at the same time determine whether the first letter of the pinyin of this character belongs to [n, f]. If it does, construct a list of mispronounced characters for this character based on the confusion of n / l (e.g., 奶 -> 来) and h / f (e.g., 飞 -> 灰). After construction, send the location, the selected Chinese character, the list of homophones & near-homophones, and the list of mispronounced characters to the Chinese Character Inspection Agent in JSON format.

[0143] Chinese Character Inspection Agent: Receive the content from the Chinese Character Expansion Generation Agent, check whether the list of homophones & near-homophones and the list of mispronounced characters are correct. If correct, send the result to the downstream Dialogue Data Construction Agent; otherwise, point out the specific errors and feedback to the upstream Chinese Character Expansion Generation Agent for correction.

[0144] Dialogue Data Construction Agent: Receive the content from the Chinese Character Inspection Agent and construct two rounds of dialogue error correction navigation data. The construction method is as follows:

[0145] 1) Replace the selected Chinese character in the location name with the $ symbol to obtain loc_1;

[0146] 2) Traverse the list of homophones & near-homophones and the list of mispronounced characters (if any), and replace the $ in loc_1 with them to obtain multiple incorrect locations, s1;

[0147] 3) Traverse s1 and construct two rounds of navigation error correction data. In the first round, use the incorrect locations in s1. In the second round, correct the locations in the first round to the original locations. The first-round expressions are: Navigate to xx, Add xx as a waypoint, Go to xx now, When will I arrive; The second-round expressions are: No / That's wrong, The character in the idiom / phrase. The JSON example is as follows:

[0148] {"Location": "Sanyuanqiao";

[0149] "Chinese Character": "元";

[0150] "First round": Navigate to Sanyuan Bridge;

[0151] "Second round": "No, it's the Yuan Dynasty."

[0152] }

[0153] Dialogue-based agent verification: Receives content from the dialogue data-constructed agent and determines whether its data is acceptable. The criteria for acceptance are as follows:

[0154] 1) Were there any typos in the locations listed in the first round?

[0155] 2) Did the correct idioms / words be used for error correction in the second round?

[0156] 3) Whether the two rounds of dialogue constitute a complete dialogue and whether the logic is correct.

[0157] Qualified data is directly stored in the database, while unqualified data is given specific reasons and fed back to the dialogue data constructing agent.

[0158] Intelligent agent interaction: The Chinese character extension generation intelligent agent, the Chinese character inspection intelligent agent, the dialogue data construction intelligent agent, and the dialogue verification intelligent agent may have multiple rounds of interaction, with the maximum number of interaction rounds set to 4.

[0159] The intelligent agent interaction process of the training data generation method provided in this application has the following technical effects:

[0160] Technical effect 1) Quickly and cost-effectively construct large amounts of high-quality data;

[0161] Technical effect 2) Using the constructed data to train a navigation model with a small number of parameters on the vehicle-mounted system, the vehicle-mounted model can respond to users quickly and accurately, improving the user experience.

[0162] Figure 5 This is a schematic diagram of the structure of the training data generation device provided in the embodiments of this application, as shown below. Figure 5 As shown, the apparatus for generating the training data includes:

[0163] Module 51 is used to obtain a preset number of popular locations related to the target city;

[0164] Module 52 is used to construct multiple incorrect names for each popular location based on the target Chinese characters in the popular location and a preset list of pronunciation errors.

[0165] The generation module 53 is used to generate multiple sets of two-round navigation correction dialogues based on multiple incorrect names of popular locations and popular locations, which are then applied to the training data of the navigation model.

[0166] In one or more embodiments, the construction module 52 constructs multiple incorrect names for popular locations based on the target Chinese characters in the popular locations and a preset list of pronunciation errors, specifically for:

[0167] Based on the target Chinese characters in popular locations, the intelligent agent generates a list of homophones and near-homophones of the target Chinese characters.

[0168] If the first letter of a homophone in the homophone list belongs to a letter in the pronunciation error list, send the target Chinese character and the homophone list to the Chinese character checking agent.

[0169] After the Chinese character checking agent verifies the correctness of the target Chinese character and the list of homophones, the dialogue data-constructed agent generates multiple incorrect names for popular locations based on the target Chinese character and the list of homophones.

[0170] In one or more embodiments, the construction module 52 constructs an intelligent agent based on dialogue data to generate multiple incorrect names for popular locations according to a list of target Chinese characters and homophones, specifically for:

[0171] By constructing an intelligent agent based on dialogue data, the target Chinese character is replaced with a character from a list of homophones or near-homophones, generating multiple incorrect names for popular locations.

[0172] In one or more embodiments, before constructing an agent based on the target Chinese character and a list of homophones and near-homophones using dialogue data to generate multiple incorrect names for popular locations, the construction module 52 is further configured to:

[0173] The intelligent agent checks the correctness of the same / near-sounding characters based on the target Chinese character and a list of homophones and near-homophones, and / or verifies the correctness of preset pronunciation errors.

[0174] If the correctness check passes, the target Chinese character and the list of homophones and near-homophones will be sent to the dialogue data construction agent;

[0175] If the correctness check fails, instruct the Chinese character extension generation agent to regenerate the list of homophones and near-homophones of the target Chinese character.

[0176] In one or more embodiments, the generation module 53 generates multiple sets of two-round navigation correction dialogues based on multiple incorrect names of popular locations and popular locations, to be applied to the training data of the navigation model, specifically for:

[0177] For each incorrect name, an intelligent agent is constructed using dialogue data to build a set of two-round navigation correction dialogues based on the incorrect name and the target Chinese characters in popular locations;

[0178] The dialogue verification agent performs an integrity check on the two-round navigation correction dialogue. If it passes the check, the two-round navigation correction dialogue is added to the training data.

[0179] In one or more embodiments, the generation module 53 is further configured to:

[0180] If the result is unsatisfactory, the two rounds of navigation correction dialogue will be fed back to the dialogue data construction agent for reconstruction until the integrity check is passed or the number of feedbacks reaches the preset number. Then, the two most recent rounds of navigation correction dialogue will be added to the training data.

[0181] In one or more embodiments, the acquisition module 51 acquires a preset number of popular locations related to the target city, specifically for:

[0182] S1, obtains a preset number of popular locations related to the target city through a location-generating agent;

[0183] S2, through the location detection agent, performs deduplication and fake location removal based on multiple popular locations and a preset location database;

[0184] S3. If the first number of popular locations after processing is less than the preset number, repeat steps S1-S3 until the first number reaches the preset number.

[0185] It should be noted that the division of the various modules in the above device is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical element, or they can be physically separated. Furthermore, these modules can be implemented entirely in software through processing element calls, or entirely in hardware. Alternatively, some modules can be implemented through processing element calls in software, while others can be implemented in hardware. Moreover, these modules can be integrated together or implemented independently. The processing element here can be an integrated circuit with signal processing capabilities. During implementation, each step of the above method or each of the above modules can be completed through the integrated logic circuits in the hardware of the processor element or through software instructions.

[0186] As can be seen from the above, the training data generation device provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effect are similar, so this embodiment will not be described in detail here.

[0187] Figure 6 A schematic diagram of the structure of the electronic device provided in this application. Figure 6 As shown, the electronic device provided in this embodiment includes at least one processor 61 and a memory 62.

[0188] Optionally, the electronic device also includes a communication component 63.

[0189] The processor 61, memory 62, and communication component 63 are connected via bus 64.

[0190] In a specific implementation, at least one processor 61 executes computer execution instructions stored in memory 62, causing at least one processor 61 to perform the above-described method.

[0191] The specific implementation process of processor 61 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0192] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.

[0193] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.

[0194] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0195] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0196] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.

[0197] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.

[0198] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.

[0199] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.

[0200] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0201] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0202] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0203] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0204] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

Claims

1. A method for generating training data, characterized in that, include: Retrieve a preset number of popular locations related to the target city; For each popular location, multiple incorrect names are constructed based on the target Chinese characters in the popular location and a preset list of pronunciation errors; Based on the multiple incorrect names of the popular locations and the popular locations themselves, multiple sets of two-round navigation correction dialogues are generated to be applied to the training data of the navigation model.

2. The method according to claim 1, characterized in that, The process involves constructing multiple incorrect names for the popular locations based on the target Chinese characters in the popular locations and a preset list of pronunciation errors, including: Based on the target Chinese characters in the popular locations, the intelligent agent generates a list of homophones and near-homophones of the target Chinese characters. If the first letter of a homophone in the homophone list belongs to a letter in the pronunciation error list, the target Chinese character and the homophone list are sent to the Chinese character checking agent. After the Chinese character checking agent verifies the correctness of the target Chinese character and the list of homophones, the dialogue data construction agent generates multiple incorrect names for the popular location based on the target Chinese character and the list of homophones.

3. The method according to claim 2, characterized in that, The process of constructing an intelligent agent based on dialogue data to generate multiple incorrect names for the popular locations according to the target Chinese characters and the list of homophones and near-homophones includes: The intelligent agent constructed using the dialogue data replaces the target Chinese characters with characters from a list of homophones and near-homophones, generating multiple incorrect names for the popular locations.

4. The method according to claim 2 or 3, characterized in that, Before constructing an agent using dialogue data to generate multiple incorrect names for the popular locations based on the target Chinese character and the list of homophones and near-homophones, the method further includes: The intelligent agent for checking Chinese characters performs correctness checks on homophones / near-homophones and / or checks on the correctness of preset pronunciation errors based on the target Chinese character and the list of homophones / near-homophones. If the correctness check passes, the target Chinese character and the list of homophones and near-homophones will be sent to the dialogue data construction agent; If the correctness check fails, instruct the Chinese character extension generation agent to regenerate the list of homophones and near-homophones of the target Chinese character.

5. The method according to any one of claims 1-3, characterized in that, The step of generating multiple sets of two-round navigation correction dialogues based on multiple incorrect names of the popular locations and the popular locations themselves, for use as training data for the navigation model, includes: For each incorrect name, an intelligent agent is constructed using dialogue data to build a set of two-round navigation correction dialogues based on the incorrect name and the target Chinese characters in the popular locations; The dialogue verification agent performs an integrity check on the two-round navigation correction dialogue. If the check is successful, the two-round navigation correction dialogue is added to the training data.

6. The method according to claim 5, characterized in that, The method further includes: If the two rounds of navigation correction dialogue are not satisfactory, the two rounds of navigation correction dialogue are fed back to the dialogue data construction agent for reconstruction until the integrity check is passed or the number of feedbacks reaches a preset number. Then, the two most recent rounds of navigation correction dialogue are added to the training data.

7. The method according to any one of claims 1-3, characterized in that, The acquisition of a preset number of popular locations related to the target city includes: S1, obtain the preset number of multiple popular locations related to the target city through a location-generating agent; S2, the location detection agent performs deduplication and fake location removal based on the multiple popular locations and the preset location database; S3, if the first number of processed popular locations is less than the preset number, repeat steps S1-S3 until the first number reaches the preset number.

8. A training data generation apparatus, characterized in that, include: The acquisition module is used to acquire a preset number of popular locations related to the target city; A construction module is used to construct multiple incorrect names for each popular location based on the target Chinese characters in the popular location and a preset list of pronunciation errors; The generation module is used to generate multiple sets of two-round navigation correction dialogues based on the multiple incorrect names of the popular locations and the popular locations themselves, for use as training data for the navigation model.

9. An electronic device, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-7.