Voice navigation method and device, electronic equipment, storage medium and vehicle
By learning POI fragments and key features through a large model, the accuracy and robustness issues of POI recognition in the voice navigation system are solved, the dependence on large-scale POI lexicons is reduced, resource consumption is optimized, and the user experience is improved.
Patent Information
- Application Number
- CN202510722456.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-19
AI Technical Summary
In the existing technology, when processing multiple rounds of interactions, the voice navigation system lacks accuracy and robustness in recognizing POI fragment features and POI key features, resulting in a poor user voice interaction experience. In addition, the reliance on a large-scale POI vocabulary leads to high resource consumption and difficult maintenance.
A large model is used to learn POI fragment features and POI key features. The model is trained through historical voice navigation data to identify and fuse POI fragments and key feature texts, reduce dependence on the POI vocabulary, adapt to information integration from multiple rounds of conversations, and generate correct POI target information.
It improves the accuracy and robustness of POI recognition, reduces dependence on online POI lexicons, optimizes resource consumption, and enhances the user's voice interaction experience.
Smart Images

Figure CN120673755A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech recognition, and in particular to a speech navigation method, a speech navigation device, an electronic device, a storage medium, and a vehicle. Background Art
[0002] In the interaction in the car cockpit, the user initiates navigation instructions through the voice user interface (VUI); the user's voice instructions are converted into text through speech recognition (ASR), and the user's navigation intention is understood through the intent understanding service (NLU).
[0003] For example, if a user says "Navigate to the Ministry of Housing and Urban-Rural Development," ASR first converts the text into the corresponding text. The downstream NLU then understands that the user's intent to complete the navigation command is to "initiate navigation" and determines the navigation destination POI (Point of Information) "Ministry of Housing and Urban-Rural Development." However, due to differences in user expression habits and scenarios, or due to issues with ASR recognition, the user's actual desired POI information is scattered across multiple sentences. Each round of NLU can only process the intent of a single sentence and cannot correlate all sentences described by the user to accurately understand the user's navigation destination. This causes inconvenience in voice interaction and a very poor voice interaction experience for users.
[0004] The existing patent CN202410533169.7 proposes an improved and optimized method: relying on the POI vocabulary for NLU semantic retrieval, the POI fragment information scattered in each round of conversation is identified, and the final POI information is combined according to the specified rules based on the defined POI representative range size labels, their weights, and interaction order.
[0005] This method has the problem that the POI vocabulary is huge, and the full amount of POI data may reach hundreds of millions or even billions. Online real-time retrieval consumes a lot of memory resources, and maintenance and updating require a lot of manpower and time. The rules of combining according to POI range weight and interaction order are mostly applicable to multiple fragment POIs of decreasing levels, which are combined into a complete POI. It is difficult to correctly judge when there are intersections between multiple POI fragments and the information needs to be reorganized. It is highly dependent on the POI vocabulary, but if the ASR recognition is wrong or the POI information cannot be retrieved in the POI database, it cannot be correctly identified and combined.
[0006] Therefore, a voice navigation solution is needed, including an NLU system for multi-round POI intent understanding based on a large model. Leveraging the large model's rich knowledge and powerful language understanding capabilities, this system learns POI features representing different ranges of tags in the vocabulary, including POI fragment features and POI key features. This ensures that POI recognition results are both accurate and robust. By learning the internal associations and combinations of various POI fragments within a complete POI, it reduces reliance on online POI lexicons, adapts to more integrated methods of acquiring POI information during multi-round conversations, and generates accurate POI target information. Summary of the Invention
[0007] The purpose of the present invention is to provide a voice navigation method, a voice navigation device, an electronic device, a storage medium and a vehicle, which at least solve the problem of advance of POI fragment features and POI key features, solve the problem of fusion of POI fragment features and POI key features, and solve one of the technical problems of single-round interaction and multi-round interaction.
[0008] The present invention provides the following solutions:
[0009] According to one aspect of the present invention, a voice navigation method is provided, the voice navigation method comprising:
[0010] Get voice navigation history data;
[0011] Based on historical voice navigation data and large-scale model learning, the POI segment features and POI key feature recognition capabilities are acquired;
[0012] Process the current voice command text based on POI segment features and POI key feature recognition capabilities;
[0013] The processing of the current voice instruction text includes extracting the POI segment feature text and the POI key feature text in the current voice instruction text according to the POI segment feature and the POI key feature;
[0014] According to the POI fragment feature text and POI key feature text, POI text information is integrated;
[0015] According to the POI text information, the preset POI vocabulary is searched to obtain the recognition result of the navigation intention and output it on the human-computer interaction terminal;
[0016] Obtaining instructions for selecting and / or confirming navigation intent recognition results on the human-computer interaction end;
[0017] According to the selected and / or confirmed instructions, the target information in the identification result of the navigation intention is locked;
[0018] Plan the driving route based on the target information in the locked navigation intention information.
[0019] Furthermore, it also includes:
[0020] Obtain semantic recognition information of a single round of voice interaction and convert it into text information;
[0021] Extracting POI segment feature text and POI key feature text from the semantics of the single-round voice interaction based on the semantic recognition information of the single-round voice interaction;
[0022] Determine whether the current voice interaction is associated with navigation semantics based on the POI segment feature text and / or POI key feature text in the semantics of the single-round voice interaction;
[0023] If,associated, then the POI fragment feature text is searched in the preset POI vocabulary to determine whether the target information can be determined;
[0024] If yes, plan the driving route according to the determined target information.
[0025] Furthermore, the step of searching a preset POI vocabulary based on the POI segment feature text to determine whether the target information can be determined includes:
[0026] If not, then fuse the POI segment feature text and / or the POI key feature text to obtain the POI text information and determine whether the current voice interaction is semantically associated with navigation based on the POI text information;
[0027] If,associated, then the POI text information is searched in the preset POI vocabulary to determine whether the target information can be determined;
[0028] If yes, plan the driving route according to the determined target information.
[0029] Furthermore, the step of searching a preset POI vocabulary based on the POI text information to determine whether the target information can be determined includes:
[0030] If,no, then start a strategy of multiple rounds of interaction;
[0031] Among them, POI segment feature text and POI key feature text in the semantics of multiple rounds of voice interaction are extracted;
[0032] Determine whether the current voice interaction is associated with navigation semantics based on the POI segment feature text and / or POI key feature text in the semantics of multiple rounds of voice interaction;
[0033] If,associated, then search the preset POI vocabulary based on the POI fragment feature text or / and the POI key feature text fused POI text information to determine whether the target information can be determined;
[0034] If yes, plan the driving route according to the determined target information.
[0035] Furthermore, the determination of whether the target information can be determined includes:
[0036] If not, then generate prompt words based on the voice navigation history data and the determined destination information;
[0037] According to the playback of the prompt word, the human-computer interaction end obtains the target selection and / or confirmation instructions;
[0038] Lock the target information according to the target location selection and / or confirmation instructions of the human-computer interaction terminal;
[0039] Plan the driving route based on the locked target information.
[0040] Furthermore, the determination of whether the target information can be determined includes:
[0041] Eliminate non-navigational semantics and text;
[0042] A strategy for initiating a single-round interaction based on removing non-navigational semantics and text;
[0043] Based on the results of executing a single-round interaction strategy, input the initiated multi-round interaction strategy;
[0044] By executing multiple rounds of interaction strategies, the results of single-round interaction strategies are supplemented, corrected or verified.
[0045] According to two aspects of the present invention, a voice navigation device is provided, the voice navigation device comprising:
[0046] Historical data module, used to obtain voice navigation historical data;
[0047] The capability acquisition module is used to acquire POI segment features and POI key feature recognition capabilities based on historical voice navigation data and large-scale model learning;
[0048] The command processing module is used to process the current voice command text based on the POI segment features and POI key feature recognition capabilities;
[0049] The processing of the current voice instruction text includes extracting the POI segment feature text and the POI key feature text in the current voice instruction text according to the POI segment feature and the POI key feature;
[0050] A fusion module is used to fuse POI fragment feature text and POI key feature text into POI text information;
[0051] The navigation intention module is used to search the preset POI vocabulary based on the POI text information, obtain the recognition result of the navigation intention and output it on the human-computer interaction terminal;
[0052] An instruction acquisition module is used to obtain instructions for selecting and / or confirming the navigation intention recognition result of the human-computer interaction terminal;
[0053] An intention locking module is used to lock the target information in the navigation intention recognition result according to the selected and / or confirmed instructions;
[0054] The path planning module is used to plan the driving path according to the target information in the locked navigation intention information.
[0055] According to three aspects of the present invention, there is provided an electronic device, comprising: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;
[0056] A computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the voice navigation method.
[0057] According to four aspects of the present invention, a computer-readable storage medium is provided, which stores a computer program executable by an electronic device. When the computer program runs on the electronic device, the electronic device executes the steps of the voice navigation method.
[0058] According to five aspects of the present invention, there is provided a vehicle comprising:
[0059] An electronic device for implementing the steps of the voice navigation method according to any one of claims 1 to 6;
[0060] a processor, the processor running a program, and executing the steps of the voice navigation method based on data output by the electronic device when the program is running;
[0061] The storage medium is used to store a program, and when the program is running, it executes the steps of the voice navigation method for the data output from the electronic device.
[0062] Through the above solution, the following beneficial technical effects are achieved:
[0063] This application achieves both accuracy and robustness in POI recognition results by learning POI features representing different range labels in the vocabulary, such as POI segment features and POI key features.
[0064] This application reduces the dependence on the online POI vocabulary by learning the internal associations and combinations of various POI fragments in the complete POI, adapts to more integrated methods of obtaining POI information in multiple rounds of conversations, and generates correct POI target information. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 This is a flow chart of a voice navigation method provided by one or more embodiments of the present invention.
[0066] Figure 2 This is a structural diagram of a voice navigation device provided by one or more embodiments of the present invention.
[0067] Figure 3 It is a schematic diagram of a voice navigation system framework according to a specific embodiment of the present invention.
[0068] Figure 4 Schematic diagram of a multi-round interactive NLU functional module according to a specific embodiment of the present invention.
[0069] Figure 5 It is a schematic diagram of the process of multi-round interactive NLU according to a specific embodiment of the present invention.
[0070] Figure 6 It is a schematic diagram of the timing of multi-round interactive NLU according to a specific embodiment of the present invention.
[0071] Figure 7 The present invention provides a block diagram of the structure of an electronic device according to one or more embodiments of the voice navigation method. DETAILED DESCRIPTION
[0072] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0073] Figure 1 This is a flow chart of a voice navigation method provided by one or more embodiments of the present invention.
[0074] like Figure 1 The voice navigation method shown includes:
[0075] Step S1, obtaining voice navigation history data;
[0076] Step S2: Based on the historical voice navigation data and large-scale model learning, the POI segment features and POI key feature recognition capabilities are acquired;
[0077] Step S3, processing the current voice command text according to the POI segment features and POI key feature recognition capabilities;
[0078] Processing the current voice command text includes extracting POI segment feature text and POI key feature text in the current voice command text according to the POI segment feature and the POI key feature;
[0079] Step S4, fusing the POI segment feature text and the POI key feature text into POI text information;
[0080] Step S5, searching the preset POI vocabulary based on the POI text information, obtaining the recognition result of the navigation intention and outputting it on the human-computer interaction terminal;
[0081] Step S6, obtaining an instruction for the human-computer interaction terminal to select and / or confirm the navigation intention recognition result;
[0082] Step S7, according to the selected and / or confirmed instructions, locking the target information in the identification result of the navigation intention;
[0083] Step S8: planning a driving route according to the target information in the locked navigation intention information.
[0084] Specifically, POI key features include POI attribute features such as "building, shopping mall, residential area, and mall." While POI key feature text alone is generally insufficient to identify a specific POI address, text accompanied by POI key features is more likely to be a POI address than general text. For example, the word "Tianhong" alone is difficult to identify as a POI address. However, if "Tianhong" is combined with POI key feature text such as "building, shopping mall, residential area, and mall," the text becomes "Tianhong Building, Tianhong Shopping Mall, Tianhong Residential Area, and Tianhong Mall," and thus can be clearly identified as a POI address.
[0085] POI fragment features include text related to a POI address, such as "Navigate to," "Want to go," "Find it," and "Nearby," as well as text that is not a POI address but can be associated with a POI address, such as "XX Department" and "XXX Gas Station." Based on POI fragment features, text within text sentences that may be POI addresses can be identified. Alternatively, the POI fragment feature text can be used as a basis or clue to search the POI vocabulary and obtain the POI address.
[0086] Fusion of POI fragment features with POI key features can reduce the capacity of the POI vocabulary. For example, to increase the redundancy of the POI vocabulary, the POI vocabulary may refer to multiple text terms such as "XX Department," "XX Department in a certain district," "XX Department in a certain district of a certain city," and "XX Department Building" as the same POI address. However, by fusing POI fragment features with POI key features, only the text data for "XX Department" and "XX Department Building" needs to be prepared in the POI vocabulary. For example, if the POI fragment feature "XX Department" is intercepted from the speech text and hits "XX Department" in the POI vocabulary, then the POI key feature "Building" is intercepted and fused to form "XX Department Building," which then hits "XX Department Building" in the POI vocabulary. The resulting POI text information is valid POI address information.
[0087] To further confirm the navigation intention, the human-machine terminal can display the determined POI address, waiting for confirmation of the navigation intention or selection of one of the POI addresses. Finally, the driving route is planned based on the target information in the locked navigation intention information.
[0088] In this embodiment, it also includes:
[0089] Obtain semantic recognition information of a single round of voice interaction and convert it into text information;
[0090] Extracting POI segment feature text and POI key feature text from the semantics of the single-round voice interaction based on the semantic recognition information of the single-round voice interaction;
[0091] Determine whether the current voice interaction is associated with navigation semantics based on the POI segment feature text and / or POI key feature text in the semantics of the single-round voice interaction;
[0092] If,associated, then the POI fragment feature text is searched in the preset POI vocabulary to determine whether the target information can be determined;
[0093] If yes, plan the driving route according to the determined target information.
[0094] Specifically, the POI address is prioritized within a single voice interaction. Limiting the search to shorter or more concentrated text helps reduce computational overhead and minimizes interference from other sentences on navigation intent recognition.
[0095] The POI segment feature text and POI key feature text can be first extracted from the sentence of a single round of voice interaction, and can be separately or fused to determine whether the current voice interaction is associated with navigation semantics; further, the POI segment feature text or the fused POI segment feature text is searched in the preset POI vocabulary to determine whether the target information can be determined; further, the driving route is planned based on the determined target information.
[0096] In this embodiment, the target information is determined by searching a preset POI vocabulary based on the POI segment feature text, including:
[0097] If not, then fuse the POI segment feature text and / or the POI key feature text to obtain the POI text information and determine whether the current voice interaction is semantically associated with navigation based on the POI text information;
[0098] If,associated, then the POI text information is searched in the preset POI vocabulary to determine whether the target information can be determined;
[0099] If yes, plan the driving route according to the determined target information.
[0100] Specifically, fusing the POI segment feature text and / or the POI key feature text includes fusing two POI segment feature texts and fusing the POI segment feature text and the POI key feature text.
[0101] For example, the POI segment feature texts "Highway Exit Gas Station" and "PetroChina Gas Station" are merged to remove duplicate fields and become "Highway Exit PetroChina Gas Station".
[0102] In this embodiment, the target information is determined by searching a preset POI vocabulary based on the POI text information, including:
[0103] If,no, then start a strategy of multiple rounds of interaction;
[0104] Among them, POI segment feature text and POI key feature text in the semantics of multiple rounds of voice interaction are extracted;
[0105] Determine whether the current voice interaction is associated with navigation semantics based on the POI segment feature text and / or POI key feature text in the semantics of multiple rounds of voice interaction;
[0106] If,associated, then search the preset POI vocabulary based on the POI fragment feature text or / and the POI key feature text fused POI text information to determine whether the target information can be determined;
[0107] If yes, plan the driving route according to the determined target information.
[0108] Specifically, in one embodiment, in the first single-round voice interaction, the POI segment feature text is "find", "nearby", "PetroChina", and the POI key feature text is "gas station". After fusion, the navigation intention of "find a nearby PetroChina gas station" is obtained. According to the navigation intention, the map annotation of the PetroChina gas station is searched, and then the driving route is planned through the selection and confirmation instructions on the human-machine terminal.
[0109] In the second single-round voice interaction, the phrase "Forget it, let's go with Shell" is recognized. Based on the fusion of the POI segment feature text and / or the POI key feature text in the semantics of the multi-round voice interaction, the navigation intent of "Find a nearby Sinopec gas station" in the first single-round voice interaction is corrected, and at least the output options of "Find a nearby Sinopec gas station" and "Find a nearby Shell gas station" are output, waiting for the selection / confirmation command from the human-machine terminal.
[0110] In this embodiment, determining whether the target information can be determined includes:
[0111] If not, then generate prompt words based on the voice navigation history data and the determined destination information;
[0112] According to the playback of the prompt word, the human-computer interaction end obtains the target selection and / or confirmation instructions;
[0113] Lock the target information according to the target location selection and / or confirmation instructions of the human-computer interaction terminal;
[0114] Plan the driving route based on the locked target information.
[0115] Specifically, historical voice navigation data includes historical data on selections and confirmations made on the human-machine interface. This historical data is used to calculate the probability of selecting a particular POI address. The currently determined destination is simply the result of this interaction. Based on the selection and / or confirmation instructions made by the human-machine interface, the target information can be locked and further iterated on the historical data.
[0116] In this embodiment, determining whether the target information can be determined includes:
[0117] Eliminate non-navigational semantics and text;
[0118] A strategy for initiating a single-round interaction based on removing non-navigational semantics and text;
[0119] Based on the results of executing a single-round interaction strategy, input the initiated multi-round interaction strategy;
[0120] By executing multiple rounds of interaction strategies, the results of single-round interaction strategies are supplemented, corrected or verified.
[0121] Specifically, in one embodiment, during the first single-round interaction, non-navigational semantics and text are removed to minimize interference with the recognition of POI segment feature text and / or POI key feature text. For example, "PetroChina is good" and "The gas card is from PetroChina" contain the POI segment feature text "PetroChina" but are not relevant to navigation intent. Such non-navigational semantics and text are removed.
[0122] Based on the results of executing a single-round interaction strategy, a multi-round interaction strategy initiated by input includes the following: in the first round of single-round interaction, the navigation intent obtained was "nearest Shell gas station", but during multiple rounds of interaction, the semantics of "PetroChina is good" and "the fuel card is from PetroChina" appeared. Further, using the semantics of "PetroChina is good" and "the fuel card is from PetroChina", the current navigation intent option of "nearest Shell gas station" can be supplemented or corrected to "nearest gas station". Another example involves verifying that the choice of "Shell gas station" in the first single-round interaction was actually the choice of "nearest gas station" by verifying that "Shell gas station" is closer to the current location than "PetroChina gas station".
[0123] Figure 2 This is a structural diagram of a voice navigation device provided by one or more embodiments of the present invention.
[0124] like Figure 2 The voice navigation device shown includes: a historical data module, a capability acquisition module, a command processing module, a fusion module, a navigation intention module, a command acquisition module, an intention locking module, and a path planning module;
[0125] Historical data module, used to obtain voice navigation historical data;
[0126] The capability acquisition module is used to acquire POI segment features and POI key feature recognition capabilities based on historical voice navigation data and large-scale model learning;
[0127] The command processing module is used to process the current voice command text based on the POI segment features and POI key feature recognition capabilities;
[0128] Processing the current voice command text includes extracting POI segment feature text and POI key feature text in the current voice command text according to the POI segment feature and the POI key feature;
[0129] A fusion module is used to fuse POI fragment feature text and POI key feature text into POI text information;
[0130] The navigation intention module is used to search the preset POI vocabulary based on the POI text information, obtain the recognition result of the navigation intention and output it on the human-computer interaction terminal;
[0131] An instruction acquisition module is used to obtain instructions for selecting and / or confirming the navigation intention recognition result of the human-computer interaction terminal;
[0132] An intention locking module is used to lock the target information in the navigation intention recognition result according to the selected and / or confirmed instructions;
[0133] The path planning module is used to plan the driving path according to the target information in the locked navigation intention information.
[0134] It is worth noting that although this system only discloses the historical data module, capability acquisition module, instruction processing module, fusion module, navigation intention module, instruction acquisition module, intention locking module, and path planning module, it does not mean that this device is limited to the above-mentioned basic functional modules. On the contrary, what the present invention wants to express is that, based on the above-mentioned basic functional modules, those skilled in the art can arbitrarily add one or more functional modules in combination with the existing technology to form an infinite number of embodiments or technical solutions. In other words, this system is open rather than closed. Just because this embodiment only discloses individual basic functional modules, it cannot be considered that the scope of protection of the claims of the present invention is limited to the above-mentioned basic functional modules.
[0135] Figure 3 It is a schematic diagram of a voice navigation system framework according to a specific embodiment of the present invention.
[0136] Figure 4 Schematic diagram of a multi-round interactive NLU functional module according to a specific embodiment of the present invention.
[0137] Figure 5 It is a schematic diagram of the process of multi-round interactive NLU according to a specific embodiment of the present invention.
[0138] Figure 6 It is a schematic diagram of the timing of multi-round interactive NLU according to a specific embodiment of the present invention.
[0139] In a specific embodiment, Figure 3 The voice navigation system framework shown includes the following: the vehicle-mounted device collects the user's voice, transmits the navigation intention voice to the cloud service, the ASR converts the voice into text and transmits it to the semantic recognition NLU service to identify the intention, and combines the cached historical multiple rounds of information with context information to generate prompt words for the large model. Finally, new POI information data is proposed and integrated and sent to the vehicle-mounted device to execute the user's instructions for POI retrieval.
[0140] In another specific embodiment, Figure 4 The multi-round interaction NLU functional module shown includes four modules: single-round interaction NLU module, context cache module, prompt word engineering module, and POI information integration large model module.
[0141] Single-round interaction NLU module: This module implements single-round semantic recognition through a traditional semantic recognition model. During in-vehicle voice interaction, semantic recognition is used to distinguish whether the current interaction semantics are navigation-related, and to exclude non-navigation semantic text (such as opening a window or turning on the air conditioner).
[0142] Context Cache: This module uses Redis to cache navigation-related voice interaction data from the last 10 rounds within 5 minutes. Each round of data includes the voice text, the time difference (in seconds) between the previous interaction and the current one, as well as single-round semantics and point of interest information, and multi-round semantics and point of interest information.
[0143] Prompt word engineering module: receives the current round of interaction data and the historical interaction data of the context cache, and generates the prompt words required for the POI information integration model.
[0144] POI information integration large model module: The multi-round voice interaction POI destination information integration assistant understands the information based on the prompt words and executes it according to its defined process, and finally produces a complete POI information data.
[0145] In another specific embodiment, Figure 5 、 6 The process and timing of the multi-round interactive NLU shown include:
[0146] After receiving the user's voice text recognized by the vehicle's ASR, the text is passed to the single-round interaction NLU for semantic understanding and recognition, determining whether the current interaction semantics are navigation semantic information. If not, it is directly sent to the vehicle to execute other non-navigation commands. Otherwise, it proceeds to the next step.
[0147] Get the historical interaction information data stored in the context cache.
[0148] After the prompt word project receives the current round of interaction information and the historical interaction information, it generates the corresponding prompt word and passes it into the multi-round voice interaction POI destination information integration.
[0149] The multi-round voice interaction POI destination information integration assistant understands the information based on the prompt words and executes it according to its defined process, ultimately producing a complete POI information data, which is finally sent to the in-vehicle system.
[0150] In this embodiment, the main text sample of the prompt word is as follows:
[0151] 1) Role: Multi-round voice interaction POI destination information integration assistant.
[0152] 2) Qualifications: Possess extensive knowledge of POI information, a deep understanding of the composition and interrelationships of various POIs, and a thorough understanding of various POI types, including landmarks, institutions, commercial locations, and transportation hubs across the country. Skilled in single-text semantic recognition, accurate identification, and extraction of POI data, as well as efficient integration of POI fragments from multi-text semantics.
[0153] 3) Task: Based on the navigation information of multiple rounds of user interactions, including the voice text of the current round of interaction and historical navigation information data, accurately extract the POI fragment data scattered in multiple rounds of conversations, and integrate them into new POI data that conforms to the POI format according to specific rules.
[0154] 4) Integration rules:
[0155] a. Combination: When the two POI fragments are combined to form a standard POI information, such as "I want to go to Kunming West" and "Bus Station", they are integrated into "Kunming West Bus Station".
[0156] b. Partial replacement: If the POI information contained in the following text replaces the previous POI information, such as "Find me a Sinopec gas station" or "Forget it, let's go with Shell", it will be integrated into "Shell gas station".
[0157] c. Overlay update: When a new round of POI information overwrites and updates historical POI information, such as "Navigate to Ministry of Housing and Urban-Rural Development" and "Beijing Ministry of Housing and Urban-Rural Development", it will eventually be integrated into "Beijing Ministry of Housing and Urban-Rural Development".
[0158] d. Supplementation: If the new round of POI information is a supplement to the previous one, it can make the final POI information more specific and complete. For example, "Navigate to Agricultural University", "Swimming Pool", "Changping Campus" can be integrated into "Agricultural University Changping Campus Swimming Pool".
[0159] 5) Execution process:
[0160] a. First, carefully determine whether the conversation text contains POI information. If so, perform accurate extraction.
[0161] b. Next, combine the current and historical POI information, as well as the time difference between historical interactions, to conduct a comprehensive analysis to confirm possible correlations.
[0162] c. Then, deeply understand the semantics of the entire multi-round conversation text and accurately determine the POI information and integration rules that need to be integrated.
[0163] d. Finally, output complete, accurate, and clear POI information based on the integration rules to ensure that users can quickly and accurately understand and use the destination information for navigation.
[0164] 6) Navigation information for multiple rounds of interaction includes:
[0165] a. The voice text of the current interaction: "Haidian District";
[0166] b. Historical navigation information data:
[0167] a) First round: semantic text: "Navigate to the Ministry of Housing and Urban-Rural Development", time difference: 30 seconds.
[0168] b) Second round: semantic text: "Beijing Municipal Ministry of Housing and Urban-Rural Development", time difference: 20 seconds.
[0169] Statement: The text used in the above embodiments is only the text of experimental data, which is used to illustrate the technical route of this application and does not involve the collection and processing of real sensitive information.
[0170] Figure 7 The present invention is a block diagram of an electronic device structure of a voice navigation method provided by one or more embodiments of the present invention.
[0171] like Figure 7 As shown, the present application provides an electronic device, comprising: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;
[0172] The memory stores a computer program, which, when executed by the processor, enables the processor to perform steps of a voice navigation method.
[0173] The present application also provides a computer-readable storage medium storing a computer program executable by an electronic device. When the computer program runs on the electronic device, the electronic device executes the steps of a voice navigation method.
[0174] The present application also provides a vehicle, comprising:
[0175] An electronic device for implementing the steps of the voice navigation method;
[0176] a processor, the processor running a program, and executing the steps of the voice navigation method based on data output by the electronic device when the program is running;
[0177] The storage medium is used to store a program, and when the program is run, the program executes the steps of the voice navigation method for data output from the electronic device.
[0178] The communication bus mentioned in the electronic device mentioned above may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not mean that there is only one bus or only one type of bus.
[0179] The electronic device includes a hardware layer, an operating system layer running on the hardware layer, and an application layer running on the operating system. The hardware layer includes hardware such as a central processing unit (CPU), a memory management unit (MMU), and memory. The operating system can be any one or more computer operating systems that control electronic devices through processes, such as the Linux operating system, the Unix operating system, the Android operating system, the iOS operating system, or the Windows operating system. In the embodiments of the present invention, the electronic device can be a handheld device such as a smartphone or a tablet computer, or an electronic device such as a desktop computer or a portable computer, which is not particularly limited in the embodiments of the present invention.
[0180] The execution subject of the electronic device control in the embodiment of the present invention can be an electronic device, or a functional module in the electronic device that can call a program and execute the program. The electronic device can obtain the firmware corresponding to the storage medium. The firmware corresponding to the storage medium is provided by the supplier. The firmware corresponding to different storage media can be the same or different, and is not limited here. After the electronic device obtains the firmware corresponding to the storage medium, it can write the firmware corresponding to the storage medium into the storage medium, specifically, burn the firmware corresponding to the storage medium into the storage medium. The process of burning the firmware into the storage medium can be implemented using existing technology and will not be described in detail in the embodiment of the present invention.
[0181] The electronic device can also obtain a reset command corresponding to the storage medium. The reset command corresponding to the storage medium is provided by the supplier. The reset commands corresponding to different storage media can be the same or different, and are not limited here.
[0182] In this case, the storage medium of the electronic device is a storage medium in which the corresponding firmware is written. The electronic device can respond to the reset command corresponding to the storage medium in which the corresponding firmware is written, thereby resetting the storage medium in which the corresponding firmware is written according to the reset command corresponding to the storage medium. The process of resetting the storage medium according to the reset command can be implemented in the existing technology and will not be described in detail in the embodiments of the present invention.
[0183] For the convenience of description, the above devices are described as various units and modules according to their functions. Of course, when implementing this application, the functions of each unit and module can be implemented in the same or multiple software and / or hardware.
[0184] Those skilled in the art will understand that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by those skilled in the art in the art to which the present invention pertains. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with those in the context of the prior art and, unless specifically defined, will not be interpreted in an idealized or overly formal sense.
[0185] For simplicity of description, the method embodiments are described as a series of actions. However, those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because certain steps can be performed in other orders or simultaneously according to the embodiments of the present invention. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.
[0186] Through the description of the above embodiments, it can be seen that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general-purpose hardware platform. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application or certain parts of the embodiments.
[0187] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A voice navigation method, characterized in that: The voice navigation method comprises: Get voice navigation history data; Based on historical voice navigation data and large-scale model learning, the POI segment features and POI key feature recognition capabilities are acquired; Process the current voice command text based on POI segment features and POI key feature recognition capabilities; The processing of the current voice instruction text includes extracting the POI segment feature text and the POI key feature text in the current voice instruction text according to the POI segment feature and the POI key feature; According to the POI fragment feature text and POI key feature text, POI text information is integrated; According to the POI text information, the preset POI vocabulary is searched to obtain the recognition result of the navigation intention and output it on the human-computer interaction terminal; Obtaining instructions for selecting and / or confirming navigation intent recognition results on the human-computer interaction end; According to the selected and / or confirmed instructions, the target information in the identification result of the navigation intention is locked; Plan the driving route based on the target information in the locked navigation intention information.
2. The voice navigation method according to claim 1, wherein: Also includes: Obtain semantic recognition information of a single round of voice interaction and convert it into text information; Extracting POI segment feature text and POI key feature text from the semantics of the single-round voice interaction based on the semantic recognition information of the single-round voice interaction; Determine whether the current voice interaction is associated with navigation semantics based on the POI segment feature text and / or POI key feature text in the semantics of the single-round voice interaction; If,associated, then the POI fragment feature text is searched in the preset POI vocabulary to determine whether the target information can be determined; If yes, plan the driving route according to the determined target information.
3. The voice navigation method according to claim 2, characterized in that: The step of searching a preset POI vocabulary based on the POI segment feature text to determine whether the target information can be determined includes: If not, then fuse the POI segment feature text and / or the POI key feature text to obtain the POI text information and determine whether the current voice interaction is semantically associated with navigation based on the POI text information; If,associated, then the POI text information is searched in the preset POI vocabulary to determine whether the target information can be determined; If yes, plan the driving route according to the determined target information.
4. The voice navigation method according to claim 3, characterized in that: The step of searching a preset POI vocabulary based on the POI text information to determine whether the target information can be determined includes: If,no, then start a strategy of multiple rounds of interaction; Among them, POI segment feature text and POI key feature text in the semantics of multiple rounds of voice interaction are extracted; Determine whether the current voice interaction is associated with navigation semantics based on the POI segment feature text and / or POI key feature text in the semantics of multiple rounds of voice interaction; If,associated, then search the preset POI vocabulary based on the POI fragment feature text or / and the POI key feature text fused POI text information to determine whether the target information can be determined; If yes, plan the driving route according to the determined target information.
5. The voice navigation method according to claim 4, characterized in that: The determination of whether the target information can be determined includes: If not, then generate prompt words based on the voice navigation history data and the determined destination information; According to the playback of the prompt word, the human-computer interaction end obtains the target selection and / or confirmation instructions; Lock the target information according to the target location selection and / or confirmation instructions of the human-computer interaction terminal; Plan the driving route based on the locked target information.
6. The voice navigation method according to claim 5, characterized in that: The determination of whether the target information can be determined includes: Eliminate non-navigational semantics and text; A strategy for initiating a single-round interaction based on removing non-navigational semantics and text; Based on the results of executing a single-round interaction strategy, input the initiated multi-round interaction strategy; By executing multiple rounds of interaction strategies, the results of single-round interaction strategies are supplemented, corrected or verified.
7. A voice navigation device, characterized in that: The voice navigation device comprises: Historical data module, used to obtain voice navigation historical data; The capability acquisition module is used to acquire POI segment features and POI key feature recognition capabilities based on historical voice navigation data and large-scale model learning; The command processing module is used to process the current voice command text based on the POI segment features and POI key feature recognition capabilities; The processing of the current voice instruction text includes extracting the POI segment feature text and the POI key feature text in the current voice instruction text according to the POI segment feature and the POI key feature; A fusion module is used to fuse POI fragment feature text and POI key feature text into POI text information; The navigation intention module is used to search the preset POI vocabulary based on the POI text information, obtain the recognition result of the navigation intention and output it on the human-computer interaction terminal; An instruction acquisition module is used to obtain instructions for selecting and / or confirming the navigation intention recognition result of the human-computer interaction terminal; An intention locking module is used to lock the target information in the navigation intention recognition result according to the selected and / or confirmed instructions; The path planning module is used to plan the driving path according to the target information in the locked navigation intention information.
8. An electronic device, characterized in that: include: A processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus; A computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the voice navigation method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that A computer program executable by an electronic device is stored, and when the computer program is run on the electronic device, the electronic device executes the steps of the voice navigation method according to any one of claims 1 to 6.
10. A vehicle, characterized in that: include: An electronic device for implementing the steps of the voice navigation method according to any one of claims 1 to 6; a processor, wherein the processor runs a program, and when the program runs, the steps of the voice navigation method according to any one of claims 1 to 6 are executed based on data output from the electronic device; A storage medium for storing a program, wherein when the program is run, the program executes the steps of the voice navigation method according to any one of claims 1 to 6 for data output from the electronic device.
Citation Information
Patent Citations
Voice navigation method and device, electronic equipment, storage medium and intelligent cabin
CN118500436A