Method and device for determining interest point from multiple rounds of dialogues, equipment and medium
By identifying the rewrite intention in the voice navigation system and generating the second slot value, the problem of incomplete correction of interest points in the voice navigation system is solved, and the user's automated processing of correcting interest points through voice is realized.
Patent Information
- Application Number
- CN202510471768.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-08-01
AI Technical Summary
When the existing voice navigation system recognizes errors, users need to manually correct the error text, which violates the original intention of voice navigation and cannot correct the positioning error of interest points through voice.
By identifying the rewrite intentions in multiple rounds of dialogue, generating the second slot value and rewriting the interest point slot, the positioning of the voice-corrected interest point is achieved.
Users can correct the wrong points of interest determined by previous conversations through voice, reduce manual operations, and improve the automatic error correction ability of the navigation system.
Smart Images

Figure CN120403597A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of artificial intelligence technology, and particularly to a method, apparatus, device, and medium for determining points of interest from multi-round conversations. Background Art
[0002] Voice navigation is a technology that uses voice commands and feedback to guide users for navigation, and is usually applied to car navigation systems and smartphone applications. It provides real-time driving routes and traffic information through GPS positioning, helping users reach their destinations safely and conveniently. Users can set it through voice commands, avoiding manual operations and improving driving safety and efficiency.
[0003] In a voice navigation task, the dialogue system needs to convert voice commands into text. However, due to reasons such as non-standard Mandarin, confusion between front and back nasal sounds, and confusion between flat and rolled tongue sounds, one or more words in the voice command may be misrecognized, resulting in the voice command being converted into an incorrect text that is not what the user wants. This makes the dialogue system locate points of interest that do not meet the user's expectations based on the incorrect text. Even if the user repeats the same voice command, due to the limitations of the voice recognition algorithm, the voice command cannot be correctly converted into the correct target text.
[0004] Currently, the common practice for the correction mechanism for voice recognition errors in the voice navigation scenario is that the user needs to manually rewrite one or more words in the incorrect text and then initiate navigation again. However, this goes against the original intention of voice navigation and still cannot truly avoid the problem of manual operations. Summary of the Invention
[0005] To overcome the problems existing in the related art, this specification provides a method, apparatus, device, and medium for determining points of interest from multi-round conversations.
[0006] According to the first aspect of the embodiments of this specification, a method for determining points of interest from multi-round conversations is provided. The method includes:
[0007] In response to receiving a voice navigation command for a non-first-round voice navigation conversation input, identifying the intent type corresponding to the voice navigation command;
[0008] When the intent type is a rewrite intent, obtaining the voice navigation command and the first slot value filled in the point-of-interest slot, and generating a second slot value based on the voice navigation command and the first slot value. The first slot value is determined by the conversation before the non-first-round voice navigation conversation;
[0009] Rewriting the first slot value in the point-of-interest slot to the second slot value. The rewritten point-of-interest slot is used to screen out corresponding target points of interest from the points of interest on the map in response to the non-first-round voice navigation conversation.
[0010] According to a second aspect of the embodiments of the present specification, there is provided a device for determining an interest point from a multi-round conversation, the device comprising:
[0011] An intent recognition module, configured to recognize an intent type corresponding to the voice navigation command in response to receiving a voice navigation command input in a non-first-round voice navigation conversation;
[0012] A second slot value determination module, configured to, when the intent type is a rewrite intent, obtain the voice navigation command and a first slot value filled in the interest point slot, and generate a second slot value according to the voice navigation command and the first slot value, where the first slot value is determined by a conversation before the non-first-round voice navigation conversation;
[0013] A rewrite module, configured to rewrite the first slot value in the interest point slot as the second slot value, and the rewritten interest point slot is used to screen out corresponding target interest points from the interest points on the map in response to the non-first-round voice navigation conversation.
[0014] According to a third aspect of the embodiments of the present specification, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, the steps of the method described in the first aspect are implemented.
[0015] According to a fourth aspect of the embodiments of the present specification, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method described in the first aspect are implemented.
[0016] The technical solutions provided by the embodiments of the present specification may include the following beneficial effects:
[0017] In the embodiments of the present specification, for a non-first-round voice navigation conversation in a multi-round conversation, the intent type corresponding to the voice navigation command input in the non-first-round voice navigation conversation is recognized. When the intent type is a rewrite intent, a second slot value is regenerated by combining the current voice navigation command and the first slot value determined by the conversation before it, and the original first slot value in the current interest point slot is rewritten as the second slot value. The rewritten interest point slot is used to screen out corresponding target interest points from the interest points on the map in response to the current non-first-round voice navigation conversation. It can be seen that in any non-first-round voice navigation conversation in a multi-round conversation, by recognizing the voice correction instruction input by the user and rewriting the original first slot value in the interest point slot as the second slot value indicated by the voice correction instruction, the purpose of the user to correct the incorrect interest point determined by the previous conversation through voice is achieved, thereby solving the problem of imperfect correction mechanism for interest points in related conversation methods in a voice navigation scenario.
[0018] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this specification. Brief Description of the Drawings
[0019] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this specification, and are used together with the specification to explain the principles of this specification.
[0020] Figure 1 is a flowchart of a method for determining an interest point from a multi-round conversation shown in accordance with an exemplary embodiment of this specification.
[0021] Figure 2 is an application scenario diagram of a method for determining an interest point from a multi-round conversation shown in accordance with an exemplary embodiment of this specification.
[0022] Figure 3 is a schematic diagram of generating a rewritten text shown in accordance with an exemplary embodiment of this specification.
[0023] Figure 4 is a schematic structural diagram of an electronic device shown in accordance with an exemplary embodiment of this specification.
[0024] Figure 5 is a block diagram of a device for determining an interest point from a multi-round conversation shown in accordance with an exemplary embodiment of this specification. Detailed Embodiments
[0025] Voice navigation is a technology that uses voice commands and feedback to guide users for navigation, and is usually applied to car navigation systems and smartphone applications. It provides real-time driving routes and traffic information through GPS positioning, helping users reach their destinations safely and conveniently. Users can set it through voice commands, avoiding manual operations and improving driving safety and efficiency.
[0026] In a voice navigation task, the dialogue system needs to convert voice commands into text. However, due to reasons such as non-standard Mandarin, confusion between front and back nasal sounds, and confusion between flat and rolled tongue sounds, one or more words in the voice command may be misrecognized, resulting in the voice command being converted into an incorrect text that is not what the user wants, causing the dialogue system to locate an interest point that does not meet the user's expectations based on the incorrect text. Even if the user repeats the same voice command again, due to the limitations of the voice recognition algorithm, the voice command cannot be correctly converted into the correct target text.
[0027] Of course, some users may also try to modify the point of interest (POI) currently located by the dialogue system of the related technology by correcting the instruction through voice input. For example, in the first round of conversation, the user inputs the voice "Navigate to **Impression**". At this time, the dialogue system or the map software called by the dialogue system can output a list of POIs related to "**Impression**". However, the user may find that the expected POI does not exist in the list of POIs. The user may try to correct the name of the POI by inputting the voice "It's the character 'yin' of music, not the character 'yin' of impression" in the next round of conversation. However, the current dialogue system does not pay attention to this requirement, that is, the user has the need to try to correct the name of the POI through voice. Accordingly, the corresponding intent recognition has not been developed either. When the user tries to correct the instruction through voice input, the dialogue system may misclassify the instruction into other intents. For example, it may misidentify the user's intent as still being a navigation intent and perform the action of extracting the POI slot from the current conversation according to the navigation task, and ultimately still cannot achieve the purpose of rewriting the POI.
[0028] Therefore, the common practice of the correction mechanism for speech recognition error problems in the current voice navigation scenario is that the user needs to manually rewrite one or more words in the error text and then initiate navigation again. However, this goes against the original intention of voice navigation and still cannot truly avoid the problem of manual operation.
[0029] In view of the above technical problems, a method for determining a POI from multi-round conversations is proposed to achieve the purpose of allowing the user to correct the incorrect POI determined from the historical conversation through voice.
[0030] Next, the embodiments of this specification will be described in detail.
[0031] Figure 1 FIG. is a flowchart of a method for determining a POI from multi-round conversations shown according to an exemplary embodiment of this specification. As Figure 1 shown, it includes steps 101-103:
[0032] Step 101: In response to receiving a voice navigation command input in a non-first-round voice navigation conversation, identify the type of intent corresponding to the voice navigation command.
[0033] Step 102: When the type of intent is a rewriting intent, obtain the voice navigation command and the first slot value filled in the POI slot, and generate a second slot value according to the voice navigation command and the first slot value. The first slot value is determined by the conversation before the non-first-round voice navigation conversation.
[0034] Step 103: Rewrite the first slot value in the POI slot as the second slot value. The rewritten POI slot is used to screen out the corresponding target POI from the POIs of the map in response to the non-first-round voice navigation conversation.
[0035] This method can be applied to the dialogue systems of voice assistants in various map navigation software, voice assistants on various intelligent devices, in-vehicle voice assistants, etc.
[0036] A point of interest (POI) can be a geographical location entity marked with a specific function or service on a map. For example, a hot pot restaurant, a hospital, a university, etc. Of course, it can also be a specific geographical location, such as a road, a street, etc.
[0037] In the application scenario of point-of-interest navigation, the user can express a demand associated with a point of interest, and the dialogue system helps them quickly find the target point of interest that meets the demand. Specifically, in the dialogue system, the user can express the demand associated with the point of interest through voice input or text input. For example, if the user wants to find the gas station closest to them, they can start a dialogue "Where is the nearest gas station?", and then the dialogue system can show a list of nearby gas stations, and the gas stations in the list can be sorted by distance. Another example is that if the user wants to find a Class III Grade A hospital, they can start a dialogue "I want to go to a Class III Grade A hospital", and then the dialogue system can show a list of Class III Grade A hospitals in the user's city.
[0038] Each round of dialogue in the multi-round dialogue in this solution may not be limited to being a dialogue for voice navigation tasks. For example, in an in-vehicle voice assistant, the user starts three rounds of dialogue. The first round of dialogue is "Open the window", the second round of dialogue is "Navigate to **First Hospital", and the third round of dialogue is "The 'First' in 'one, two, three, four', not the 'Medical' in 'hospital'". It can be seen that there can be other task-type dialogues in the multi-round dialogue. Of course, other rounds of dialogue except for voice navigation dialogue may not be limited to being in the form of voice interaction, and can also be in the form of text input. Also, it should be noted that the non-first-round voice navigation dialogue in the multi-round dialogue in this solution can be any non-first-round voice navigation dialogue in the multi-round dialogue.
[0039] Intent recognition can be to determine the hidden intention or purpose behind the user's input text or voice by analyzing it, and classify it into a predefined intent type, so as to drive the execution of the subsequent task corresponding to this intent type. For example, assume that the user inputs "Book a flight to Shanghai tomorrow" in this round of dialogue, then the content input by the user can be classified into the flight booking intent, and thus corresponding actions can be triggered, such as calling the flight booking interface to execute the flight booking task.
[0040] The intention of rewriting can be generally understood as the intention of a user to correct the name of a point of interest in the list of points of interest output by the dialogue system through voice guidance when the name in the list does not meet the expectations in the navigation point of interest scenario. For example, for the user input of "It's the character for'music' instead of the character for 'impression'", the intention type of this input can be recognized as the "rewriting intention", while for the user input of "Open the sunroof", the intention type of this input can be recognized as the "open window intention". Of course, in addition to the rewriting intention of this solution, the dialogue system may also perform other intention recognitions on the user input at the same time, and this solution does not impose any restrictions on the methods of other intention recognitions and how to execute tasks after recognition.
[0041] In a dialogue system, a slot can be a predefined variable used to extract key information related to the intention from the user input. For example, "time", "location", and "service type", etc. The slot value is the specific data corresponding to the slot in the user input, that is, the actual content filled into the slot. For example, in the intention type of "flight reservation", the slot of "departure city" can be filled with "Beijing". Correspondingly, the point of interest slot can be a slot for filling in the point of interest. For example, if the point of interest information extracted from the user input is "restaurant", the point of interest slot can be specifically filled with "restaurant"; if the point of interest information extracted from the user input is "hospital", the point of interest slot can be specifically filled with "hospital".
[0042] In one embodiment, for non-first-round voice navigation dialogues in a multi-round dialogue, the non-first-round voice navigation dialogue can be any non-first-round voice navigation dialogue in the multi-round dialogue, and the following execution logics for each non-first-round voice navigation dialogue are the same. Next, only a certain non-first-round voice navigation dialogue in the multi-round dialogue will be introduced:
[0043] In response to the voice navigation command input in the non-first-round voice navigation dialogue, identify the intention type corresponding to the voice navigation command. In the case where the intention type is the rewriting intention, obtain the voice navigation command and the first slot value filled in the point of interest slot, and generate a second slot value based on the voice navigation command and the first slot value. The first slot value is determined by the dialogue before the non-first-round voice navigation dialogue. Among them, the same point of interest slot can be processed for each round of dialogue in the multi-round dialogue.
[0044] Rewrite the first slot value in the point of interest slot as the second slot value. The rewritten point of interest slot is used to filter out the corresponding target point of interest from the points of interest on the map in response to a non-first-round voice navigation dialogue. Of course, for the point of interest slot in any state before the non-first-round voice navigation dialogue, it can be used to filter out the corresponding target point of interest from the points of interest on the map in response to the current round of dialogue when the slot value is filled in / rewritten therein. Among them, this solution does not limit the method of filtering out the corresponding target point of interest from the points of interest on the map according to the rewritten point of interest slot. For example, it can be that the dialogue system calls the navigation map software and jumps to the target point of interest displayed in the navigation map software, or directly displays the target point of interest filtered out from the points of interest on the map in the dialogue system.
[0045] Figure 2 This is an application scenario diagram for determining a point of interest from a multi-round dialogue shown in this specification according to an exemplary embodiment. As Figure 2 shown, assume that the user starts two rounds of dialogue in the dialogue system. The content input in the first round of dialogue is "Navigate to **Impression**". At this time, the first slot value extracted from the input is "**Impression**" and filled into the point of interest slot. This point of interest slot is used by the dialogue system to return a point of interest list 20 to the user in response to the user's input. The user finds that the point of interest determined by the dialogue system is "**Impression**" instead of "**Audio-Visual**" that the user wants to input, resulting in the point of interest list 20 not having the point of interest expected by the user at all. Therefore, the user can input the content in the second round of dialogue as "The 'yin' of music, not the 'yin' of impression", and the dialogue system can recognize that the intention type corresponding to the content is a rewrite intention, and generate a second slot value according to the content input in the second round and the first slot value. The second slot value can be "**Audio-Visual**", rewrite the "**Impression**" in the point of interest slot as "**Audio-Visual**", and the rewritten point of interest slot is used by the dialogue system to return a point of interest list 21 to the user in response to the second round of dialogue.
[0046] In this embodiment, for non-first-round voice navigation dialogues in a multi-round dialogue, the intent type corresponding to the voice navigation command input in the non-first-round voice navigation dialogue is recognized. When the intent type is the rewriting intent, a second slot value is regenerated by combining the current voice navigation command and the first slot value determined from the previous dialogue, and the original first slot value in the current point of interest slot is rewritten as the second slot value. The rewritten point of interest slot is used to filter out the corresponding target point of interest from the points of interest on the map in response to the current non-first-round voice navigation dialogue. It can be seen that in any non-first-round voice navigation dialogue in a multi-round dialogue, this solution recognizes the voice correction instruction input by the user and rewrites the original first slot value in the point of interest slot as the second slot value indicated by the voice correction instruction, so as to achieve the purpose of the user correcting the incorrect point of interest determined in the previous dialogue by voice, thereby solving the problem of the imperfect correction mechanism for points of interest in related dialogue methods in the voice navigation scenario.
[0047] In one embodiment, when recognizing the intent type corresponding to the voice navigation command, the voice of the voice navigation command can be converted into a voice navigation text, and the intent type corresponding to the voice navigation text is recognized by a first deep learning model.
[0048] In one embodiment, the first deep learning model can be a classification model such as BERT, Robustly Optimized BERT Approach (RoBERTa), or Text Convolutional Neural Network (TEXTCNN). This specification does not limit the model framework of the first deep learning model. Of course, in order to train the first deep learning model to have the ability to recognize the rewriting intent, this specification provides a first rewriting dataset for training the first deep learning model. Any sample in the first rewriting dataset includes a rewritten text and a first label, where the rewritten text is used to express the user's intent to rewrite the first slot value, and the first label is used to classify the intent type.
[0049] The first tag can be represented by 1 to i, where 1 represents the type of rewriting intention, and 2 to i represent other intention types, where i is a positive integer not less than 2. The rewritten text can be obtained by summarizing the sentence patterns followed by the rewritten phrases from all possible rewritten phrases used by the user, and designing corresponding rewritten texts for each sentence pattern. For example, some sentence patterns followed by the rewritten phrases may include: "is **not**", "is not **is**", "change **to**", "what I want to say is **, rather than**", "change to / be / replace with **", etc. Of course, other sentence patterns may also be included, and this specification does not impose any restrictions on this. In addition, multiple rewritten texts can be designed for the same sentence pattern. For example, the rewritten texts for the "is **not**" sentence pattern can be "is the 'yin' of music, not the 'yin' of impression", "is the 'ka' of coffee, not the 'kai' of flying an airplane", etc.
[0050] In one embodiment, when generating the rewritten text, a standard navigation data set can be obtained; any sample in the standard navigation data set includes a voice navigation text and a point of interest slot value extracted from the voice navigation text. For each point of interest slot value, homophones and / or word combinations associated with the characteristics of the point of interest slot value are generated, and based on at least one of the point of interest slot value, the homophones and word combinations corresponding to the point of interest slot value, and a preset sentence pattern, a rewritten text corresponding to the preset sentence pattern is generated.
[0051] Through the synthesis method of the rewritten text in this embodiment, for different preset sentence patterns, by combining with the homophones and word combinations associated with the characteristics of the generated point of interest slot values, a large number of rewritten texts can be quickly generated, thereby improving the generation efficiency and generalization of the rewritten text.
[0052] Figure 3 This specification shows a schematic diagram of generating a rewritten text according to an exemplary embodiment. As Figure 3As shown, in any sample in the standard navigation dataset, the voice navigation text can be "Navigate to ** Restaurant," and the corresponding POI slot value can be "** Restaurant." Alternatively, the voice navigation text can be "Navigate to ** Home," and the corresponding POI slot value can be "** Home." For each POI slot value in the standard navigation dataset, a pseudo POI slot value associated with the characteristics of the POI slot value can be generated by generating homophones and word groups for the POI slot value. The pseudo POI slot value can be generally understood as a word similar to the POI slot value, such as a homophone or synonym. For example, for the POI slot value "along the way," homophones corresponding to each character in the POI slot value can be generated. For example, homophones for "along" are "yan," "yan," and "yan," and homophones for "tu" are "tu," "tu," and "tu." Word groups similar to "along the way" can also be generated, such as "along the road," "along the street," "forehead," and "long distance." For example, for the interest point slot value "coffee", homophones corresponding to each character in the interest point slot value can be generated, for example, the homophones of "咖" can be "加", "家" and "嘉", etc., and the homophones of "咖啡" can be "非", "费" and "肥", etc. It is also possible to generate word groups similar to the "coffee", such as "咖啡", "Garfield", "飞", "异人" and "芳菲", etc. Based on the interest point slot value, the homophones and word groups corresponding to the interest point slot value, and the preset sentence pattern, a rewritten text corresponding to the preset sentence pattern is generated. For example, the preset sentence pattern is "replace ** with **", the interest point slot value is "市" and "政", and the corresponding word group is "电视", which is split into "电" and the homophone "视" corresponding to "政". By combining at least one of the above word groups and characters with different preset sentence patterns, rewritten texts with different speech techniques can be generated. For example, if the preset sentence pattern is "Rewrite ** to **", the corresponding rewritten text can be generated as "Rewrite the city in municipal affairs to the video in television". For example, if the preset sentence pattern is "I am talking about ** instead of **", the corresponding rewritten text can be generated as "I am talking about television instead of municipal affairs". For another example, if the preset sentence pattern is "I am talking about **", the corresponding rewritten text can be generated as "I am talking about the video in television".
[0053] In one embodiment, when identifying the intent type corresponding to the voice navigation command, the voice of the voice navigation command can be converted into voice navigation text, and the voice navigation text can be matched with a preset sentence pattern. If the match is successful, the intent type corresponding to the voice navigation command is identified as a rewrite intention. If the match fails, the intent type corresponding to the voice navigation text is identified through the first deep learning model.
[0054] In this embodiment, the method of identifying intent by rule matching is faster and more accurate, but it is difficult to process uncommon new sentence patterns. Although the deep learning model is less efficient and slower than the rule matching method, it has stronger generalization ability and can process uncommon new sentence patterns. By using the complementary method of rule matching and deep learning model in this embodiment, the speed and accuracy of intent identification can be comprehensively improved.
[0055] In one embodiment, the preset sentence patterns can be sentence patterns such as "is ** not **", "** is ** not **", "change ** to **", "the *th ** rewrite as **", "** is ** of **", "change to / be / replace with **", etc. Of course, there may also be other sentence patterns for rewriting words, and this specification does not limit this.
[0056] When matching the voice navigation text with the preset sentence patterns, the voice navigation text can be respectively matched with the regular expressions corresponding to each preset sentence pattern. If it matches successfully with any regular expression, the intent type corresponding to the voice navigation command is identified as the rewrite intent. For example, assuming the voice navigation text is "The character for'music' is 'yin', not the character for 'impression' which is 'yin'", it can be matched with the regular expression corresponding to "is ** not **". If the match is successful, the intent type corresponding to the voice navigation command is identified as the rewrite intent. If it fails to match with all the regular expressions corresponding to the preset sentence patterns, the intent type corresponding to the voice navigation text can be identified by the first deep learning model.
[0057] In one embodiment, when generating the second slot value according to the voice navigation command and the first slot value, the voice navigation text converted from the voice navigation command and the first slot value can be input into the second deep learning model, and the second slot value output by the second deep learning model can be obtained. Among them, the second deep learning model is trained by a second rewrite dataset. The input data of any sample in the second rewrite dataset includes the rewritten text and the interest point slot value before rewriting corresponding to the rewritten text, and the output result is the interest point slot value after rewriting.
[0058] Among them, the second deep learning model can be a large language model, and this specification does not impose any restrictions on the model framework of the second deep learning model. For the same pair of rewritten interest point slot values and the interest point slot values before rewriting, different rewritten texts can be set according to preset sentence patterns. For example, the input data can be "Wanshun Road, what I want to say is Mancao instead of Wannian", and the output data is "Manshun Road". Among them, "Wanshun Road" is the interest point slot value before rewriting, "What I want to say is Mancao instead of Wannian" is the rewritten text, and "Manshun Road" is the interest point slot value after rewriting. For example, the input data can also be "Wanshun Road, rewrite the first Wannian "Wan" to "Man" in "Mancao", and the output data is "Manshun Road". For example, the input data can also be "Wanshun Road, I am talking about Mancao instead of Wannian", and the output result is "Manshun Road".
[0059] In one embodiment, the slot value matching library may store the correspondence between each first slot value and each second slot value obtained from each user's previous use of the dialogue function, wherein the correspondence is associated with the user identifier. For the first round of voice navigation dialogue in a multi-round dialogue, the voice navigation command input in the first round of voice navigation dialogue is obtained. If the first slot value corresponding to the user identifier extracted from the voice navigation command is recorded in the slot value matching library, the second slot value corresponding to the first slot value is entered into the point of interest slot. The point of interest slot is used to filter the corresponding target point of interest from the points of interest on the map in response to the first round of voice navigation dialogue.
[0060] For example, suppose user A previously entered the voice navigation command "Navigate to Audio and Video City" into the dialogue system. The dialogue system extracts the first slot value from the input as "Impression City" and sets it to the POI slot. The dialogue system then filters target POIs related to "Impression City" from the POIs on the map. After the user discovers that the target POI is not the expected POI, they initiate another dialogue with "It's the sound of music, not the print of impression." The dialogue system changes the first slot value of the POI slot to the second slot value, "Audio and Video City," and filters target POIs related to "Audio and Video City" from the POIs on the map. At this point, the dialogue system can store the correspondence between "Impression City" and "Audio and Video City," along with the identifier of the user who issued the correction instruction.
[0061] When the same user A inputs the voice navigation command "Navigate to Audio-Visual City" in the dialogue system again, even if the first slot value extracted by the dialogue system from the input again is "Impression City", at this time, the first slot value corresponding to the user identifier of user A extracted from the voice navigation command can be recorded in the slot value matching library, and the second slot value corresponding to the first slot value can be filled in the point of interest slot, that is, "Audio-Visual City" instead of the extracted "Impression City" is filled in the point of interest slot.
[0062] In this embodiment, to address the problem that the text after speech conversion is not the target text expected by the user due to personalized pronunciation for different users or homophones, etc., by storing the first slot value before rewriting and the second slot value after rewriting for each user in the slot value matching library, when the user uses the dialogue system again next time, the rewritten second slot value can be directly used without the user having to perform "secondary correction" again, thereby reducing the number of interactions between the user and the dialogue system and being applicable to scenarios where the user frequently uses a certain point of interest.
[0063] Corresponding to the embodiments of the foregoing method, this specification also provides embodiments of an apparatus and a terminal to which the apparatus is applied.
[0064] Figure 4 It is a schematic structural diagram of an electronic device shown according to an exemplary embodiment of this specification. As Figure 4 shown, at the hardware level, the electronic device 400 includes a processor 402, an internal bus 404, a network interface 406, a memory 408, and a non-volatile memory 410. Of course, there may also be other hardware required for other services. One or more embodiments of this specification can be implemented in a software manner. For example, the processor 402 reads the corresponding computer program from the non-volatile memory 410 into the memory 408 and then runs it. Of course, in addition to the software implementation manner, one or more embodiments of this specification do not exclude other implementation manners, such as a logic device or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic module and can also be hardware or a logic device.
[0065] Figure 5 It is a block diagram of a device for determining a point of interest from a multi-round dialogue shown according to an exemplary embodiment of this specification. As Figure 5 shown, the device can be applied to the electronic device 400 shown as Figure 4 shown to implement the technical solution of this specification. The device includes:
[0066] An intent recognition module 502, configured to recognize an intent type corresponding to the voice navigation command in response to receiving a voice navigation command input in a non-first-round voice navigation dialogue;
[0067] A second slot value determination module 504, configured to, when the intent type is a rewrite intent, obtain the voice navigation command and the first slot value filled in the point of interest slot, and generate a second slot value according to the voice navigation command and the first slot value, where the first slot value is determined by a dialogue before the non-first-round voice navigation dialogue;
[0068] A rewriting module 506, configured to rewrite the first slot value in the point of interest slot into the second slot value, and the rewritten point of interest slot is used to filter out corresponding target points of interest from the points of interest on the map in response to the non-first-round voice navigation dialogue.
[0069] Optionally, the intent recognition module 502 is specifically configured to convert the voice of the voice navigation command into a voice navigation text; identify an intent type corresponding to the voice navigation text through a first deep learning model; or, match the voice navigation text with a preset sentence pattern. If the match is successful, the intent type corresponding to the voice navigation command is identified as a rewriting intent. If the match fails, the intent type corresponding to the voice navigation text is identified through the first deep learning model.
[0070] Optionally, the intent recognition module 502 is specifically configured to match the voice navigation text with regular expressions corresponding to respective preset sentence patterns.
[0071] Optionally, any sample in the first rewriting dataset of the first deep learning model includes a rewritten text and a first label; wherein, the rewritten text is used to express the user's intent to rewrite the first slot value, and the first label is used to classify the intent type.
[0072] Optionally, the apparatus further includes a rewritten text generation module 508, configured to obtain a standard navigation dataset; any sample in the standard navigation dataset includes a voice navigation text and a point of interest slot value extracted from the voice navigation text; for each point of interest slot value, homophones and / or word combinations associated with the characteristics of the point of interest slot value are generated, and a rewritten text corresponding to the preset sentence pattern is generated based on at least one of the point of interest slot value, the homophones and word combinations corresponding to the point of interest slot value, and the preset sentence pattern.
[0073] Optionally, the second slot value determination module 504 is specifically configured to input the voice navigation text converted from the voice navigation command and the first slot value into a second deep learning model, and obtain a second slot value output by the second deep learning model; wherein, the second deep learning model is trained through a second rewriting dataset, and the input data of any sample in the second rewriting dataset includes a rewritten text and a point of interest slot value before rewriting corresponding to the rewritten text, and the output result is a point of interest slot value after rewriting.
[0074] Optionally, a corresponding relationship between each first slot value and each second slot value obtained each time a user uses the dialogue function is stored in the slot value matching library, and the corresponding relationship is associated with the user identifier. The apparatus further includes a voice adaptation module 510, configured to, for the first-round voice navigation dialogue in a multi-round dialogue, obtain the voice navigation command input in the first-round voice navigation dialogue, and fill the second slot value corresponding to the first slot value into the point of interest slot when the first slot value corresponding to the user identifier extracted from the voice navigation command is recorded in the slot value matching library, where the point of interest slot is used to screen out a corresponding target point of interest from the points of interest on the map in response to the first-round voice navigation dialogue.
[0075] For the implementation processes of the functions and actions of each module in the above apparatus, refer to the implementation processes of the corresponding steps in the above method for details, which will not be elaborated here.
[0076] For the apparatus embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The apparatus embodiment described above is only illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution in this specification. A person of ordinary skill in the art can understand and implement it without creative efforts.
[0077] This specification also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of any one of the foregoing methods for determining a point of interest from a multi-round dialogue provided by this application are implemented.
[0078] Specifically, computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices (such as EPROM, EEPROM, and flash memory devices), magnetic disks (such as internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks.
[0079] This specification also provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of any one of the foregoing methods for determining a point of interest from a multi-round dialogue are implemented.
Claims
1. A method for determining points of interest from multi-round conversations, characterized in that, The method includes: In response to receiving a voice navigation command input in a non-first-round voice navigation dialogue, identifying an intent type corresponding to the voice navigation command; When the intent type is a rewriting intent, obtaining the voice navigation command and a first slot value filled in an interest point slot, and generating a second slot value according to the voice navigation command and the first slot value, where the first slot value is determined by a dialogue before the non-first-round voice navigation dialogue; Rewriting the first slot value in the interest point slot as the second slot value, and the rewritten interest point slot is used to filter out a corresponding target interest point from the interest points on the map in response to the non-first-round voice navigation dialogue.
2. The method according to claim 1, wherein The identifying the intent type corresponding to the voice navigation command includes: Converting the voice of the voice navigation command into voice navigation text; Identifying the intent type corresponding to the voice navigation text through a first deep learning model; or, matching the voice navigation text with a preset sentence pattern, if the match is successful, identifying the intent type corresponding to the voice navigation command as a rewriting intent, if the match fails, identifying the intent type corresponding to the voice navigation text through the first deep learning model.
3. The method according to claim 2, wherein The matching the voice navigation text with a preset sentence pattern includes: Matching the voice navigation text with regular expressions corresponding to each preset sentence pattern respectively.
4. The method according to claim 2, wherein Any sample in a first rewriting dataset for training the first deep learning model includes a rewritten text and a first label; wherein, the rewritten text is used to express the user's intent to rewrite the first slot value, and the first label is used to classify the intent type.
5. The method according to claim 4, wherein Generating the rewritten text includes: Obtaining a standard navigation dataset; any sample in the standard navigation dataset includes voice navigation text and an interest point slot value extracted from the voice navigation text; For each interest point slot value, generating a homophone and / or a compound word associated with the feature of the interest point slot value, and generating a rewritten text corresponding to the preset sentence pattern based on at least one of the interest point slot value, the homophone and the compound word corresponding to the interest point slot value, and the preset sentence pattern.
6. The method according to claim 1, wherein The generating the second slot value according to the voice navigation command and the first slot value includes: Inputting the voice navigation text converted from the voice navigation command and the first slot value into a second deep learning model, and obtaining the second slot value output by the second deep learning model; wherein, the second deep learning model is trained through a second rewriting dataset, and the input data of any sample in the second rewriting dataset includes a rewritten text and an interest point slot value before rewriting corresponding to the rewritten text, and the output result is an interest point slot value after rewriting.
7. The method according to any one of claims 1-6, characterized in that, The method further includes: Obtaining the voice navigation command input in the first-round voice navigation dialogue, and when a first slot value corresponding to a user identifier extracted from the voice navigation command is recorded in a slot value matching library, filling a second slot value corresponding to the first slot value into the interest point slot, and the interest point slot is used to filter out a corresponding target interest point from the interest points on the map in response to the first-round voice navigation dialogue. Among them, the slot value matching library stores the corresponding relationships between the respective first slot values and second slot values obtained each time each user uses the dialogue function, and the corresponding relationships are associated with user identifiers.
8. An apparatus for determining points of interest from multi-round conversations, characterized in that, The device includes: An intent recognition module, configured to recognize an intent type corresponding to the voice navigation command in response to receiving a voice navigation command for a non-first-round voice navigation dialogue input; A second slot value determination module, configured to, when the intent type is a rewriting intent, obtain the voice navigation command and the first slot value filled in the point of interest slot, and generate a second slot value according to the voice navigation command and the first slot value, where the first slot value is determined by a dialogue before the non-first-round voice navigation dialogue; A rewriting module, configured to rewrite the first slot value in the point of interest slot as the second slot value, and the rewritten point of interest slot is used to screen out corresponding target points of interest from the points of interest on the map in response to the non-first-round voice navigation dialogue.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the steps of the method according to any one of claims 1-7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the steps of the method according to any one of claims 1-7 are implemented.