Hybrid Multilingual Navigation Voice Command Processing Method, Apparatus, and Electronic Device

By determining the user's location to apply location-specific voice processing strategies and segmenting voice commands, the method addresses the challenge of mixed language navigation commands, ensuring accurate navigation intent understanding.

CN116092492BActive Publication Date: 2025-07-15IFLYTEK CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310085223.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-17
Publication Date
2025-07-15
Estimated Expiration
2043-01-17

AI Technical Summary

Technical Problem

When existing voice assistants process navigation voice commands in mixed multilinguals, they cannot accurately identify the user's navigation intentions, resulting in comprehension bias.

Method used

In the navigation scenario, the country or region where the user is currently located is determined in advance, and the electronic fence information is used to determine it, the voice command is cut into place name segments and non-place name segments, and the voice processing strategy matching the region is called for identification, and finally the navigation intention is understood based on the recognition results of the two.

Benefits of technology

Accurate recognition and understanding of mixed multilingual navigation voice commands is achieved, avoiding the cost of building pronunciation dictionary and model adjustments, and improving the reliability and efficiency of navigation intentions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116092492B_ABST
    Figure CN116092492B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, apparatus and electronic device for processing mixed multi-language navigation voice commands. First, it determines the country or region where the user is currently located in advance, cuts the user input voice command into a place name segment and a non-place name segment, calls a voice processing strategy matching the country or region where the user is located to recognize the voice command corresponding to the place name segment, and finally combines the recognition results of the place name segment and the non-place name segment to understand the navigation intention. By determining the current country or region where the user is located instead of conventional positioning, a voice processing strategy matching the local language can be determined in advance, and through the cutting of the input voice, targeted recognition processing can be performed on the mixed multi-language voice commands, so as to provide a more reliable and accurate navigation intention understanding result. The present invention does not need to cost to build a dictionary, nor does it need to make a large number of parameter adjustments to the existing model, and can more economically and efficiently process the situation where mixed languages appear in the navigation scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of voice interaction technology, and in particular, to a method and device for processing hybrid multi-language navigation voice commands and an electronic device. Background Art

[0002] Currently, intelligent voice technology has been fully popularized, and voice products are widely used in smart furniture, mobile devices, and vehicle-mounted fields. Among them, in the process of voice interaction in overseas navigation scenarios, there are often situations where the main language and the place name language are inconsistent. For example, in the case of "navigate to XX place name" in English + local language, such as "Navigate to (English) красную площадь (Russian)" and other such voice commands involving navigation in hybrid multi-languages. However, existing voice assistant products usually cannot accurately recognize voice commands in such scenarios, resulting in a deviation in understanding the navigation intent of users.

[0003] Currently, for the specific problems in the above specific scenarios, existing technologies mostly, in principle, build on the original pronunciation dictionary, comprehensively consider the pronunciation habits of several hybrid languages. This process also requires collecting a large amount of audio annotation data, and large-scale parameter adjustment is also required at the model training level.

[0004] However, due to the differences between languages, there are still limitations in building a multi-language pronunciation dictionary. Especially for the recognition support ability of data in the tens of millions of levels such as multiple countries and multiple place names, it does not meet the expectations. Summary of the Invention

[0005] In view of the above, the present invention aims to provide a method and device for processing hybrid multi-language navigation voice commands and an electronic device to solve the drawbacks of using the existing method of building a pronunciation dictionary and adjusting the model for hybrid multi-language navigation voice commands.

[0006] The technical solution adopted by the present invention is as follows:

[0007] In a first aspect, the present invention provides a method for processing hybrid multi-language navigation voice commands, which includes:

[0008] In a navigation scenario, pre-determine the country or region where the user is currently located;

[0009] Cut the voice command input by the user into a place name segment and a non-place name segment;

[0010] Call a voice processing strategy matching the country or region where the user is currently located to recognize the voice command corresponding to the place name segment;

[0011] Combine the recognition results of the place name segment and the non-place name segment to understand the navigation intent.

[0012] In at least one possible implementation manner, the step of pre-determining the country or region where the user is currently located includes: obtaining electronic fence information corresponding to national borders or regional boundaries from a navigation map.

[0013] In at least one possible implementation manner, the processing method further includes: updating the obtained electronic fence information by using the position information of the user; or dynamically adjusting the period for obtaining the electronic fence information according to the position information, moving speed information, and distance information from the electronic fence of the user.

[0014] In at least one possible implementation manner, the step of splitting the voice command input by the user into a place name segment and a non-place name segment includes:

[0015] Splitting is performed by means of semantic understanding of the speech recognition content, or by means of classifying the speech recognition content, or by using the differences in the audio dimension of the voice command.

[0016] In at least one possible implementation manner, after recognizing the voice command corresponding to the place name segment, it is determined whether the recognition result meets the preset confidence requirement. If not, the voice command corresponding to the place name segment is recognized by using the same speech processing algorithm as that of the non-place name segment.

[0017] In at least one possible implementation manner, after recognizing the voice command corresponding to the place name segment, when it is determined that the recognition result does not correspond to the place name of the current country or region, the recognition result is corrected and several place names close to the recognition result are provided for the user to confirm.

[0018] In at least one possible implementation manner, the processing method further includes: before splitting the place name segment and the non-place name segment, detecting whether the voice command input by the user contains a place name. If not, the input complete voice command is directly recognized and its intention is understood by using the speech processing strategy corresponding to the main language.

[0019] In a second aspect, the present invention provides a hybrid multi-language navigation voice command processing device, which includes:

[0020] A location determination module, configured to pre-determine the country or region where the user is currently located in a navigation scenario;

[0021] A command splitting module, configured to split the voice command input by the user into a place name segment and a non-place name segment;

[0022] A place name recognition module, configured to call a speech processing strategy matching the country or region where the user is currently located to recognize the voice command corresponding to the place name segment;

[0023] A navigation intention understanding module, configured to perform navigation intention understanding by combining the recognition results of the place name segment and the non-place name segment.

[0024] In a third aspect, the present invention provides an electronic device, which includes:

[0025] One or more processors, a memory, and one or more computer programs, where the memory may adopt a non-volatile storage medium, and the one or more computer programs are stored in the memory. The one or more computer programs include instructions that, when executed by the device, cause the device to execute the method as described in the first aspect or any possible implementation manner of the first aspect.

[0026] The main concept of the present invention is to pre-determine the country or region where the user is currently located in a navigation scenario. After the user inputs a voice command, the command is segmented into a place name segment and a non-place name segment, and a voice processing strategy matching the country or region where the user is currently located is called to recognize the voice command corresponding to the place name segment. Finally, navigation intention understanding is performed by combining the recognition results of the place name segment and the non-place name segment. By determining the current country or region instead of the conventional positioning idea, a voice processing strategy matching the local language can be pre-determined, and through the segmentation of the input voice, targeted recognition processing can be performed on the input mixed multi-language voice command, thereby providing a more reliable and accurate navigation intention understanding result. The present invention does not need to consume costs to build a dictionary, nor does it need to perform a large number of parameter adjustments on the existing model, and can more economically and efficiently handle the situation of mixed languages in the navigation scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described below in conjunction with the accompanying drawings, where:

[0028] Figure 1 is a flowchart of an embodiment of the method for processing mixed multi-language navigation voice commands provided by the present invention;

[0029] Figure 2 is a schematic diagram of an embodiment of the device for processing mixed multi-language navigation voice commands provided by the present invention;

[0030] Figure 3 is a schematic diagram of an embodiment of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0032] For the case of mixed non-primary language navigation voice commands in a navigation scenario, the present invention proposes the following embodiments of at least one method for processing mixed multi-language navigation voice commands, as Figure 1 shown, which may specifically include:

[0033] Step S1: In a navigation scenario, pre-determine the country or region where the user is currently located;

[0034] In actual operation, this process can be implemented based on various means. For example, using the coordinate information of the user's location, specifically, the longitude and latitude of the user's current location can be obtained through navigation technologies such as satellite positioning. This method is relatively conventional. Although the implementation barrier is low, there are also certain drawbacks. That is, single coordinate positioning may have deviations, or in real application scenarios, it is difficult for a moving user to obtain its stable coordinate information reliably. Especially during vehicle driving, the user's moving speed is relatively high. If combined with areas where different countries are densely distributed, such as in Europe, it is difficult to reliably and accurately determine the country or region where the user is currently located.

[0035] Based on such considerations, in some preferred embodiments of the present invention, it is proposed to preferably use the electronic fence information corresponding to the national border or regional boundary in the navigation map to perform weighted excitation for subsequent mixed language processing. That is, in the concept of the preferred embodiment, it is to determine the country or region where the user is currently located based on the concept of "range". This can more reliably meet the requirements of the navigation scenario compared to single coordinate positioning, especially for specific scenarios such as driving or crossing national borders / regional boundaries.

[0036] Generating an electronic fence representing national or regional boundaries on a map and obtaining the above-mentioned electronic fence information from the map are both existing mature technologies, and the present invention will not elaborate or limit them here. Additionally, it can be further explained that based on the concept of determining the user's country or region by "range" mentioned above, in some preferred embodiments of the present invention, the obtained electronic fence information can be updated in combination with the user's location information (which can be from the aforementioned longitude and latitude coordinates). Of course, based on this, the cycle of obtaining the electronic fence information can be dynamically adjusted by further combining the user's moving speed information and the distance information from the electronic fence, so as to achieve reliable and accurate switching of the country / region where the user is located in some scenarios, such as cross-border driving. For example, when it is determined based on the above information that the user is getting closer to a national border, the electronic fence information can be obtained once every 15 seconds (which can be adjusted as needed).

[0037] Step S2: Cut the voice command input by the user into a place name segment and a non-place name segment;

[0038] Those skilled in the art can at least understand two aspects. Firstly, the scenario targeted by the present invention is navigation. Correspondingly, the user voice received by the system is related to navigation and is relatively likely to include voice commands such as querying navigation for a destination. Secondly, generally speaking, the default language of the navigation system must be the language commonly used by the user (such as the mother tongue), which the present invention calls the main language, indicating that this language is the most important basic language for the navigation to perform voice processing.

[0039] Based on the above two aspects, in actual operation, when receiving the navigation-related voice provided by the user, through the default main language, the input voice can be recognized and transcribed, and according to the established semantic understanding strategy, the input voice command can be split into a place name segment and a non-place name segment. Since this is a mature technology, the present invention only makes a schematic introduction: for example, after the audio of the input voice is preprocessed, it is transcribed through a voice recognition model of a certain main language, and the word segmentation and sentence separation processing are completed. Then, in combination with the intention understanding model, the possible place name parts and non-place name parts in the original input command are divided according to the preset expression templates, semantic slots, etc.; in addition, a method of classifying the voice recognition content can also be considered. During implementation, a classifier trained according to the established target can be used to classify the voice recognition content, so as to obtain the place name part and the non-place name part therein; in some better embodiments of the present invention, also considering the above-mentioned first and second aspects, preferably, the audio parts different from the main language can be directly determined as the place name segment through the differences in the audio dimension. This method can significantly improve the overall processing efficiency, and there are also mature voice and acoustic processing solutions in the art for reference. No matter which of the above cutting methods is adopted, the technical tools themselves are not the focus of the present invention, so no further elaboration or limitation is made.

[0040] Step S3: Call the speech processing strategy matching the country or region where the user is currently located to recognize the speech command corresponding to the place name segment;

[0041] Step S4: Understand the navigation intention by combining the recognition results of the place name segment and the non-place name segment.

[0042] Regarding the above two steps, it can be pointed out here that the recognition of the non-place name segment, as described above, can be obtained through the speech processing model related to the main language default in the navigation system, and in different implementation scenarios, its recognition timing is not limited; for the recognition of the place name segment, the speech processing model determined for the current international / region through the previous steps is preferably used for recognition processing. Then, considering the recognition results of both, they can be input into the mature natural language processing algorithm in the field to understand the navigation expectation input by the user and perform corresponding conventional navigation operations such as path planning. The present invention will not elaborate on this.

[0043] Furthermore, it can be further explained that in some preferred embodiments of the present invention, considering the special circumstances in the real scenario, if the speech processing model of the local country / region is directly used, reliable place name recognition results may not be obtained. Here, the following preferred solutions are proposed to address this.

[0044] (1) After using the speech processing strategy matching the local international / region in step S3 to perform place name segment recognition processing, determine whether the recognition result meets the preset confidence requirement (in actual operation, it can be compared according to a score mechanism). If it meets, continue to execute step S4; if it does not meet, use the same speech processing algorithm as the non-place name segment to recognize the speech command corresponding to the place name segment, and then execute step S4. That is to say, in this embodiment, from the perspective of the system, if the speech processing model of the local country / region cannot obtain the place name segment speech with a confidence that meets the requirements, it is considered that this place name segment speech is probably still in the main language. Therefore, the speech processing model for processing the non-place name segment is called to recognize this place name segment.

[0045] (2) After the recognition process of the place name segment is performed using the voice processing strategy matched with the local international / region in step S3, if it is determined that the recognition result cannot correspond to the place name of the country / region where it is located, it is preferred to correct the recognition result and provide one (the closest to the recognition result) local place name or multiple local place names (the top N closest to the recognition result) for the user to confirm, and then step S4 is executed. The corrected place name provided here can be in the form of voice and / or text output, and the method of correcting the recognition result involved in this embodiment can refer to the existing voice and text correction methods. That is to say, this embodiment of the system considers that if the recognition result of the place name segment obtained by the voice processing model of the local country / region may not be accurately matched to the corresponding local place name due to the user's inaccurate pronunciation of the local language, then a strategy of assisting the user in correction is adopted. In addition, the idea of matching the recognition result with the local place name is mentioned here, that is to say, in some other embodiments of the present invention, it is possible but not limited to constructing place name databases of different countries / regions in the local language in the early stage, including the addresses of cities, scenic spots, roads, buildings, etc. When the recognition result of the place name segment is obtained through the voice processing model of the local country / region, it can be matched with this place name database.

[0046] Finally, it can also be added that when step S2 was described above, it was mentioned that in the navigation scenario, most of the voice commands input by the user are related to addresses. However, the present invention takes into account that in actual applications, there are also navigation voice commands that do not input place names. In this regard, in some other preferred solutions, the present invention proposes that before the segmentation of the place name segment and the non-place name segment, it can be first detected whether the voice command input by the user contains a place name. For example, through conventional and mature technologies such as language recognition, understanding, and classification, it can be preliminarily determined whether the voice command contains place name content. If the preliminary judgment result is negative, the voice processing strategy corresponding to the default main language of the system can be directly used to recognize and understand the complete voice command input, and the link of segmentation and calling the voice processing algorithm of the country / region where it is located can be skipped, so that the user's needs can be responded to faster for such situations.

[0047] Based on the above-mentioned various embodiments and combined with the examples mentioned earlier, the following reference is provided to assist in explaining the working process of the solution of the present invention: A user whose native language is English is driving locally in Russia. When using a navigation application, the navigation system with the default main language being English will determine that the current user is located within Russia based on the electronic fence representing the Russian national border provided by the high-precision map module in the current navigation. (The core goal of this process is to serve the subsequent corresponding language recognition model. In the specific implementation of the solution, at this time, the Russian recognition model can be pre-selected for subsequent invocation). When the user inputs the mixed-language voice interaction instruction "Navigate to (English, meaning navigate to) красную площадь (Russian, meaning Red Square)", the navigation system identifies the English part and the non-English part at the acoustic level, and according to the established strategy, determines the English part as a non-place-name segment and the non-English part as a place-name segment.

[0048] Next, based on the fact that it has been previously determined by the electronic fence that the current user is located in Russia, the constructed Russian speech recognition model is called to perform speech recognition processing on the audio segment corresponding to the place-name segment "красную площадь". Of course, it can be understood that for the non-place-name segment part, the default language of the navigation system, that is, the English speech recognition model, can be directly used for recognition processing.

[0049] Finally, according to the audio time sequence of the input voice, the recognition texts output by the above two speech recognition models are spliced and sent into the pre-constructed natural language understanding algorithm for text-based semantic analysis processing. Here, it can be supplemented that the process of natural language understanding can be independent of the language type. In actual operation, different language datasets only need to be collected according to real needs to train the existing intention understanding model to obtain the intention understanding result; and for the purpose of simplifying the training of the natural language understanding algorithm, the previously recognized text "Navigate to красную площадь" can also be first converted into the text "Navigate to Red Square" using the preset Russian-to-English translation service, and then sent into the intention understanding model that has only been trained with English samples. Thus, through the above entire process, the navigation system can determine that the user expects to drive from the current location to Red Square in Moscow, and then give the corresponding planned route and other relevant navigation information.

[0050] In summary, the main idea of the present invention is to pre-determine the country or region where the user is currently located in the navigation scenario. After the user inputs a voice command, the command is segmented into a place name segment and a non-place name segment. A voice processing strategy matching the country or region where the user is currently located is called to recognize the voice command corresponding to the place name segment. Finally, the navigation intention is understood by combining the recognition results of the place name segment and the non-place name segment. By determining the current country or region instead of the conventional positioning idea, a voice processing strategy matching the local language can be pre-determined, and through the segmentation of the input voice, targeted recognition processing can be performed on the input mixed multi-language voice command, thereby providing a more reliable and accurate navigation intention understanding result. The present invention does not need to consume costs to construct a dictionary for voice recognition, nor does it need to perform a large number of parameter adjustments on the existing voice recognition model, and can more economically and efficiently process the situation where mixed languages appear in the navigation scenario.

[0051] Corresponding to the above embodiments and preferred solutions, the present invention also provides an embodiment of a mixed multi-language navigation voice command processing device, as Figure 2 shown, which may specifically include the following components:

[0052] The location area determination module 1 is used to pre-determine the country or region where the user is currently located in the navigation scenario;

[0053] The command segmentation module 2 is used to segment the voice command input by the user into a place name segment and a non-place name segment;

[0054] The place name recognition module 3 is used to call a voice processing strategy matching the country or region where the user is currently located to recognize the voice command corresponding to the place name segment;

[0055] The navigation intention understanding module 4 is used to understand the navigation intention by combining the recognition results of the place name segment and the non-place name segment.

[0056] It should be understood that the above Figure 2The division of each component in the shown hybrid multilingual navigation voice command processing device is only a division of logical functions. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these components can all be implemented in the form of software called by processing elements; they can also all be implemented in the form of hardware; or some components can be implemented in the form of software called by processing elements, and some components can be implemented in the form of hardware. For example, a certain above-mentioned module can be a separately established processing element, or can be integrated in a certain chip of an electronic device. The implementation of other components is similar. In addition, all or part of these components can be integrated together or can be independently implemented. In the implementation process, each step of the above method or each of the above components can be completed through the integrated logic circuit of the hardware in the processor element or the instructions in the form of software.

[0057] For example, the above components can be one or more integrated circuits configured to implement the above method, such as: one or more application specific integrated circuits (hereinafter referred to as: ASIC), or, one or more microprocessors (hereinafter referred to as: DSP), or, one or more field programmable gate arrays (hereinafter referred to as: FPGA), etc. Again, these components can be integrated together and implemented in the form of a system-on-a-chip (hereinafter referred to as: SOC).

[0058] Based on the above embodiments and their preferred solutions, those skilled in the art can understand that in actual operation, the technical concept involved in the present invention can be applied to various implementation manners. The present invention uses the following carriers as illustrative explanations:

[0059] (1) An electronic device. Specifically, the device may include: one or more processors, a memory, and one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions that, when executed by the device, cause the device to execute the steps / functions of the foregoing embodiments or equivalent implementation manners.

[0060] The electronic device can specifically be an electronic device related to a computer, such as but not limited to various interactive terminals, electronic products, mobile terminals, etc.

[0061] Figure 3A schematic structural diagram of an embodiment of the electronic device provided by the present invention. Specifically, the electronic device 900 includes a processor 910 and a memory 930. Among them, the processor 910 and the memory 930 can communicate with each other through an internal connection path to transmit control and / or data signals. The memory 930 is used to store a computer program, and the processor 910 is used to call and run the computer program from the memory 930. The above-mentioned processor 910 and the memory 930 can be integrated into a processing device, and more commonly, they are independent components. The processor 910 is used to execute the program code stored in the memory 930 to implement the above functions. Specifically, the memory 930 can also be integrated in the processor 910, or independent of the processor 910.

[0062] In addition, in order to make the functions of the electronic device 900 more complete, the device 900 may further include one or more of an input unit 960, a display unit 970, an audio circuit 980, a camera 990, and a sensor 901, etc. The audio circuit may further include a speaker 982, a microphone 984, etc. Among them, the display unit 970 may include a display screen.

[0063] Furthermore, the above-mentioned device 900 may further include a power supply 950 for supplying electrical energy to various components or circuits in the device 900.

[0064] It should be understood that for the operations and / or functions of each component in the device 900, reference can be specifically made to the descriptions of the embodiments of the method, system, etc. in the foregoing text. To avoid repetition, the detailed descriptions are appropriately omitted here.

[0065] It should be understood, Figure 3 The processor 910 in the shown electronic device 900 may be a system-on-chip (SOC). The processor 910 may include a central processing unit (Central Processing Unit; hereinafter referred to as: CPU), and may further include other types of processors, such as: a graphics processing unit (Graphics Processing Unit; hereinafter referred to as: GPU), etc., which will be specifically introduced later.

[0066] In summary, the various processors or processing units inside the processor 910 can cooperate with each other to implement the previous method process, and the corresponding software programs of each part of the processor or processing unit can be stored in the memory 930.

[0067] (2) A computer data storage medium, on which a computer program or the above-mentioned device is stored. When the computer program or the above-mentioned device is executed, the computer is made to execute the steps / functions of the foregoing embodiments or equivalent embodiments.

[0068] In several embodiments provided by the present invention, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer data storage medium. Based on such an understanding, some technical solutions of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product as described below.

[0069] In particular, it should be pointed out that the storage medium may refer to a server or a similar computer device. Specifically, that is, the computer program or the above-mentioned device is stored in a storage device in the server or a similar computer device.

[0070] (3) A computer program product (the product may include the above-mentioned device). When the computer program product runs on a terminal device, it causes the terminal device to execute the hybrid multilingual navigation voice instruction processing method of the foregoing embodiment or an equivalent implementation manner.

[0071] From the description of the above embodiments, those skilled in the art can clearly understand that all or part of the steps in the above implementation methods can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the above computer program product may include, but is not limited to, an APP.

[0072] Continuing from the foregoing, the above device / terminal may be a computer device, and the hardware structure of the computer device may specifically further include: at least one processor, at least one communication interface, at least one memory, and at least one communication bus; the processor, communication interface, and memory can all complete communication with each other through the communication bus. Among them, the processor may be a central processing unit CPU, DSP, microcontroller, or digital signal processor, and may also include a GPU, an embedded neural network processor (Neural-network Process Units; hereinafter referred to as: NPU), and an image signal processor (Image Signal Processing; hereinafter referred to as: ISP). The processor may also include an application-specific integrated circuit ASIC, or one or more integrated circuits configured to implement the embodiments of the present invention. In addition, the processor may have the function of operating one or more software programs, and the software programs may be stored in a storage medium such as a memory; and the foregoing memory / storage medium may include: non-volatile memory, such as a non-removable disk, a USB flash drive, a mobile hard disk, an optical disc, etc., as well as a read-only memory (hereinafter referred to as: ROM), a random access memory (hereinafter referred to as: RAM), etc.

[0073] In the embodiments of the present invention, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent the situations of A existing alone, A and B existing simultaneously, and B existing alone. Wherein A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one of the following" and its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, and c can represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, and c can be single or multiple.

[0074] Those skilled in the art can realize that the various modules, units, and method steps described in the embodiments disclosed in this specification can be implemented by a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0075] Moreover, the modules, units, etc. described as separate components may or may not be physically separated, that is, they can be located in one place, or they can be distributed to multiple places, such as the nodes of a system network. Specifically, some or all of the modules and units can be selected according to actual needs to achieve the purpose of the above-mentioned embodiment solutions. Those skilled in the art can understand and implement it without creative labor.

[0076] The above has detailedly described the structure, features, and effects of the present invention according to the embodiments shown in the drawings. However, the above are only the preferred embodiments of the present invention. It should be noted that for the technical features involved in the above embodiments and their preferred modes, those skilled in the art can reasonably combine and match them into a variety of equivalent solutions without departing from and without changing the design concept and technical effects of the present invention; Therefore, the present invention is not limited by the scope shown in the drawings. As long as the changes made in accordance with the concept of the present invention, or the equivalent embodiments modified into equivalent changes, still do not exceed the spirit covered by the specification and the drawings, they should all be within the protection scope of the present invention.

Claims

1. A method for processing mixed multilingual navigation voice commands, characterized in that, including: In a navigation scenario, pre-determine the country or region where the user is currently located, and select a speech processing strategy for the currently located country or region for subsequent invocation; Cut the speech command input by the user into a place name segment and a non-place name segment, including cutting by means of semantic understanding of the speech recognition content; Invoke a speech processing strategy that matches the country or region where the user is currently located, recognize the speech command corresponding to the place name segment, and determine whether the recognition result meets the preset confidence requirement. If not, use the same speech processing algorithm as the non-place name segment to recognize the speech command corresponding to the place name segment; Understand the navigation intention by combining the recognition results of the place name segment and the non-place name segment.

2. The hybrid multilingual navigation voice command processing method according to claim 1, wherein The pre-determining the country or region where the user is currently located includes: obtaining electronic fence information corresponding to the national border or regional boundary from the navigation map.

3. The hybrid multilingual navigation voice command processing method according to claim 2, wherein The processing method further includes: updating the obtained electronic fence information using the user's location information; or dynamically adjusting the period of obtaining the electronic fence information according to the user's location information, moving speed information, and distance information from the electronic fence.

4. The hybrid multilingual navigation voice command processing method according to claim 1, wherein After recognizing the speech command corresponding to the place name segment, when it is determined that the recognition result cannot correspond to the place name of the current country or region due to the user's inaccurate pronunciation of the local language, correct the recognition result and provide several place names close to the recognition result for the user to confirm.

5. The method for processing hybrid multilingual navigation voice commands according to any one of claims 1 to 4, characterized in that, The processing method further includes: before cutting the place name segment and the non-place name segment, detecting whether the speech command input by the user contains a place name. If not, directly use the speech processing strategy corresponding to the main language to recognize and understand the intention of the input complete speech command.

6. A hybrid multilingual navigation voice command processing device, characterized in that, including: A location determination module, configured to pre-determine the country or region where the user is currently located in a navigation scenario, and select a speech processing strategy for the currently located country or region for subsequent invocation; An instruction cutting module, configured to cut the speech command input by the user into a place name segment and a non-place name segment, including cutting by means of semantic understanding of the speech recognition content; A place name recognition module, configured to invoke a speech processing strategy that matches the country or region where the user is currently located, recognize the speech command corresponding to the place name segment, and determine whether the recognition result meets the preset confidence requirement. If not, use the same speech processing algorithm as the non-place name segment to recognize the speech command corresponding to the place name segment; A navigation intention understanding module, configured to understand the navigation intention by combining the recognition results of the place name segment and the non-place name segment.

7. An electronic device, characterized in that, including: One or more processors, a memory, and one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions that, when executed by the electronic device, cause the electronic device to execute the hybrid multilingual navigation speech command processing method according to any one of claims 1 to 5.

8. A computer data storage medium, characterized in that, The computer data storage medium stores a computer program, which, when running on a computer, causes the computer to execute the hybrid multi-language navigation voice instruction processing method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Apparatus and method for recognizing voice and text

    CN104282302A

  • Novel shared bicycle management method and system

    CN107093277A

  • Voice recognition device and voice recognition method

    CN107112007A