Navigation voice online generation method and device, electronic equipment and storage medium

By identifying the category of the navigation voice broadcast text and segmenting it into short sentences, and using the server to generate navigation voice, the problems of excessively long navigation voice generation time and unnatural sound have been solved. This has enabled fast and high-quality navigation voice broadcast, improving user experience and navigation accuracy.

CN121506086APending Publication Date: 2026-02-10BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511342759.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Traditional navigation voice broadcast systems suffer from problems such as excessively long generation times and unnatural voices, which may cause users to miss key information in complex traffic environments.

Method used

By identifying the sentence categories in the navigation voice broadcast text, a targeted text segmentation strategy is used to break long sentences into shorter ones. These shorter sentences are then sent to the server for voice generation, utilizing the server's high-performance processing capabilities to generate high-quality navigation voice.

Benefits of technology

It enables rapid generation and high-quality output of navigation voice, improves the naturalness of the voice and the user's auditory experience, and ensures the accuracy of navigation information and driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121506086A_ABST
    Figure CN121506086A_ABST
Patent Text Reader

Abstract

The invention provides a navigation voice online generation method and device, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence, in particular to the technical fields of natural language processing, text-to-voice technology, voice broadcasting and the like. The method comprises the following steps: acquiring a broadcast text of navigation voice to be generated, wherein the broadcast text comprises at least one whole sentence; identifying the category of the whole sentence in the broadcast text, and adopting a text segmentation strategy corresponding to the category according to the category of the whole sentence to obtain segmented short sentences corresponding to the whole sentence; and sending the segmented short sentences corresponding to the whole sentence to a server, and generating navigation voice corresponding to the whole sentence through the server. According to the scheme disclosed by the invention, the generation speed of the navigation voice is accelerated, the requirement of real-time navigation is met, and meanwhile, the naturalness of the voice and the auditory experience of the user are also improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of artificial intelligence, in particular to the technical field of natural language processing, text-to-speech technology, and voice broadcasting, and more particularly to a navigation voice online generation method and device, an electronic device, and a storage medium. BACKGROUND

[0002] In modern navigation systems, providing accurate and timely voice broadcasting is crucial for improving driving experience and safety. With the development of technology, users have increasingly high requirements for the naturalness, fluency, and real-time performance of navigation voice. However, traditional navigation voice broadcasting systems often have some limitations, such as limited acoustic model size, inability to highly restore tone color, excessive performance consumption, and long generation time of large model voice. These problems make it difficult for navigation voice broadcasting to meet user needs in terms of real-time performance and voice quality, especially in complex traffic environments, which may cause users to miss critical navigation information. SUMMARY

[0003] The present disclosure provides a navigation voice online generation method, device, electronic device, and storage medium.

[0004] According to an aspect of the present disclosure, a navigation voice online generation method is provided, which includes:

[0005] obtaining a broadcast text of a navigation voice to be generated, the broadcast text containing at least one complete sentence;

[0006] identifying the category of the complete sentence in the broadcast text, and adopting a text segmentation strategy corresponding to the category of the complete sentence according to the category of the complete sentence to obtain a segmented short sentence corresponding to the complete sentence;

[0007] sending the segmented short sentence corresponding to the complete sentence to a server to generate a navigation voice corresponding to the complete sentence through the server.

[0008] According to another aspect of the present disclosure, a navigation voice online generation device is provided, which includes:

[0009] an obtaining module configured to obtain a broadcast text of a navigation voice to be generated, the broadcast text containing at least one complete sentence;

[0010] a segmentation module configured to identify the category of the complete sentence in the broadcast text, and adopt a text segmentation strategy corresponding to the category of the complete sentence according to the category of the complete sentence to obtain a segmented short sentence corresponding to the complete sentence;

[0011] a sending module configured to send the segmented short sentence corresponding to the complete sentence to a server to generate a navigation voice corresponding to the complete sentence through the server.

[0012] According to a third aspect of the present disclosure, an electronic device is provided, comprising:

[0013] at least one processor; and

[0014] a memory connected with the at least one processor in communication; wherein,

[0015] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of the above technical solutions.

[0016] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to perform the method of any one of the above technical solutions.

[0017] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method of any one of the above technical solutions.

[0018] The present disclosure provides a navigation voice online generation method, device, equipment and storage medium. The present disclosure acquires navigation voice broadcast text containing multiple complete sentences, intelligently identifies the category to which each complete sentence belongs, and adopts a targeted text segmentation strategy to effectively segment long sentences into shorter and easier-to-process short sentences. This fine processing method not only optimizes the text processing process and ensures that each short sentence can be processed in the most suitable way for its content, but also sends these segmented short sentences to the server, which can make the server generate navigation voice more efficiently. Such a method not only ensures that the online generated voice time is less than 1 second, but also meets the audio broadcast effect according to the sentence semantics and punctuation, thereby meeting the user's auditory experience expectations. This not only speeds up the generation of navigation voice and meets the real-time navigation demand, but also improves the naturalness of the voice and the auditory experience of the user, making the navigation voice clearer and more accurate, which helps the driver better understand and follow the navigation instructions, thereby improving the accuracy of navigation and the safety of driving.

[0019] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0020] The accompanying drawings are used to better understand the present scheme and do not limit the present disclosure. Among them:

[0021] Figure 1is an exemplary navigation voice generation flowchart of one embodiment;

[0022] Figure 2 is an exemplary navigation voice generation flowchart of another embodiment;

[0023] Figure 3 is a step diagram of a navigation voice online generation method in an embodiment of the present disclosure;

[0024] Figure 4 is a whole diagram of a navigation voice online generation method in an embodiment of the present disclosure;

[0025] Figure 5 is a principle block diagram of a navigation voice online generation device in an embodiment of the present disclosure;

[0026] Figure 6 is a block diagram of an electronic device for implementing a navigation voice online generation method in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0027] Exemplary embodiments of the present disclosure are described herein with reference to the accompanying drawings, which are included to provide a thorough understanding of embodiments of the present disclosure by a person of ordinary skill in the art, and should be considered in connection with the following detailed description, and should be considered as merely exemplary. Accordingly, those of ordinary skill in the art will recognize that modifications and other implementations are possible without departing from the scope and spirit of the present disclosure, and the implementations included herein are intended to cover all such modifications and equivalents. Also, in the following description, descriptions of well-known functions and constructions are omitted for clarity and conciseness.

[0028] Referring to Figure 1 , Figure 1 is an exemplary navigation voice generation flowchart of one embodiment. The generation manner of the navigation voice is to generate and broadcast a complete sentence directly online during navigation. This method can generate complete and coherent voice instructions, and provide a natural auditory experience. However, this method has a problem of long generation time, which exceeds 1 second. In navigation broadcasting, 1 second is the time limit that can be perceived by a user. If the time exceeds this limit, the user may feel a delay, thereby affecting the real-time performance of navigation and the user experience.

[0029] Referring to Figure 2 , Figure 2 is an exemplary navigation voice generation flowchart of another embodiment. The generation manner of the navigation voice is to cut a complete sentence script, cut it into short sentences according to punctuation, and shorten the length of the script. The advantage of this method is that it can quickly generate voice, and meet the strict requirements of navigation broadcasting on time. However, unlimited cutting into short sentences will result in incoherent voice broadcasting, that is, there is a pause in the broadcasting effect, which damages the auditory experience of the user and makes it difficult for the user to understand the navigation information smoothly.

[0030] In summary, the technical problems existing in the prior art mainly focus on two aspects: one is that the generation time is too long when generating a complete sentence voice online, and the other is that the effect is unnatural when cutting the script into short sentences for broadcasting, which affects the user experience.

[0031] To solve the above technical problems, the present disclosure provides a navigation voice online generation method, as shown in Figure 3 Figure 3 is a step schematic diagram of the navigation voice online generation method in the embodiments of the present disclosure, the method is applied to a client, and the method comprises:

[0032] Step S301, obtaining a broadcast text of a navigation voice to be generated, the broadcast text containing at least one complete sentence.

[0033] Specifically, the "broadcast text" refers to the original text information used to generate the navigation voice, which usually contains a series of navigation instructions or prompts, such as "turn right in 500 meters ahead", "go straight at the next intersection", etc. "Multiple complete sentences" means that the broadcast text is composed of several complete sentences, each of which conveys one or more specific navigation information. When obtaining the broadcast text of the navigation voice to be generated, the broadcast text is usually created by the navigation system through integrating the destination information input by the user, the route data in the map database and the real-time traffic update, using a preset text template or an automatic generation algorithm, and then issued to the client by the navigation system, so that the broadcast text can be obtained.

[0034] Step S302, identifying the category of the complete sentence in the broadcast text, and adopting a text segmentation strategy corresponding to the category of the complete sentence according to the category of the complete sentence to obtain a segmented short sentence corresponding to the complete sentence.

[0035] Specifically, the "broadcast text" refers to a series of text instructions used to generate the navigation voice, and the "category of the complete sentence" refers to the type of these text instructions classified according to their content. The category can include with variable and without variable. The implementation process of this scheme includes: first, analyzing each complete sentence in these broadcast texts to determine which category they belong to. This identification process can be realized by using natural language processing (NLP) technology through a preset classification rule or a machine learning model. Once the category of each complete sentence is determined, a text segmentation strategy matching the category will be adopted. The text segmentation strategy refers to the method of decomposing long sentences into shorter, more understandable and manageable sentence fragments, which will vary according to the category of the sentence to maintain the integrity and accuracy of the information. Among them, the text segmentation strategy is created in advance, and different text segmentation strategies are adopted for different categories of complete sentences.

[0036] ​In this way, through targeted text segmentation, segmented short sentences corresponding to the original complete sentences can be generated, which are not only easier to quickly generate speech, but also can ensure the clarity of information and the user's auditory experience when broadcasting. Such a processing method helps to improve the quality and real-time performance of navigation speech, while meeting the user's demand for accuracy and easy understanding of navigation information.

[0037] In step S303, the segmented short sentences corresponding to the complete sentences are sent to the server, and the navigation speech corresponding to the complete sentences is generated by the server.

[0038] Specifically, "segmented short sentences" refer to short text fragments processed by text segmentation strategies, which are extracted from the original broadcast text, and each fragment contains navigation information that is easy to understand and process. "Server" here refers to a remote computing server that has powerful text-to-speech (TTS) processing capabilities. The specific implementation process of this scheme includes: after obtaining these segmented short sentences, the segmented short sentences are sent to the server through the network. After the server receives these short sentences, it will use advanced speech synthesis technology to convert text into natural and fluent speech output.

[0039] In the process of converting text to speech by the server, this process usually includes several key steps: first, the server analyzes the received text to understand its semantics and context; second, according to the text content and preset voice parameters, the server selects the appropriate voice model for speech synthesis; then, the server generates a speech file, which may involve adjusting the speech speed, tone, and volume, etc. to ensure the naturalness and audibility of the speech; finally, the generated speech file is sent back to the client, i.e. the user's navigation device.

[0040] In this way, even on resource-limited client devices, high-quality navigation speech broadcasting can be achieved. This method not only reduces the computational burden of the client, but also takes advantage of the high-performance processing capabilities of the server, thereby ensuring the rapid generation and high-quality output of navigation speech and improving the user's navigation experience.

[0041] The present disclosure provides a navigation voice online generation method and device, electronic equipment and storage medium. The present disclosure obtains a navigation voice broadcast text containing multiple complete sentences, intelligently identifies the category to which each complete sentence belongs, and effectively divides long sentences into shorter and easier-to-process short sentences using targeted text segmentation strategies. This refined processing method not only optimizes the text processing process and ensures that each short sentence can be processed in the most suitable way for its content, but also sends the segmented short sentences to the server, which can make the server generate navigation voice more efficiently. Such a method can not only ensure that the online generated voice time is less than 1 second, but also meet the audio broadcast effect according to the sentence semantics and punctuation, thereby meeting the user's auditory experience expectations. This not only speeds up the generation of navigation voice and meets the real-time navigation demand, but also improves the naturalness of the voice and the auditory experience of the user, making the navigation voice clearer and more accurate, which helps the driver better understand and follow the navigation instructions, thereby improving the accuracy of navigation and the safety of driving.

[0042] In some optional embodiments, identifying the category to which each complete sentence in the broadcast text belongs includes:

[0043] If the broadcast text contains elements that change according to actual navigation conditions, the category of the complete sentence is a variable complete sentence.

[0044] If the broadcast text does not contain elements that change according to actual navigation conditions, the category of the complete sentence is a non-variable complete sentence.

[0045] Specifically, the "variable complete sentence" refers to a complete sentence in the broadcast text that contains elements that change according to actual navigation conditions (such as the actual position of the user, traffic conditions, estimated arrival time, etc.). For example, "500" in "turn left 500 meters ahead" is a variable that changes dynamically according to actual conditions. In contrast, the "non-variable complete sentence" refers to a complete sentence in the broadcast text that does not contain any elements that change according to navigation conditions. These sentences usually contain fixed information, such as "You have arrived at the destination".

[0046] The specific implementation process of this scheme includes: after obtaining the broadcast text, analyzing the broadcast text to identify whether each complete sentence in the text belongs to a "variable complete sentence" or a "non-variable complete sentence". This is usually done through natural language processing (NLP) technology, which specifically checks whether each sentence contains variable elements to determine the category of each complete sentence.

[0047] In this way, by identifying whether the broadcast text contains elements that change according to actual navigation conditions, and classifying the text as a variable-containing sentence or a non-variable-containing sentence, this classification method enables the navigation system to adopt the most appropriate processing strategy for different types of text. This differentiated processing not only optimizes the efficiency and naturalness of voice broadcast, but also improves the accuracy of information transmission and user satisfaction, ultimately enhancing the overall performance and practicality of the navigation system.

[0048] In some optional embodiments, according to the category of the sentence, a text segmentation strategy corresponding to the category is adopted to obtain segmented short sentences corresponding to the sentence, including:

[0049] If the category of the sentence is a variable-containing sentence, a first text segmentation strategy is adopted to obtain segmented short sentences corresponding to the sentence, the first text segmentation strategy being associated with the position of punctuation marks in the sentence, the length of static elements, and the length of dynamic elements;

[0050] If the category of the sentence is a non-variable-containing sentence, a second text segmentation strategy is adopted to obtain segmented short sentences corresponding to the sentence, the second text segmentation strategy being associated with the position of punctuation marks in the sentence, and whether the number of words before and after the punctuation marks satisfies a specific condition.

[0051] Specifically, a "variable-containing sentence" refers to those text sentences that change dynamically during navigation according to actual conditions (such as user location, traffic conditions, etc.), such as sentences containing variable information such as distance, direction, or time; while a "non-variable-containing sentence" refers to those sentences with fixed content that do not change with navigation conditions. For "variable-containing sentences", a "first text segmentation strategy" is adopted, which considers the position of punctuation marks in the sentence, the length of static elements (i.e. text parts that do not change), and the length of dynamic elements (i.e. text parts that change) when segmenting the text, to ensure that the segmented short sentences not only retain complete semantics, but also adapt to real-time updates of dynamic information. Conversely, for "non-variable-containing sentences", a "second text segmentation strategy" is adopted, which mainly relies on the position of punctuation marks and whether the number of words before and after the punctuation marks satisfies a specific condition to segment the text, with the aim of maintaining the integrity and readability of the sentence. Through these two different text segmentation strategies, different types of navigation text can be flexibly processed to ensure that the generated navigation voice is accurate and natural, while improving the efficiency of voice broadcast and the user's auditory experience.

[0052] In this way, by distinguishing between "sentences with variables" and "sentences without variables" and applying corresponding text segmentation strategies, fine processing of navigation voice broadcast text is realized. For "sentences with variables", the first text segmentation strategy takes into account the position of punctuation marks, the length of static and dynamic elements, ensuring that when dynamic information is updated, the segmented short sentences can still maintain the integrity and accuracy of the semantics, so that the broadcast of navigation information is both flexible and real-time. For "sentences without variables", the second text segmentation strategy is used, which is segmented according to the position of punctuation marks and specific word conditions, optimizing the readability of the text and the naturalness of the broadcast. This differentiated processing method significantly improves the broadcast effect of navigation voice, making the broadcast content clearer and easier to understand, while improving the efficiency and fluency of the broadcast.

[0053] In some optional embodiments, if the category of the sentence is a sentence with variables, the first text segmentation strategy is used to obtain the segmented short sentences corresponding to the sentence, including:

[0054] Identifying punctuation marks in the sentence with variables;

[0055] Cutting the sentence with variables according to the position of the punctuation marks to obtain multiple short sentences;

[0056] Aggregating static short sentences in the multiple short sentences, and performing adsorption processing on dynamic short sentences in the multiple short sentences according to the length of the short sentences to obtain the segmented short sentences corresponding to the sentence, the static short sentences being short sentences not containing variables, and the dynamic short sentences being short sentences containing variables.

[0057] Specifically, "sentences with variables" refer to text sentences that contain elements that will change according to actual navigation conditions (such as distance, time, etc.), and these sentences need to be dynamically updated to reflect real-time navigation information. "Punctuation marks" are symbols in the text, such as commas, periods, etc., which serve to separate and emphasize in the sentence.

[0058] The implementation process of this scheme includes: after obtaining the complete sentence with variables, text analysis technology is used to identify the positions of punctuation marks in these sentences. Next, using the punctuation marks as natural dividing points, the long sentence is segmented into multiple shorter sentences. This helps separate dynamically changing information from static information, making each short sentence more concise and clear, facilitating subsequent speech synthesis and playback. Then, text elements in multiple short sentences that do not change with navigation conditions (such as directional instructions "turn left," "turn right," etc.) are merged together to maintain the coherence and consistency of information. Simultaneously, dynamic short sentences within the multiple sentences are processed according to their length to ensure that they are both accurate and conform to the user's auditory habits during playback. Finally, segmented short sentences corresponding to the complete sentence are obtained. These short sentences contain the necessary navigation information and have undergone optimization to adapt to dynamic updates. In this way, the navigation system can process and broadcast navigation information more effectively, ensuring that the voice commands received by the user are both clear and timely, thereby improving navigation accuracy and the user's driving experience.

[0059] By identifying punctuation marks within sentences with variables and segmenting the sentences based on their positions, effective segmentation of the navigation text is achieved, generating multiple short sentences. This method not only improves the flexibility of text processing but also maintains information consistency and coherence by aggregating static sentences, making the navigation instructions received by the user clearer and easier to understand. Simultaneously, the absorption of dynamic sentences ensures that information changing according to actual navigation conditions (such as distance and time) is accurately and promptly updated and broadcast, thereby enhancing the real-time nature and accuracy of navigation information. This processing significantly improves the user's auditory experience, making the navigation voice more natural and fluent, adaptable to various navigation scenarios, and providing personalized voice guidance. Furthermore, this refined text segmentation and processing strategy helps improve the efficiency and quality of voice broadcasting, ensuring that users receive accurate and easily understandable navigation information under rapidly changing navigation conditions, thus improving the usability of the navigation system and user satisfaction.

[0060] In some optional embodiments, static phrases from multiple short sentences are aggregated, and dynamic phrases from multiple short sentences are snapped together according to their length to obtain segmented phrases corresponding to the whole sentence, including:

[0061] Identify static short sentences among multiple short sentences, and aggregate the static short sentences to obtain aggregated static sentences;

[0062] Identify dynamic short sentences among multiple short sentences. If the length of the current dynamic short sentence is less than the set number of characters, then merge the current dynamic short sentence with other dynamic short sentences to obtain the merged dynamic sentence.

[0063] Based on the aggregated static sentences and the adsorbed dynamic sentences, the segmented short sentences corresponding to the whole sentences are obtained.

[0064] Specifically, "static phrases" refer to fixed information in the broadcast text that does not change with navigation conditions, such as directional instructions ("turn left", "go straight"), while "dynamic phrases" refer to information that changes according to real-time navigation data, such as distance ("after 500 meters") and time ("expected arrival in 5 minutes").

[0065] The specific implementation process of this scheme includes: identifying static short sentences among multiple short sentences, that is, using text analysis technology to identify all unchanged text content from the segmented short sentences. Next, the static short sentences are aggregated, that is, these static information are merged into a coherent, aggregated static sentence. This maintains the integrity and consistency of fixed information in navigation instructions.

[0066] For dynamic phrases, the system identifies dynamic phrases among multiple phrases and compares their length with a preset word count threshold. If the current dynamic phrase's length is less than the set word count (e.g., four characters), it is merged with other dynamic phrases to form a more complete and easily understood merged dynamic sentence. This merging process helps optimize the fluency and naturalness of the speech, while ensuring that the speech of dynamic information is not too fragmented or discontinuous.

[0067] Finally, based on the aggregated static sentences and the absorbed dynamic sentences, segmented phrases corresponding to the complete sentences can be obtained. These segmented phrases contain both complete static navigation information and dynamically changing information, making the final broadcast content both accurate and coherent. In this way, the navigation system can provide clearer, more understandable, and more natural voice guidance, thereby improving the user's navigation experience and driving safety.

[0068] To facilitate understanding of the solution in this embodiment, the following example illustrates how to segment a sentence with variables: for instance, the sentence is "Prepare to depart, exit the current highway, and head towards Shucheng, Tongcheng, and Chizhou";

[0069] ① First, identify the punctuation marks in the whole sentence with variables, and cut the whole sentence according to the punctuation marks to get “Prepare to depart”, “Exit the current highway”, “To Shucheng”, “Tongcheng”, and “Chizhou direction”. Among them, “Prepare to depart” and “Exit the current highway” are static short sentences, while “To Shucheng”, “Tongcheng”, and “Chizhou direction” are dynamic short sentences.

[0070] ② Aggregate the static phrases from multiple short sentences. The static phrases are "preparing to depart" and "exiting the current highway". Aggregating the static phrases will give "preparing to exit the current highway".

[0071] ③ For dynamic short sentences in multiple short sentences, the short sentence is snapped up according to the sentence length. For example, if the short sentence length is less than 4, snapping is required. The dynamic short sentence "to Shucheng" has less than 4 characters, so it is snapped up with the following short sentence to get "to Shucheng Tongcheng". "Chizhou direction" has exactly 4 characters, so it is not snapped up.

[0072] After the above steps, the final short sentences are "Preparing to depart from the current expressway", "Towards Shucheng and Tongcheng", and "Towards Chizhou".

[0073] In this way, by identifying and aggregating static phrases from multiple short sentences, and intelligently absorbing dynamic phrases, optimized segmentation of navigation voice broadcast text is achieved. This processing method ensures the integrity and coherence of static information, enabling the aggregated static sentences to clearly convey unchanging navigation instructions, such as directions and location names. Simultaneously, for dynamic phrases, when their length is less than a preset number of characters, they are absorbed with other dynamic phrases to form absorbed dynamic sentences. This helps avoid redundancy and discontinuity in the broadcast process, ensuring that dynamic information such as distance and time is broadcast accurately and smoothly. The solution proposed in this application improves the clarity and naturalness of navigation voice, making it easier for users to understand and follow navigation instructions, thereby enhancing driving safety. Furthermore, by reducing pauses and discontinuities in the broadcast process, the efficiency of voice broadcast is improved, making the transmission of navigation information more efficient. In summary, this refined text processing and segmentation strategy makes navigation voice broadcast more in line with users' auditory habits and expectations.

[0074] In some optional embodiments, the current dynamic phrase is snapped together with other dynamic phrases to obtain a snapped dynamic phrase, including:

[0075] If there is no other dynamic phrase preceding the current dynamic phrase, then the current dynamic phrase is snapped together with the next dynamic phrase to obtain the snapped dynamic phrase.

[0076] Specifically, "dynamic phrases" refer to information segments in navigation text that change based on real-time data. For example, "500 meters later" in "turn right after 500 meters" is a dynamic phrase because it changes as the user's location updates. "Snapping" is a text processing technique that merges two or more adjacent dynamic phrases into a more coherent phrase or sentence to optimize the fluency and naturalness of the voice broadcast.

[0077] The specific implementation process of this scheme includes: First, identifying all dynamic phrases in the broadcast text. Then, for each dynamic phrase, checking whether its preceding element is also a dynamic phrase. If the preceding element is not a dynamic phrase, that is, "the preceding element of the current dynamic phrase does not contain any other dynamic phrases," then this current dynamic phrase is "absorbed" with the immediately following dynamic phrase. This absorption operation merges the two dynamic phrases into a more complete and coherent dynamic sentence, for example, merging "500 meters later" and "turn right" into "turn right after 500 meters."

[0078] This adsorption process ensures the continuity and clarity of navigation voice broadcasts when conveying dynamically changing information, avoiding unnatural or incomprehensible broadcasts caused by fragmented information. Simultaneously, this optimized broadcasting method also enhances the user's auditory experience, making navigation instructions more intuitive and easier to understand, thereby improving the usability of the navigation system and user satisfaction. Furthermore, the formation of dynamic sentences helps reduce pauses and interruptions during broadcasting, making the entire navigation voice broadcast process smoother and further improving the efficiency and effectiveness of navigation information delivery.

[0079] In this way, by merging a current dynamic phrase with the following dynamic phrase when there are no other dynamic phrases preceding it, forming a dynamic link, the coherence and naturalness of the navigation voice broadcast are significantly improved. This strategy optimizes the integration of dynamic information, avoiding potential interruptions and jumps during broadcasting, making navigation instructions smoother and easier to understand. Specifically, when a dynamic phrase in the navigation text, such as "500 meters later," is not immediately followed by a static phrase or punctuation mark but by a preceding dynamic phrase, it is intelligently merged with a subsequent dynamic phrase, such as "turn right," to generate a dynamic link like "turn right after 500 meters." This processing not only reduces pauses in the broadcast and improves the efficiency of information transmission but also enhances the user's auditory experience, enabling users to grasp navigation information more quickly and accurately, thereby improving driving safety and user satisfaction with the navigation system. Furthermore, the formation of this dynamic link also helps improve the adaptability of the voice broadcast, allowing it to flexibly respond to various navigation scenarios and provide users with more personalized and adaptable voice guidance.

[0080] In some optional embodiments, the current dynamic phrase is snapped together with other dynamic phrases to obtain a snapped dynamic phrase, which further includes:

[0081] If there is no other dynamic phrase after the current dynamic phrase, then the current dynamic phrase is snapped together with the previous dynamic phrase to obtain the snapped dynamic phrase.

[0082] Specifically, after identifying all dynamic phrases in the broadcast text, if no other dynamic phrase follows the current dynamic phrase, then the current dynamic phrase is "attached" to the previous dynamic phrase. For example, if the text contains "towards City A" and "City B," since "City B" is two characters, less than four characters, and no other element follows "City B," then "towards City A" and "City B" are attached to generate the dynamic phrase "towards City A, City B." It should be noted that this is only an example for ease of understanding and does not represent the actual situation.

[0083] This adsorption process improves the naturalness and fluency of the voice prompts, making navigation instructions easier to understand and follow. Secondly, by reducing pauses and interruptions during the broadcast, it increases the efficiency of information transmission, thereby enhancing the user's auditory experience. Furthermore, this dynamic sentence structure helps improve the adaptability of the navigation system, enabling it to flexibly respond to various navigation scenarios and provide users with more personalized and adaptable voice guidance. Finally, this optimized broadcast method also helps reduce the user's cognitive burden, allowing them to focus more on driving, thus improving driving safety and the usability of the navigation system.

[0084] In some optional embodiments, the current dynamic phrase is snapped together with other dynamic phrases to obtain a snapped dynamic phrase, which further includes:

[0085] If both the preceding and following elements of the current dynamic phrase exist in other dynamic phrases, then the current dynamic phrase will be snapped together with other dynamic phrases with fewer characters to obtain a snapped dynamic phrase.

[0086] Specifically, after identifying all dynamic phrases in the broadcast text, for each dynamic phrase, it checks whether its preceding and following elements are also dynamic phrases. If "both the preceding and following elements of the current dynamic phrase contain other dynamic phrases," then the word counts of these adjacent dynamic phrases are compared. Based on the word count comparison results, the current dynamic phrase is "snap-in" with other dynamic phrases with fewer words.

[0087] This solution improves the naturalness and fluency of voice prompts, making navigation instructions easier to understand and follow. Secondly, by reducing pauses and interruptions during playback, it increases the efficiency of information transmission, thereby enhancing the user's auditory experience. Furthermore, this dynamic sentence structure helps improve the adaptability of the navigation system, enabling it to flexibly respond to various navigation scenarios and provide users with more personalized and adaptable voice guidance. Finally, this optimized playback method helps reduce the user's cognitive burden, allowing them to focus more on driving, thus improving driving safety and the usability of the navigation system. In this way, the navigation system can process and broadcast navigation information more effectively, ensuring that the voice instructions received by the user are both clear and timely, thereby improving navigation accuracy and the user's driving experience.

[0088] In some optional embodiments, if the category of the entire sentence is an unvariable sentence, then a second text segmentation strategy is adopted to obtain segmented short sentences corresponding to the entire sentence, including:

[0089] If the entire sentence belongs to the category of "unspecified sentence", then identify the punctuation marks in the entire sentence;

[0090] If the number of characters before and after a punctuation mark meets the preset number of characters, then the sentence is split at the position of the punctuation mark that meets the condition, resulting in a segmented short sentence corresponding to the whole sentence.

[0091] Specifically, "unvarnished complete sentences" refers to complete sentences in the navigation announcements that do not contain any elements that change with actual navigation conditions, such as "There is a residential area ahead, please slow down." "Punctuation marks" are symbols in the text, such as commas and periods, which serve to separate and emphasize elements within sentences. "Preset word count" refers to the word limit set by the system to optimize the naturalness and fluency of the voice announcements; for example, it could be 6 words.

[0092] The specific implementation process of this scheme includes: First, identifying the first punctuation mark in a complete sentence belonging to the "unvarnished complete sentence" category. Next, checking whether the number of characters before and after the first punctuation mark meets a preset character count condition. This preset character count condition is likely based on best practices for voice playback to ensure that each segmented short sentence is neither too long nor too short, thereby improving the clarity of the playback and the user's auditory experience. For example, the preset character count condition could be to meet 6 characters.

[0093] If the number of characters before and after the first punctuation mark is 6, the entire sentence will be split at the position of the first punctuation mark, resulting in a corresponding short sentence. For example, if a sentence without variables is "There is a residential area ahead, please be careful of pedestrians passing by", and the number of characters before and after the comma meets the preset condition, then the sentence will be split into two short sentences: "There is a residential area ahead" and "Please be careful of pedestrians passing by".

[0094] It should be noted that if the number of characters before and after the first punctuation mark is less than 6, then look for the second punctuation mark. If the number of characters before and after the second punctuation mark is 6, then the sentence is cut at the position of the second punctuation mark. Generally, only the entire sentence without variables is cut in one stroke.

[0095] In this way, when the number of characters before and after punctuation marks meets preset conditions, text segmentation is performed at these locations to generate segmented short sentences. This method makes navigation instructions more direct and easier to follow because the segmented short sentences better align with users' information receiving and processing habits, enhancing the clarity and comprehensibility of the broadcast content. Simultaneously, this segmentation strategy based on punctuation and character count limitations improves the naturalness and fluency of the broadcast, avoiding auditory fatigue or comprehension difficulties that may result from excessively long sentences, thus enhancing the user's auditory experience. Furthermore, this strategy helps improve broadcast efficiency because it allows the system to process and broadcast information quickly, ensuring users receive timely navigation guidance. Overall, this feature-based processing not only optimizes the quality of navigation voice broadcasts but also improves driving safety and the usability of the navigation system, providing users with a more personalized and adaptive navigation experience.

[0096] In some optional embodiments, the segmented short sentences corresponding to the whole sentence are sent to the server, and the server generates navigation voice corresponding to the whole sentence, including:

[0097] The segmented short sentences corresponding to the whole sentence are sent to the server in parallel, and the server generates the navigation voice corresponding to the whole sentence.

[0098] Specifically, "segmented phrases" refer to shorter sentence fragments obtained from the original navigation text after processing according to a specific text segmentation strategy. These phrases contain various parts of the navigation information, such as direction indicators, distance tips, and location names, and are designed in an easy-to-understand and process format. "Server" here refers to a remote computing server with text-to-speech (TTS) conversion capabilities, enabling it to convert text information into naturally pleasing speech output.

[0099] The specific implementation process of this scheme includes: First, according to preset rules or strategies, the original navigation text is segmented into multiple short sentences. These short sentences are arranged in the order they appear in the original text to ensure the coherence and logic of the information. Then, these segmented short sentences are sent to the server in parallel. After receiving these short sentences, the server uses its built-in TTS engine to convert each short sentence into a corresponding speech segment. During this process, the server may adjust parameters such as speech rate, tone, and volume as needed to ensure that the generated speech is both clear and natural.

[0100] Finally, the server splices the generated audio segments together in the order they were sent to form a complete navigation audio message, which is then sent back to the client, i.e., the user's navigation device. This allows the user to hear coherent and clear navigation instructions, enabling them to understand and execute navigation information more accurately.

[0101] In this way, the solution achieves efficient generation and playback of navigation voice. Since the text-to-speech conversion is completed on the server, the computational burden on the client is reduced, and the powerful processing capabilities of the server ensure high-quality and natural-sounding voice. Furthermore, by sending segmented short sentences to the server in parallel, the server can quickly process and respond to client requests, generating and playing navigation voice in a timely manner, thus improving the real-time performance of the playback. This method not only enhances the user's navigation experience but also contributes to improving driving safety and efficiency.

[0102] In this way, because the segmented short sentences are sent to the server in parallel, this mechanism first ensures the coherence and logic of the broadcast content. Secondly, by utilizing advanced text-to-speech (TTS) technology on the server side, high-quality navigation voice can be generated quickly, significantly improving broadcast quality. The server-side TTS engine can intelligently adjust the speech rate, tone, and volume to generate clear and natural navigation voice, greatly enhancing the user's auditory experience. Furthermore, this mechanism enhances the real-time nature of the broadcast; the server can quickly process client requests, promptly generate and broadcast navigation voice, ensuring users can obtain key navigation information in a timely manner, thereby improving driving safety. Simultaneously, by performing text-to-speech conversion on the server, this solution optimizes resource utilization and reduces the computational burden on client devices, enabling even resource-limited devices to provide smooth and efficient navigation services.

[0103] For a better overall understanding of the process of initiating this application, please refer to [link / reference]. Figure 4 As shown, Figure 4This is an overall schematic diagram of the online navigation voice generation method in this embodiment. First, the complete sentence scripts required for navigation are obtained. These scripts may originate from user input, a map database, or preset templates. Next, these scripts are divided into two categories: sentences with variables and sentences without variables. For sentences without variables (text without dynamically changing elements), if the number of characters before and after a punctuation mark in the sentence is greater than or equal to 6, it is directly segmented at that punctuation mark to generate short sentences, which are then sent to the server. For sentences with variables (text containing elements that change according to navigation conditions), they are first fully segmented, and then further segmented based on whether the number of characters before and after a punctuation mark exceeds 4. Specifically, when processing the segmented short sentences, static short sentences are identified and aggregated, while dynamic short sentences with fewer than 4 characters are snapped together to form aggregated static sentences and snapped dynamic sentences. These processed short sentences are then sent to the server, where the server-side text-to-speech (TTS) engine quickly generates high-quality navigation voice based on this information.

[0104] This application significantly improves the quality and efficiency of navigation voice broadcasting by combining text classification, intelligent text segmentation, element aggregation and adsorption, and efficient server-side speech synthesis. This method not only solves the problem of excessively long generation time in existing technologies but also overcomes the drawback of unnatural broadcasting effects, providing users with a faster, more natural, and more auditory-friendly navigation voice service, thereby enhancing user experience and driving safety.

[0105] The following describes an apparatus embodiment of this application, which can be used to execute the online navigation voice generation method in the above embodiments of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the online navigation voice generation method described above.

[0106] This disclosure also provides an online navigation voice generation device 500, such as... Figure 5 As shown, it includes:

[0107] The acquisition module 501 is used to acquire the broadcast text of the navigation voice to be generated, and the broadcast text contains at least one complete sentence;

[0108] The segmentation module 502 is used to identify the category of the whole sentence in the broadcast text, and to use the text segmentation strategy corresponding to the category to obtain the segmented short sentences corresponding to the whole sentence.

[0109] The sending module 503 is used to send the segmented short sentences corresponding to the whole sentence to the server, and the server generates the navigation voice corresponding to the whole sentence.

[0110] In some optional embodiments, the segmentation module 502 identifies the category to which an entire sentence in the broadcast text belongs, including;

[0111] If the broadcast text contains elements that change according to the actual navigation conditions, the entire sentence belongs to the category of "entire sentence with variables";

[0112] If the broadcast text does not contain any elements that change according to the actual navigation conditions, the entire sentence belongs to the category of "unvariable sentence".

[0113] In some optional embodiments, the segmentation module 502 uses a text segmentation strategy corresponding to the category to which the whole sentence belongs to obtain segmented short sentences corresponding to the whole sentence, including:

[0114] If the category of the whole sentence is a sentence with variables, the first text segmentation strategy is adopted to obtain the segmented short sentences corresponding to the whole sentence. The first text segmentation strategy is related to the position of punctuation marks, the length of static elements and dynamic elements in the whole sentence.

[0115] If the category of the whole sentence is an unspecified whole sentence, then the second text segmentation strategy is adopted to obtain the segmented short sentences corresponding to the whole sentence. The second text segmentation strategy is related to the position of the punctuation marks in the whole sentence and whether the number of words before and after the punctuation marks meets specific conditions.

[0116] In some optional embodiments, if the segmentation module 502 classifies the entire sentence as a sentence with variables, it employs a first text segmentation strategy to obtain segmented short sentences corresponding to the entire sentence, including:

[0117] Identify punctuation marks in sentences with variables;

[0118] The entire sentence with variables is divided into multiple shorter sentences according to the position of punctuation marks;

[0119] The static short sentences from multiple short sentences are aggregated, and the dynamic short sentences from multiple short sentences are snapped together according to their length to obtain the segmented short sentences corresponding to the whole sentence. The static short sentences are short sentences that do not contain variables, and the dynamic short sentences are short sentences that contain variables.

[0120] In some optional embodiments, the segmentation module 502 aggregates static phrases from multiple sentences and performs adsorption processing on dynamic phrases from multiple sentences according to their length to obtain segmented phrases corresponding to the whole sentence, including:

[0121] Identify static short sentences among multiple short sentences, and aggregate the static short sentences to obtain aggregated static sentences;

[0122] Identify dynamic short sentences among multiple short sentences. If the length of the current dynamic short sentence is less than the set number of characters, then merge the current dynamic short sentence with other dynamic short sentences to obtain the merged dynamic sentence.

[0123] Based on the aggregated static sentences and the adsorbed dynamic sentences, the segmented short sentences corresponding to the whole sentences are obtained.

[0124] In some optional embodiments, the segmentation module 502 snaps the current dynamic phrase with other dynamic phrases to obtain a snapped dynamic phrase, including:

[0125] If there is no other dynamic phrase preceding the current dynamic phrase, then the current dynamic phrase is snapped together with the next dynamic phrase to obtain the snapped dynamic phrase.

[0126] In some optional embodiments, the segmentation module 502 merges the current dynamic phrase with other dynamic phrases to obtain a merged dynamic phrase, and further includes:

[0127] If there is no other dynamic phrase after the current dynamic phrase, then the current dynamic phrase is snapped together with the previous dynamic phrase to obtain the snapped dynamic phrase.

[0128] In some optional embodiments, the segmentation module 502 merges the current dynamic phrase with other dynamic phrases to obtain a merged dynamic phrase, and further includes:

[0129] If both the preceding and following elements of the current dynamic phrase exist in other dynamic phrases, then the current dynamic phrase will be snapped together with other dynamic phrases with fewer characters to obtain a snapped dynamic phrase.

[0130] In some optional embodiments, if the segmentation module 502 classifies the entire sentence as an unvariable sentence, it employs a second text segmentation strategy to obtain segmented short sentences corresponding to the entire sentence, including:

[0131] If the entire sentence belongs to the category of "unspecified sentence", then identify the punctuation marks in the entire sentence;

[0132] If the number of characters before and after a punctuation mark meets the preset number of characters, then the sentence is split at the position of the punctuation mark that meets the condition, resulting in a segmented short sentence corresponding to the whole sentence.

[0133] In some optional embodiments, the sending module 503 sends the segmented short sentences corresponding to the whole sentence to the server, and the server generates the navigation voice corresponding to the whole sentence, including:

[0134] The segmented short sentences corresponding to the whole sentence are sent to the server in parallel, and the server generates the navigation voice corresponding to the whole sentence.

[0135] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0136] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0137] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0138] like Figure 6 As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may also store various programs and data required for the operation of the device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0139] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 608, such as various types of monitors, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0140] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as the online navigation voice generation method. For example, in some embodiments, the online navigation voice generation method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the applet distribution described above can be performed. Alternatively, in other embodiments, the computing unit 601 can be configured to perform the online navigation voice generation method by any other suitable means (e.g., by means of firmware).

[0141] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0142] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to the processor or controller of a general-purpose computer, special-purpose computer, or other programmable navigation voice online generation device, such that when executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0143] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0144] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0145] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0146] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0147] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this disclosure can be achieved, and this is not limited herein.

[0148] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for online generation of navigation voice, wherein, The method includes: Obtain the text to be used for generating navigation voice, wherein the text contains at least one complete sentence; Identify the category of the entire sentence in the broadcast text, and adopt the text segmentation strategy corresponding to the category to obtain the segmented short sentences corresponding to the entire sentence; The segmented short sentences corresponding to the complete sentence are sent to the server, and the server generates navigation voice corresponding to the complete sentence.

2. The method according to claim 1, wherein, The process of identifying the category of a complete sentence in the broadcast text includes: If the broadcast text contains elements that change according to the actual navigation conditions, then the entire sentence belongs to the category of a sentence with variables; If the broadcast text does not contain any elements that change according to the actual navigation conditions, then the entire sentence belongs to the category of an unvariable sentence.

3. The method according to claim 2, wherein, The step of using a text segmentation strategy corresponding to the category to which the whole sentence belongs to obtain segmented short sentences corresponding to the whole sentence includes: If the category of the whole sentence is a sentence with variables, then the first text segmentation strategy is adopted to obtain the segmented short sentences corresponding to the whole sentence. The first text segmentation strategy is related to the position of the punctuation marks, the length of static elements and dynamic elements in the whole sentence. If the category of the whole sentence is a whole sentence without variables, then the second text segmentation strategy is adopted to obtain the segmented short sentences corresponding to the whole sentence. The second text segmentation strategy is related to the position of the punctuation marks in the whole sentence and whether the number of characters before and after the punctuation marks meets specific conditions.

4. The method according to claim 3, wherein, If the category of the entire sentence is a sentence with variables, then the first text segmentation strategy is adopted to obtain the segmented short sentences corresponding to the entire sentence, including: Identify punctuation marks in the sentence with variables; The sentence with variables is divided into multiple short sentences according to the position of the punctuation marks; The static short sentences in the plurality of short sentences are aggregated, and the dynamic short sentences in the plurality of short sentences are snapped according to the sentence length to obtain the segmented short sentences corresponding to the whole sentence. The static short sentences are short sentences that do not contain variables, and the dynamic short sentences are short sentences that contain variables.

5. The method according to claim 4, wherein, The process of aggregating static short sentences from the plurality of short sentences and absorbing dynamic short sentences from the plurality of short sentences according to their length to obtain segmented short sentences corresponding to the whole sentence includes: Identify static short sentences among the multiple short sentences, and aggregate the static short sentences to obtain aggregated static sentences; Identify dynamic short sentences among the multiple short sentences. If the length of the current dynamic short sentence is less than a set number of characters, then the current dynamic short sentence is merged with other dynamic short sentences to obtain a merged dynamic sentence. Based on the aggregated static sentences and the adsorbed dynamic sentences, segmented short sentences corresponding to the whole sentences are obtained.

6. The method according to claim 5, wherein, The step of absorbing the current dynamic phrase with other dynamic phrases to obtain the absorbed dynamic phrase includes: If there are no other dynamic phrases preceding the current dynamic phrase, then the current dynamic phrase is snapped together with the next dynamic phrase to obtain the snapped dynamic phrase.

7. The method according to claim 5, wherein, The step of absorbing the current dynamic phrase with other dynamic phrases to obtain the absorbed dynamic phrase further includes: If there is no other dynamic phrase after the current dynamic phrase, then the current dynamic phrase is snapped together with the previous dynamic phrase to obtain the snapped dynamic phrase.

8. The method according to claim 5, wherein, The step of absorbing the current dynamic phrase with other dynamic phrases to obtain the absorbed dynamic phrase further includes: If both the preceding and following elements of the current dynamic phrase exist in other dynamic phrases, then the current dynamic phrase is snapped together with other dynamic phrases with fewer characters to obtain a snapped dynamic phrase.

9. The method according to claim 3, wherein, If the category of the entire sentence is an unspecified sentence, then the second text segmentation strategy is adopted to obtain the segmented short sentences corresponding to the entire sentence, including: If the entire sentence belongs to the category of "unspecified sentence", then identify the punctuation marks in the entire sentence; If the number of characters before and after the punctuation mark meets the preset number of characters, then the sentence is segmented at the position of the punctuation mark that meets the condition to obtain a segmented short sentence corresponding to the whole sentence.

10. The method according to any one of claims 1 to 9, wherein, The step of sending the segmented short sentences corresponding to the whole sentence to the server, and generating navigation voice corresponding to the whole sentence through the server, includes: The segmented short sentences corresponding to the complete sentence are sent to the server in parallel, and the server generates the navigation voice corresponding to the complete sentence.

11. An online navigation voice generation device, wherein, The device includes: The acquisition module is used to acquire the broadcast text of the navigation voice to be generated, wherein the broadcast text contains at least one complete sentence; The segmentation module is used to identify the category of the whole sentence in the broadcast text, and to adopt the text segmentation strategy corresponding to the category of the whole sentence to obtain the segmented short sentences corresponding to the whole sentence; The sending module is used to send the segmented short sentences corresponding to the whole sentence to the server, and the server generates the navigation voice corresponding to the whole sentence.

12. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method of any one of claims 1-10.

13. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-10.

14. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-10.

Citation Information

Patent Citations

  • Voice navigation method, device and system

    CN104197946A

  • Navigation voice broadcasting method and apparatus

    CN105788588A

  • Voice broadcasting method, mobile payment device, storage medium and computer device

    CN117577092A

  • Text segmentation method and device and terminal equipment

    CN119920233A

  • Text error correction method and device, electronic equipment and storage medium

    CN120336535A