Navigation voice broadcast method, device and equipment and computer readable storage medium
By filtering and sorting navigation voice content based on vehicle driving control mode and environmental information, the problem of chaotic navigation voice broadcasting has been solved, achieving orderly broadcasting and timely delivery of important information, thereby improving driving safety and passenger experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-04
- Publication Date
- 2026-03-24
AI Technical Summary
Navigation voice prompts may be confusing in certain locations, interfering with the driver's ability to obtain important information and affecting driving safety and passenger experience.
Based on the vehicle's driving control mode and the current location's environmental information, target voice content that matches the voice broadcast mode is selected, and the broadcast priority is sorted to generate a sequence of target voice content for broadcast.
Ensure that navigation voice broadcasts are neat and orderly, enabling drivers to obtain important information in a timely manner, thereby improving driving safety and passenger experience.
Smart Images

Figure CN115900750B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, specifically to a navigation voice broadcasting method, apparatus, device, and computer-readable storage medium. Background Technology
[0002] With the improvement of electronic map information, electronic navigation functions have been gradually applied in various regions, providing users with great convenience for travel. However, in electronic navigation functions, navigation voice can further provide users with guidance information, especially during vehicle operation. Navigation voice is generally combined with navigation routes, providing drivers with traffic-related voice prompts while driving, reducing drivers' reliance on visual navigation functions, and improving driving safety to a certain extent.
[0003] However, when related technologies broadcast navigation voice prompts, the content of the broadcasts generally depends on the vehicle's location during travel, such as rest reminders near highway service areas, turn-off prompts at intersections, and navigation voice prompts for tourist areas. In some special locations, multiple navigation voice prompts may need to be broadcast at the same time, which makes the navigation voice prompts relatively chaotic, interfering with the driver's acquisition of important voice information, easily misleading the vehicle's driving, reducing driving safety, and affecting the passenger experience due to the overly complicated navigation voice prompts. Summary of the Invention
[0004] This application provides a navigation voice broadcasting method, apparatus, device, and computer-readable storage medium, which enables the navigation voice to be broadcast in a neat and orderly manner, which is beneficial for the driver to obtain important voice information and avoids the impact of complicated navigation voice on the passenger experience.
[0005] This application provides a navigation voice broadcast method, including:
[0006] Determine the current driving control mode of the vehicle, and determine the voice broadcast mode corresponding to the driving control mode;
[0007] Obtain environmental information corresponding to the current target location of the vehicle, and obtain a set of voice content corresponding to the environmental information;
[0008] Target audio content that conforms to the audio playback mode is selected from the audio content set, and the target audio content is sorted by playback priority to obtain a target audio content sequence.
[0009] The target speech content in the target speech content sequence is broadcast in sequence.
[0010] Accordingly, embodiments of this application provide a navigation voice broadcasting device, including:
[0011] The determining unit is used to determine the current driving control mode of the vehicle and the voice broadcast mode corresponding to the driving control mode.
[0012] The acquisition unit is used to acquire environmental information corresponding to the current target location of the vehicle, and to acquire a set of voice content corresponding to the environmental information;
[0013] The sorting unit is used to filter out target speech content that conforms to the speech broadcasting mode from the speech content set, and sort the target speech content according to the broadcasting priority to obtain a target speech content sequence.
[0014] The broadcasting unit is used to broadcast the target speech content in the target speech content sequence in sequence.
[0015] In some embodiments, the acquiring unit is further configured to:
[0016] Multiple information categories are identified from the environmental information;
[0017] Determine the voice attribute category to which each information category belongs;
[0018] Based on the environmental information and voice attribute category, select the corresponding voice content to be processed, and construct the voice content set corresponding to the voice content to be processed.
[0019] In some embodiments, the acquiring unit is further configured to:
[0020] Identify the road structures and target building categories corresponding to the road condition information, and determine the target distance between the road structures and the target location;
[0021] Query multiple traffic condition voice contents corresponding to the traffic condition voice category from the voice content library, and extract the target building voice contents corresponding to the target building category from the multiple traffic condition voice contents;
[0022] Extract architectural speech templates from the target architectural speech content, and generate speech content to be processed based on the target distance and architectural speech templates.
[0023] In some embodiments, the environmental information includes road sign information, the voice attribute category includes road sign voice category, and the acquisition unit is further configured to:
[0024] Identify the indicative semantics in the road sign information;
[0025] Query the voice content library for multiple road sign voice contents corresponding to the road sign voice category;
[0026] The target road sign voice content corresponding to the indication semantics is matched from the multiple road sign voice contents, and the target road sign voice content is determined as the voice content to be processed.
[0027] In some embodiments, the sorting unit is further configured to:
[0028] Determine the playable category associated with the voice broadcasting mode;
[0029] Select target audio content that matches the playable category from the audio content set.
[0030] In some embodiments, the sorting unit is further configured to:
[0031] Determine the priority weight of each target speech content at the target location;
[0032] Based on the priority weight of each target speech content, the playback priority among the target speech content is determined;
[0033] The target speech content is sorted according to the broadcast priority to obtain a target speech content sequence.
[0034] In some embodiments, the target voice content includes traffic voice content and weather voice content, and the sorting unit is further configured to:
[0035] Identify road and building information in the environmental information, and determine the first risk coefficient corresponding to the road and building information;
[0036] Identify the climate information in the environmental information and determine the second risk coefficient corresponding to the climate information;
[0037] Based on the first risk coefficient and the second risk coefficient, the priority weights of the road condition voice content and the climate voice content are determined respectively.
[0038] In some embodiments, the navigation voice broadcasting device further includes an updating unit for:
[0039] Read the content of the played voice messages from the list of played voice messages;
[0040] When the target speech content is identified to be consistent with the already played speech content, the duration of the playback interval of the already played speech content at the current time is determined;
[0041] The target speech content sequence is updated according to the broadcast interval duration to obtain the updated target speech content sequence;
[0042] The broadcasting unit is further configured to sequentially broadcast the target speech content in the updated target speech content sequence.
[0043] Furthermore, this application also provides a computer device, including a processor and a memory, wherein the memory stores a computer program, and the processor is used to run the computer program in the memory to implement the steps in the navigation voice broadcasting method provided in this application.
[0044] Furthermore, embodiments of this application also provide a computer-readable storage medium storing multiple instructions adapted for loading by a processor to execute steps in any of the navigation voice broadcasting methods provided in embodiments of this application.
[0045] Furthermore, embodiments of this application also provide a computer program product, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps in any of the navigation voice broadcasting methods provided in embodiments of this application.
[0046] This application embodiment can determine the vehicle's current driving control mode and the corresponding voice broadcast mode; obtain environmental information corresponding to the vehicle's current target location and a set of voice content corresponding to the environmental information; filter target voice content that conforms to the voice broadcast mode from the voice content set, and sort the target voice content by broadcast priority to obtain a target voice content sequence; and broadcast the target voice content in the target voice content sequence in sequence. Therefore, this solution can optimize the voice content that needs to be broadcast while the vehicle is in motion. Specifically, first, the vehicle's current location and driving control mode are determined, and the voice broadcast mode under the current driving control mode is determined. Simultaneously, the set of voice content to be broadcast is determined based on the environmental information of the current location; then, the voice content in the voice content set is filtered according to the voice broadcast mode, and the playback priority of the filtered voice content is sorted; finally, the voice content is broadcast according to the sorted broadcast priority. This ensures that the navigation voice is orderly and organized during broadcast, which is beneficial for the driver to obtain important voice information and avoids complicated navigation voice affecting the passenger experience. Attached Figure Description
[0047] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0048] Figure 1This is a schematic diagram of a navigation voice broadcast system provided in an embodiment of this application;
[0049] Figure 2 A flowchart illustrating the steps of the navigation voice broadcasting method provided in this application embodiment;
[0050] Figure 3 This is a schematic flowchart of another step of the navigation voice broadcasting method provided in the embodiments of this application;
[0051] Figure 4 This is a schematic diagram of the architecture of the navigation voice broadcasting device provided in the embodiments of this application;
[0052] Figure 5 This is a schematic diagram of the navigation voice broadcasting device provided in the embodiments of this application;
[0053] Figure 6 This is a schematic diagram of the structure of the computer device provided in the embodiments of this application. Detailed Implementation
[0054] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0055] This application provides a navigation voice broadcast method, apparatus, device, and computer-readable storage medium. This application will describe the navigation voice broadcast apparatus from the perspective of a navigation voice broadcast device, which can be integrated into a computer device. This computer device can be a terminal device, specifically a terminal device mounted on a vehicle, i.e., an in-vehicle terminal. Furthermore, the terminal device can also be other types of devices, such as a television, smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, smart wearable device, etc.; however, it is not limited to these.
[0056] For example, see Figure 1 This is a schematic diagram of a navigation voice broadcast system provided in an embodiment of this application. The control system is applicable to both manual and automatic driving scenarios, and is not specifically limited to navigation voice broadcast scenarios on highways or in parking lots. This scenario includes a terminal or a server.
[0057] Specifically, the terminal can be an in-vehicle terminal, used to determine the current driving control mode of the vehicle and the corresponding voice broadcast mode; obtain environmental information corresponding to the target location and obtain a set of voice content corresponding to the environmental information; filter out target voice content that matches the voice broadcast mode from the voice content set, sort the target voice content by broadcast priority, and obtain a target voice content sequence; and broadcast the target voice content in the target voice content sequence in sequence.
[0058] It should be noted that when the navigation voice broadcast system includes a server, a communication connection can be established between the in-vehicle terminal and the server. The server can determine the vehicle's current driving control mode and the corresponding voice broadcast mode; obtain the environmental information corresponding to the target location and the corresponding voice content set; filter the target voice content that matches the voice broadcast mode from the voice content set, and sort the target voice content by broadcast priority to obtain a target voice content sequence; then, send the target voice content sequence to the in-vehicle terminal, so that the in-vehicle terminal broadcasts the target voice content in the target voice content sequence in sequence.
[0059] For example, in the case of autonomous driving, although the driver can relinquish control of the vehicle, the in-vehicle terminal can provide navigation voice prompts so that the driver or passengers can understand problems during the driving process. Specifically, the in-vehicle terminal can determine the vehicle's current target location and driving control mode, which may be autonomous driving mode. Furthermore, autonomous driving mode can be further subdivided into adaptive cruise control, lane centering assist, and highway autonomous navigation driving modes, etc., without limitation here. Different driving control modes correspond to different voice broadcast modes, such as driving behavior-related broadcast modes like road condition voice, or combined broadcast modes like road condition voice and weather voice. Then, it collects surrounding environmental information at the target location, such as road conditions, weather, road signs, and surrounding objects, using camera components, LiDAR components, and humidity sensors, to select one or more voice contents based on the environmental information that can trigger voice broadcast. Next, voice contents that do not conform to the broadcast mode are filtered. For example, in the combined road condition and weather voice broadcast mode, tourist area introduction voices are filtered, and the filtered voice contents are sorted in order of broadcast, resulting in a sorted target voice content sequence. Finally, according to the sorting relationship, the target voice contents (such as road condition voice and weather voice) in the target voice content sequence are broadcast. The above are merely examples under autonomous driving mode and are not intended to limit the implementation of this application.
[0060] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the preferred order of the embodiments.
[0061] In this embodiment, the description will focus on a navigation voice broadcasting device, which can be integrated into a computer device such as a terminal device or a server. See also Figure 2 , Figure 2 The following is a flowchart illustrating the steps of a navigation voice broadcast method provided in this application embodiment. Taking a terminal device as an example, which is a terminal mounted on a vehicle, the specific process of the navigation voice broadcast method when the processor on the terminal device executes the program instructions corresponding to the navigation voice broadcast method is as follows:
[0062] 101. Determine the current driving control mode of the vehicle and the corresponding voice broadcast mode.
[0063] This application mainly relates to voice broadcast services applied to vehicles, such as voice broadcast services in electronic navigation or autonomous driving modes. In a voice broadcast service, there may be multiple voice contents that need to be broadcast at the same time. In order to play some or all categories of voice contents in an orderly manner later, the voice broadcast mode of the vehicle can be used to determine the broadcast mode.
[0064] The driving control mode can be the vehicle's driving mode, such as manual driving mode, driver assistance mode, or autonomous driving mode. Furthermore, the driving mode can be further subdivided, such as adaptive cruise control, lane centering assist, and highway autonomous navigation driving.
[0065] To determine the current driving control mode of the vehicle, the process can be as follows: acquire the vehicle's driving control logic from the previous moment, and determine the current driving control mode based on that logic. It's understandable that different driving control modes correspond to different driving control logics. For example, electronic navigation mode actually provides driving behavior instructions, but the actual driving behavior is still manually controlled by the driver, who manipulates relevant hardware to input control signals. Its control logic typically includes two parts: listening to control signals and responding to control information; there is no explicit overall control logic. In contrast, autonomous driving mode can initially construct an overall control logic based on a pre-built planned path, and then precisely adjust the accuracy of this overall control logic by combining real-time location and environmental information. For instance, if the overall control logic includes "vehicle turning" control logic in a lane turning section, the precise control logic corresponding to the "steering heading angle" and "rate of change of heading angle" can be determined by combining the actual road condition information of the vehicle in that lane turning section. Therefore, the vehicle's driving control mode can be determined based on the vehicle's driving control logic, such as whether the vehicle is in manual driving or assisted driving mode, or autonomous driving mode (which can be further subdivided into adaptive cruise control, highway autonomous navigation driving, etc.).
[0066] Optionally, different voice broadcast modes can be set for different driving control modes, meaning different driving control modes correspond to different voice broadcast modes. The types, number, and order of voice messages broadcast in different voice broadcast modes may differ, allowing occupants to capture key voice information in the corresponding driving control mode.
[0067] The voice broadcast mode can be a way of playing navigation voice information, which is associated with the playable categories. That is, the voice broadcast mode limits the categories of voices that can be broadcast in this mode. For example, in autonomous driving mode, the voice broadcast mode can be a detailed (multi-category) broadcast mode, which allows the broadcast of multiple categories of navigation voice information. For example, it can play voices related to driving behavior, such as road conditions and actions that the vehicle needs to perform (such as turning left or going straight), or it can play voices related to weather (such as rain, light intensity, and temperature).
[0068] By using the above methods, the corresponding voice broadcast mode can be determined according to the current driving mode of the vehicle, so that the corresponding type of voice can be broadcast according to the voice broadcast mode, so that the passengers in the vehicle can receive relevant voice information.
[0069] 102. Obtain the environmental information corresponding to the current target location of the vehicle, and obtain the set of voice content corresponding to the environmental information.
[0070] In this embodiment of the application, after determining the corresponding voice broadcast mode, it is also necessary to determine the voice content to be broadcast at each time point or each location. Specifically, when determining the voice content to be broadcast, the corresponding voice content to be broadcast at the corresponding time can be selected based on the environmental information of the vehicle's real-time location.
[0071] The environmental information can be the surrounding environment of the vehicle's location, which can serve as the trigger for navigation voice prompts, allowing the selection of the appropriate navigation voice for playback. It should be noted that this environmental information can include multiple categories, such as but not limited to road conditions, weather, road signs, greenbelts, and specific areas (e.g., scenic spots).
[0072] It should be noted that environmental information can be obtained through relevant sensors, such as camera components or infrared sensor components, to collect road condition information, climate information, road sign information, green belt information, and information about specific areas (such as scenic spots).
[0073] In addition, high-precision electronic maps can be used to obtain road condition information, road sign information, green belt information, and information on specific areas (such as scenic spots), as well as climate information in conjunction with weather service applications. It should be noted that, to improve the accuracy of environmental information, camera components can be used in conjunction with high-precision maps and weather service applications to acquire environmental information.
[0074] The voice content set contains one or more voice content items. Each voice content item is determined based on the current location and environmental information. Specifically, it selects the voice content that can be triggered and played when the vehicle is in its current location, taking into account the environmental conditions it is currently facing. For example, if the environmental information at the vehicle's location includes road conditions (such as buildings), weather, road signs, rest areas, etc., this environmental information is used as the trigger condition for the navigation voice, so that the corresponding voice content can be selected as the navigation voice that can be played at the current location.
[0075] In some implementations, the playable speech category can be determined based on environmental information, and the corresponding speech content can be selected by combining the environmental information and the speech category. For example, "obtaining the set of speech content corresponding to the environmental information" in step 102 may include:
[0076] (102.1) Identify multiple information categories from environmental information;
[0077] (102.2) Determine the voice attribute category to which each information category belongs;
[0078] (102.3) Select the corresponding speech content to be processed based on the environmental information and speech attribute category, and construct the speech content set corresponding to the speech content to be processed.
[0079] The information category can be an environmental information category, i.e., a classification of environmental information. For example, taking a vehicle at a target location on a highway as an example, the environmental information includes road condition information, weather information, road sign information, service area information before the tunnel, etc., which are located not far ahead of the lane. The corresponding information categories can include road condition information category, weather information category, road sign information category, service area information category, etc. The above is only an example and does not constitute a specific limitation for implementing this application. Other environmental information may also be included.
[0080] Furthermore, each information category is a general term for environmental information of the corresponding category, representing the attributes of the information. For example, in the climate information category, environmental information such as light rain, heavy rain, torrential rain, sunshine, fog, and haze belong to the climate information attribute, so the corresponding information category is the climate information category. Similarly, different road signs contain different directional information; information such as going straight, turning, U-turn, lane merging, lane diverging, and distance to xx city in xxx kilometers belong to the road sign information attribute, so the corresponding information category is the road sign information category. Other environmental information classifications are related to the attributes of the information, and will not be elaborated upon here.
[0081] The voice attribute category can be the voice category corresponding to the information category. For example, the traffic information category corresponds to the traffic voice category, the climate information category corresponds to the climate voice category, the road sign information category corresponds to the road sign voice category, and the specific area information category corresponds to the area introduction information category.
[0082] The audio content to be processed is the audio content that needs to be filtered and sorted.
[0083] Specifically, when acquiring the voice content corresponding to environmental information, the information categories contained in the current location's environmental information can be determined first, and the voice attribute categories corresponding to each information category can be determined. Then, the voice content corresponding to the specific environmental information can be found based on the voice attribute categories. In particular, when searching for the voice content corresponding to specific environmental information, the main focus is on finding voice content that matches the actual environmental information of the current location. For ease of understanding, the following will use actual environmental information as an example to describe the voice content acquisition process.
[0084] In some implementations, speech content can be generated using speech templates, specifically by combining hard environment information with corresponding speech templates. For example, if the environmental information includes road condition information and the speech attribute category includes road condition speech category, then step (102.3) of "selecting the corresponding speech content to be processed based on the environmental information and speech attribute category" can include: identifying the road buildings and the corresponding target building category corresponding to the road condition information, and determining the target distance between the road buildings and the target location; querying multiple road condition speech contents corresponding to the road condition speech category from the speech content library, and extracting the target building speech content corresponding to the target building category from the multiple road condition speech contents; extracting the building-type speech template from the target building speech content, and generating the speech content to be processed based on the target distance and the building-type speech template.
[0085] For example, if environmental information includes road condition information, then the voice attribute category includes the road condition voice category. The road condition information can be building information in the lane, such as tunnels, bridges, overpasses, downhill slopes, continuous curves, etc. When identifying building information in road condition information, environmental images can be captured in real time using a camera component. A lane-building recognition model is then used to identify the building information within these images, thus identifying road buildings and their categories. Simultaneously, a LiDAR component measures the target distance between the road building and the vehicle. Next, multiple road condition voice contents (or road condition voice statements) matching the road condition voice category are identified from a voice content library (which can be understood as a navigation voice corpus). Then, the target building voice content corresponding to the target building category is selected to avoid errors in voice content acquisition due to identical or similar voice information (such as word information in statements). Finally, a building-type voice template corresponding to the target building voice content is extracted. This template can be understood as an incomplete voice content template, containing unfilled semantic slots but including the voice word information of the road building. Therefore, when generating complete voice content, the corresponding information can be filled in, such as filling the building-type voice template with "target distance" to obtain the building voice content corresponding to the lane building. This building voice content is then used as the voice content to be processed.
[0086] It should be noted that the identification of road structures and their categories in road condition information can also be determined using the precision information contained in high-precision maps. Furthermore, when determining the target distance between a vehicle and a lane structure, the distribution relationship between lane structure information and the vehicle's current location information within the planned path in electronic navigation can be considered.
[0087] In some implementations, the corresponding speech content can be obtained based on the semantic information in the environmental information. For example, the environmental information includes road sign information, and the speech attribute category includes road sign speech category. The step (102.3) of "selecting the corresponding speech content to be processed based on the environmental information and speech attribute category" may include: identifying the indicative semantics in the road sign information; querying multiple road sign speech contents corresponding to the road sign speech category from the speech content library; matching the target road sign speech content corresponding to the indicative semantics from the multiple road sign speech contents, and determining the target road sign speech content as the speech content to be processed.
[0088] For example, taking road sign information as environmental information, different road sign information contains different directional semantics. The specific process of obtaining the road sign audio content corresponding to the road sign information is as follows: First, an environmental information image is acquired. The road sign semantic recognition model is used to identify whether the environmental information image contains road sign information. When the environmental information image contains road sign information, the directional semantics or road sign semantic tags corresponding to the road sign information are identified, such as the "turn" directional semantics. Then, multiple road sign audio contents corresponding to the road sign audio category are queried from the audio content library. Finally, the target road sign audio content corresponding to the "turn" directional semantics can be selected from the multiple road sign audio contents, such as the target road sign audio content corresponding to the "turn" directional semantics, the target road sign audio content corresponding to the "go straight" directional semantics, etc., as the audio content to be processed.
[0089] In some implementations, environmental information includes climate information. Climate attributes (such as rain, sunshine, and fog) can be identified from environmental information images, or climate attributes can be determined using climate information provided by a climate service application. Then, multiple climate voice contents corresponding to the climate voice category are queried from a voice content library. The target climate voice content corresponding to the climate attribute is matched from these multiple climate voice contents, and this target climate voice content is identified as the voice content to be processed. The implementation description of this climate voice content can be referred to the aforementioned implementation process of "road sign voice content," and will not be elaborated upon here.
[0090] Using the above methods, the playable voice content can be selected based on the environmental information of the vehicle's real-time location, thereby identifying multiple categories of voice content. This allows for the subsequent selection and sorting of more important voice content from these categories.
[0091] 103. Select target speech content that matches the speech playback pattern from the speech content set, and sort the target speech content according to the playback priority to obtain the target speech content sequence.
[0092] In this embodiment of the application, in order to ensure that the voice content to be broadcast in a neat and orderly manner and in accordance with the current driving control mode of the vehicle, the voice content can be filtered and sorted to generate a target voice content sequence, so that the voice content can be broadcast in accordance with the target voice content sequence in the future, so as to avoid the phenomenon that complicated voice content will cause trouble to the passengers in the vehicle during the broadcast.
[0093] In some implementations, when filtering voice content in the voice content set, the filtering can be based on the vehicle's current driving control mode. Different driving control modes have different voice broadcast modes; that is, the voice content is filtered by the voice broadcast mode to ensure that the broadcast voice content conforms to the vehicle's current voice broadcast mode. Specifically, step 103, "filtering target voice content that conforms to the voice broadcast mode from the voice content set," may include: determining the playable category associated with the voice broadcast mode, which is the category of voices allowed to be broadcast under the current voice broadcast mode; and filtering target voice content that conforms to the playable category from the voice content set.
[0094] For example, taking the high-speed autonomous navigation driving mode as the driving control mode, the corresponding voice broadcast mode could be a detailed broadcast mode. In this mode, multiple categories of voice content can be broadcast. For instance, in addition to providing voice broadcasts about road conditions related to autonomous driving safety, there could be voice broadcasts about road signs, weather, and introductions to specific locations. Therefore, first, the playable categories associated with this voice broadcast mode are determined. Since each voice content in the voice content set has a corresponding voice attribute category, voice content that does not conform to the playable categories can be filtered based on the matching between the playable categories and the voice attribute categories, thereby obtaining the target voice content that conforms to this voice broadcast mode. The above is merely an example.
[0095] For example, in lane centering assist mode as a driving control mode, the corresponding voice broadcast mode can be a simple broadcast mode. This simple broadcast mode can broadcast only one category of voice messages, such as road condition messages. For instance, it can only provide voice broadcasts of road conditions and the vehicle's actual deviation distance. Therefore, other unnecessary voice content can be filtered out according to the voice broadcast mode to ensure that the driver pays attention to road condition information and improve driving safety.
[0096] In this embodiment of the application, after filtering the voice content in the voice content set, the playback priority of the target voice content to be played can be sorted so that the target voice content to be played has a sequential order, so that the driver and passengers can obtain the more important voice information first.
[0097] In some implementations, the target speech content can be sorted according to the priority weights among the various target speech contents to obtain a target speech content sequence. Specifically, step 103, "sorting the target speech content by playback priority to obtain a target speech content sequence," may include:
[0098] (103.1) Determine the priority weight of each target speech content at the target location;
[0099] (103.2) Determine the playback priority among target speech contents based on the priority weight of each target speech content;
[0100] (103.3) Sort the target speech content according to the broadcast priority to obtain the target speech content sequence.
[0101] The priority weight can represent the importance of the voice content at the target location, and it can serve as the basis for ranking multiple target voice contents. For example, if driving safety is the priority, and multiple target voice contents include road condition voice, weather voice, and voice containing information about tourist areas near the lane, then the road condition voice content related to driving behavior has a higher priority than the weather voice content, and the weather voice content has a higher priority than the voice content containing information about tourist areas near the lane.
[0102] Specifically, after determining the priority weight of each target speech content at the target location, the playback priority among multiple target speech contents can be determined according to the order of priority weights. It can be understood that the target speech content with a larger priority weight has a higher playback priority. Furthermore, by sorting the multiple target speech contents according to their playback priority, a sequence of target speech contents is obtained.
[0103] In some implementations, to improve driving safety, the priority weights of corresponding voice content can be determined based on the risks present in environmental information. For example, the target voice content includes road condition voice content and climate voice content. Step (103.1) may include: identifying road and building information in the environmental information and determining a first risk coefficient corresponding to the road and building information; identifying climate information in the environmental information and determining a second risk coefficient corresponding to the climate information; and determining the priority weights of the road condition voice content and the climate voice content based on the first risk coefficient and the second risk coefficient, respectively.
[0104] Specifically, the system identifies road condition and climate information contained in the environmental information. For example, it collects environmental images of the vehicle at the target location and uses a trained recognition model to identify them. Specifically, it uses a lane structure recognition model to identify lane structure information and its category in the environmental image, and a climate information recognition model to identify climate information (such as visible weather conditions like rain or fog) and climate type in the environmental image. Then, it determines the corresponding risk coefficients, which can be determined by referring to a pre-defined risk coefficient table. For example, this risk coefficient table might be as follows:
[0105]
[0106] Therefore, after determining that the vehicle has identified environmental information at the target location, the corresponding risk coefficient can be found based on the specific identified environmental information. For example, the first risk coefficient corresponding to the lane and building information of the road conditions in the environmental information can be identified, and the second risk coefficient corresponding to the climate information can be determined, so as to determine the priority weight of the road condition voice content and the climate voice content according to the risk coefficient.
[0107] It should be noted that when determining priority weights, it is advisable to first determine whether the first risk coefficient and the second risk coefficient are the same. On the one hand, if the first risk coefficient and the second risk coefficient are different, the risk coefficient can be directly used as the priority weight value for road condition voice content and weather voice content. On the other hand, if the first risk coefficient and the second risk coefficient are equal, the risk coefficient can be weighted by combining the corresponding category weight coefficients to determine the priority weights corresponding to road condition voice content and weather voice content respectively.
[0108] For example, taking interchange lanes as road construction information, the first risk coefficient is 0.6. Taking "rain" as climate information, the second risk coefficient is 0.6. Since the first and second risk coefficients are the same, to determine the priority weights of traffic condition voice content and climate voice content respectively, a weighting is performed based on the category weight coefficient and the risk coefficient. Multiplying the first risk coefficient "0.6" by the category weight coefficient "0.9" for road construction results in a priority weight of 0.54 for traffic condition voice content associated with road construction information; while multiplying the second risk coefficient "0.6" by the category weight coefficient "0.8" for climate information results in a priority weight of 0.48 for traffic condition voice content associated with road construction information. This facilitates subsequent sorting of each target voice content according to its priority weight.
[0109] Using the above methods, the voice content can be filtered and sorted according to the vehicle's current driving control mode to generate a target voice content sequence, so that subsequent voice broadcasts can be performed according to the target voice content sequence, ensuring that the subsequently broadcast voice content is neat and orderly.
[0110] 104. Broadcast the target speech content in the target speech content sequence in sequence.
[0111] After obtaining the target speech content sequence with a broadcast order in this embodiment, each target speech content in the sequence can be broadcast sequentially, making the broadcast order between the target speech content neat and orderly. For example, taking the high-speed autonomous navigation driving mode in autonomous driving mode as an example, assuming the detailed broadcast mode corresponding to this driving mode allows multiple categories of speech content to be broadcast at the same time, such as road condition category, weather category, nearby tourist area information category, road sign category, etc., assuming that the broadcast priority of the above categories of speech content is sorted according to driving safety, the broadcast priority of the speech content is obtained as road condition category, weather category, road sign category, nearby tourist area information category, and each speech content is broadcast in sequence according to the broadcast priority.
[0112] In some implementations, although the identified environmental information may differ due to changes in the vehicle's position while moving, some environmental information may remain the same. To avoid repetitive playback of identical voice content, a playback interval can be set for each voice content to control the playback of identical voice content. Specifically, before the step "playing the target voice content in the target voice content sequence sequentially," the following may be included:
[0113] (A) Read the content of the played voice messages from the list of played voice messages;
[0114] (B) When the target speech content is identified to be consistent with the already broadcast speech content, determine the broadcast interval of the already broadcast speech content at the current time;
[0115] (C) Update the target speech content sequence according to the broadcast interval duration to obtain the updated target speech content sequence;
[0116] Then step 104, “broadcasting the target speech content in the target speech content sequence in sequence,” may include: broadcasting the target speech content in the updated target speech content sequence in sequence.
[0117] Understandably, after the vehicle terminal plays each voice message, it can generate a playback record in the list of played voice messages. This playback record can specifically include the text, information category, and voice attribute category of the corresponding voice message.
[0118] Specifically, when updating the target voice content sequence based on the broadcast interval, the order of the target voice content within the sequence can be updated. This involves identifying voice content whose broadcast interval is less than a preset threshold and reordering it to the last position in the target voice content sequence, resulting in the updated sequence. This maximizes the interval between already broadcast voice content, ensuring that content broadcast within a short period is rebroadcast after the maximum possible interval, thus improving the driver's or passengers' experience with navigation voice prompts.
[0119] Specifically, updating the target voice content sequence based on the broadcast interval can involve deleting previously broadcast target voice content. Specifically, the broadcast interval is compared with a preset threshold. When a broadcast interval shorter than the threshold is detected, the target voice content with a shorter interval is designated as pending management and deleted, resulting in an updated target voice content sequence. This allows each voice content to have a broadcast cooldown period, effectively preventing the same voice content from being broadcast frequently within a short period, thus improving the driver's or passengers' experience with navigation voice.
[0120] Using the above method, each target speech content in the target speech content sequence can be played sequentially according to the playback order relationship between multiple target speech content in the target speech content sequence, so that the playback order between the target speech content is neat and orderly.
[0121] By implementing any one or a combination of implementation methods in the embodiments of this application, the application scenario of navigation voice broadcasting can be realized.
[0122] As can be seen from the above, the embodiments of this application can determine the current driving control mode of the vehicle and the corresponding voice broadcast mode; obtain the environmental information corresponding to the current target location of the vehicle and obtain the voice content set corresponding to the environmental information; filter out the target voice content that conforms to the voice broadcast mode from the voice content set, and sort the target voice content according to the broadcast priority to obtain the target voice content sequence; and broadcast the target voice content in the target voice content sequence in sequence. Therefore, this solution can optimize the voice content that needs to be broadcast while the vehicle is in motion. Specifically, firstly, the current location and driving control mode of the vehicle are determined, and the voice broadcast mode under the current driving control mode is determined. Simultaneously, the set of voice content to be broadcast is determined based on the environmental information of the current location; then, the voice content in the voice content set is filtered according to the voice broadcast mode, and the playback priority of the filtered voice content is sorted; finally, the voice content is broadcast according to the sorted broadcast priority. This ensures that the navigation voice is orderly and organized during broadcast, which is beneficial for the driver to obtain important voice information and avoids complicated navigation voice affecting the passenger experience.
[0123] Based on the method described in the above embodiments, the following examples will provide further detailed explanations.
[0124] This application uses a navigation voice broadcasting device as an example to further describe the navigation voice broadcasting method provided in this application. Among them, Figure 3 This is a schematic flowchart of another step in the navigation voice broadcasting method provided in the embodiments of this application. Figure 4 This is a schematic diagram of the architecture of the navigation voice broadcasting device provided in the embodiments of this application. For ease of understanding, the embodiments of this application are combined with… Figure 3-4 Describe it.
[0125] In this embodiment, the description will focus on a navigation voice broadcasting device, which can be integrated into a computer device such as an in-vehicle terminal. When the processor on the in-vehicle terminal executes the program instructions corresponding to the data transmission method, the specific flow of the navigation voice broadcasting method is as follows:
[0126] 201. The vehicle terminal determines the vehicle's current target location and driving control mode.
[0127] The target location refers to the vehicle's real-time location information, which can be obtained from the vehicle's current location information through location services, such as determining it based on the Global Positioning System (GPS).
[0128] This driving control mode refers to the vehicle's driving mode, such as manual driving mode, driver assistance mode, and autonomous driving mode. Furthermore, driving modes can be further subdivided, such as adaptive cruise control, lane centering assist, and highway autonomous navigation driving. The specific driving control mode is determined based on the vehicle's driving control logic.
[0129] 202. The vehicle terminal determines the voice broadcast mode corresponding to the driving control mode.
[0130] Specifically, different voice broadcast modes can be set for different driving control modes, meaning different driving control modes correspond to different voice broadcast modes. The types, number, and order of voice messages broadcast under different voice broadcast modes may differ, allowing occupants to capture key voice information in the appropriate driving control mode.
[0131] The voice broadcast mode can be a way of playing navigation voice information, which is associated with the playable categories. That is, the voice broadcast mode limits the voice categories that are allowed to be broadcast under this mode.
[0132] 203. The vehicle terminal obtains environmental information corresponding to the current target location of the vehicle.
[0133] The environmental information can be the surrounding environment of the vehicle's location. This environmental information can include multiple categories of environmental information, such as but not limited to road condition information, climate information, road sign information, green belt information, and information about specific areas (such as scenic spots).
[0134] When acquiring environmental information, it can be obtained through relevant sensors, such as camera components or infrared sensor components, to collect road condition information, climate information, road sign information, green belt information, and information about specific areas (such as scenic spots).
[0135] In addition, high-precision electronic maps can be used to obtain road condition information, road sign information, green belt information, and information on specific areas (such as scenic spots), as well as climate information in conjunction with weather service applications. It should be noted that, to improve the accuracy of environmental information, camera components can be used in conjunction with high-precision maps and weather service applications to acquire environmental information.
[0136] 204. The vehicle terminal identifies multiple information categories from the environmental information and determines the voice attribute category to which each information category belongs.
[0137] This information category can be an environmental information category, that is, a classification of environmental information. For example, environmental information includes road condition information, weather information, road sign information, service area information before the tunnel, etc., which can be classified into categories such as road condition information, weather information, road sign information, and service area information.
[0138] The voice attribute category can be the voice category corresponding to the information category. For example, the traffic information category corresponds to the traffic voice category, the climate information category corresponds to the climate voice category, the road sign information category corresponds to the road sign voice category, and the specific area information category corresponds to the area introduction information category.
[0139] 205. The vehicle terminal selects the corresponding voice content to be processed based on the environmental information and voice attribute category, and constructs a set of voice content corresponding to the voice content to be processed.
[0140] For example, when obtaining the voice content to be processed based on environmental information and voice attribute categories, the voice content can be generated using voice templates. For instance, environmental information includes road condition information, and voice attribute categories include road condition voice categories. The specific process for obtaining the voice content to be processed is as follows: identify the road buildings and target building categories corresponding to the road condition information, and determine the target distance between the road buildings and the target location; query multiple road condition voice contents corresponding to the road condition voice categories from the voice content library, and extract the target building voice contents corresponding to the target building categories from the multiple road condition voice contents; extract building-type voice templates from the target building voice contents, and generate the voice content to be processed based on the target distance and the building-type voice templates.
[0141] For example, the corresponding speech content to be processed can also be obtained based on the semantic information in the environmental information. For instance, if the environmental information includes road sign information, then the speech attribute category includes the road sign speech category. The specific process for obtaining the speech content to be processed is as follows: identify the indicative semantics in the road sign information; query multiple road sign speech contents corresponding to the road sign speech category from the speech content library; match the target road sign speech contents corresponding to the indicative semantics from the multiple road sign speech contents, and determine the target road sign speech contents as the speech content to be processed.
[0142] 206. The vehicle terminal determines the playable category associated with the voice broadcast mode and selects the target voice content that matches the playable category from the voice content set.
[0143] Specifically, to avoid overwhelming passengers with complex voice prompts, the voice content in the audio feed can be filtered, specifically based on the vehicle's current driving control mode. Different driving control modes have different voice broadcast modes; therefore, the voice content is filtered according to these modes to ensure that the broadcasted audio content matches the target voice content of the currently playable category within the vehicle's current voice broadcast mode.
[0144] 207. The vehicle terminal sorts the target voice content according to the priority of broadcasting to obtain the target voice content sequence.
[0145] Specifically, after filtering the audio content in the audio content set, the priority of the target audio content to be played can be sorted so that the target audio content to be played has a sequence when played, so that the driver and passengers can obtain the more important audio information first.
[0146] To improve driving safety, priority weights can be determined based on the risks associated with environmental information, and multiple target voice contents to be broadcast can be prioritized according to these priority weights. The specific process can be as follows: determine the priority weight of each target voice content at the target location; determine the broadcast priority among the target voice contents based on their priority weights; and sort the target voice contents according to their broadcast priority to obtain a sequence of target voice contents.
[0147] 208. The vehicle-mounted terminal sequentially broadcasts the target audio content in the target audio content sequence. In this embodiment, after obtaining the target audio content sequence with a broadcast order, each target audio content in the sequence can be broadcast sequentially, ensuring a neat and orderly broadcast order among the target audio content.
[0148] For example, taking the high-speed autonomous navigation driving mode in autonomous driving mode as an example, suppose the corresponding detailed broadcast mode in this driving mode allows multiple categories of voice content to be broadcast at the same time, such as road condition category, weather category, nearby tourist area information category, road sign category, etc. Suppose that the broadcast priority of the above categories of voice content is sorted according to driving safety, and the broadcast priority of the voice content is obtained in the order of road condition category, weather category, road sign category, nearby tourist area information category, and each voice content is broadcast in order of this broadcast priority.
[0149] By performing steps 201-208 above, combined with... Figure 4 This can achieve the following scenarios:
[0150] In this embodiment of the application, the vehicle terminal may include, in terms of architecture, an environmental data module, a data processing module, a voice processing module, and a voice broadcasting module.
[0151] The environmental data module is used to acquire the absolute and relative position information of the current autonomous vehicle in real time, while simultaneously sensing and acquiring information about the surrounding road environment and the status of other vehicles. Specifically, it can receive data from sensing components, such as millimeter-wave radar, ultrasonic radar, lidar, and cameras, to perceive and preprocess data such as road information and the status of other vehicles in the vehicle's surrounding environment. Furthermore, it can also receive target location information from maps and navigation positioning.
[0152] The data processor is used to filter and determine the trigger signal for real-time navigation voice broadcasting based on upstream environmental data. Specifically, it performs further filtering and post-processing on the environmental data transmitted from the environmental data module, and interacts with the upstream environmental data module to process data in real time, providing timely feedback and transmission to other upstream and downstream modules.
[0153] The voice processing module is used to set the categories and priorities of navigation voice broadcast content, and to determine the relevant prompts to be broadcast based on the current driving behavior and driving mode, taking into account both autonomous driving safety and driving experience. In addition, the voice processor can receive environmental data information transmitted from the upstream environmental data module, and combine this information with the data processor's transmission to determine and sort the categories and order of navigation voice broadcast content, filtering out irrelevant information and low-priority voice broadcasts to avoid excessive and cluttered navigation voice broadcasts interfering with the passenger's driving experience.
[0154] The voice broadcast module is a downstream display device for the voice content, and it is not limited to view display devices or voice broadcast devices. It broadcasts the filtered and sorted navigation voice content.
[0155] As can be seen from the above, the embodiments of this application can determine the current driving control mode of the vehicle and the corresponding voice broadcast mode; obtain the environmental information corresponding to the current target location of the vehicle and obtain the voice content set corresponding to the environmental information; filter out the target voice content that conforms to the voice broadcast mode from the voice content set, and sort the target voice content according to the broadcast priority to obtain the target voice content sequence; and broadcast the target voice content in the target voice content sequence in sequence. Therefore, this solution can optimize the voice content that needs to be broadcast while the vehicle is in motion. Specifically, firstly, the current location and driving control mode of the vehicle are determined, and the voice broadcast mode under the current driving control mode is determined. Simultaneously, the set of voice content to be broadcast is determined based on the environmental information of the current location; then, the voice content in the voice content set is filtered according to the voice broadcast mode, and the playback priority of the filtered voice content is sorted; finally, the voice content is broadcast according to the sorted broadcast priority. This ensures that the navigation voice is orderly and organized during broadcast, which is beneficial for the driver to obtain important voice information and avoids complicated navigation voice affecting the passenger experience.
[0156] The implementation process or description of the above embodiments is equivalent to or similar to the description of the preceding embodiments. For details, please refer to the description of the foregoing embodiments, which will not be repeated here.
[0157] To better implement the above methods, this application also provides a navigation voice broadcasting device, which can be integrated into computer equipment, such as in-vehicle terminals and other computer equipment.
[0158] For example, such as Figure 5 As shown, the navigation voice broadcasting device may include a determining unit 501, an acquiring unit 502, a sorting unit 503, and a broadcasting unit 504.
[0159] The determining unit 501 is used to determine the current driving control mode of the vehicle and the corresponding voice broadcast mode of the driving control mode.
[0160] The acquisition unit 502 is used to acquire environmental information corresponding to the current target location of the vehicle, and to acquire a set of voice content corresponding to the environmental information;
[0161] The sorting unit 503 is used to filter out target speech content that conforms to the speech broadcasting mode from the speech content set, and sort the target speech content according to the broadcasting priority to obtain the target speech content sequence.
[0162] The broadcasting unit 504 is used to broadcast the target speech content in the target speech content sequence in sequence.
[0163] In some embodiments, the acquisition unit 502 is further configured to: identify multiple information categories from the environmental information; determine the voice attribute category to which each information category belongs; select the corresponding voice content to be processed according to the environmental information and the voice attribute category, and construct a set of voice content corresponding to the voice content to be processed.
[0164] In some embodiments, the acquisition unit 502 is further configured to: identify the road buildings and the corresponding target building categories corresponding to the road condition information, and determine the target distance between the road buildings and the target location; query multiple road condition voice contents corresponding to the road condition voice categories from the voice content library, and extract the target building voice contents corresponding to the target building categories from the multiple road condition voice contents; extract building-type voice templates from the target building voice contents, and generate voice contents to be processed based on the target distance and the building-type voice templates.
[0165] In some embodiments, the environmental information includes road sign information, and the voice attribute category includes road sign voice category. The acquisition unit 502 is further configured to: identify the indicative semantics in the road sign information; query multiple road sign voice contents corresponding to the road sign voice category from the voice content library; match the target road sign voice contents corresponding to the indicative semantics from the multiple road sign voice contents, and determine the target road sign voice contents as the voice contents to be processed.
[0166] In some implementations, the sorting unit 503 is further configured to: determine the playable category associated with the voice broadcasting mode; and filter out target voice content that matches the playable category from the voice content set.
[0167] In some embodiments, the sorting unit 503 is further configured to: determine the priority weight of each target speech content at the target location; determine the broadcast priority among the target speech content based on the priority weight of each target speech content; and sort the target speech content according to the broadcast priority to obtain a target speech content sequence.
[0168] In some embodiments, the target voice content includes road condition voice content and climate voice content. The sorting unit 503 is further configured to: identify road and building information in the environmental information and determine a first risk coefficient corresponding to the road and building information; identify climate information in the environmental information and determine a second risk coefficient corresponding to the climate information; and determine the priority weights of the road condition voice content and the climate voice content according to the first risk coefficient and the second risk coefficient, respectively.
[0169] In some embodiments, the navigation voice broadcasting device further includes an updating unit, configured to: read the broadcast voice content from the broadcast voice list; when the target voice content is identified to be consistent with the broadcast voice content, determine the broadcast interval duration of the broadcast voice content at the current time; update the target voice content sequence according to the broadcast interval duration to obtain the updated target voice content sequence;
[0170] Then the broadcasting unit 504 is also used to broadcast the target speech content in the updated target speech content sequence in sequence.
[0171] As can be seen from the above, the embodiments of this application can optimize the voice content that needs to be broadcast while the vehicle is in motion. Specifically, firstly, the current position and driving control mode of the vehicle are determined, and the voice broadcast mode under the current driving control mode is determined. At the same time, the set of voice content to be broadcast is determined based on the environmental information of the current position. Then, the voice content in the set of voice content is filtered according to the voice broadcast mode, and the playback priority of the filtered voice content is sorted. Finally, the voice content is broadcast according to the sorted playback priority. In this way, the navigation voice can be broadcast in a neat and orderly manner, which is beneficial for the driver to obtain important voice information and avoids the impact of complicated navigation voice on the passenger experience.
[0172] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0173] This application also provides a computer device, such as... Figure 6 As shown, it illustrates a structural schematic diagram of the computer device involved in the embodiments of this application, specifically:
[0174] The computer device may include components such as a processor 601 with one or more processing cores, a memory 602 with one or more computer-readable storage media, a power supply 603, and an input unit 604. Those skilled in the art will understand that... Figure 6 The computer device structure shown does not constitute a limitation on the computer device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:
[0175] The processor 601 is the control center of the computer device. It connects various parts of the computer device via various interfaces and lines, and performs various functions and processes data by running or executing software programs and / or modules stored in the memory 602, and by calling data stored in the memory 602, thereby providing overall monitoring of the computer device. Optionally, the processor 601 may include one or more processing cores; preferably, the processor 601 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 601.
[0176] The memory 602 can be used to store software programs and modules. The processor 601 executes various functional applications and navigation voice broadcasts by running the software programs and modules stored in the memory 602. The memory 602 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 602 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 602 may also include a memory controller to provide the processor 601 with access to the memory 602.
[0177] The computer device also includes a power supply 603 that supplies power to the various components. Preferably, the power supply 603 can be logically connected to the processor 601 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 603 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0178] The computer device may also include an input unit 604, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0179] Although not shown, the computer device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 601 in the computer device loads the executable files corresponding to the processes of one or more application programs into the memory 602 according to the following instructions, and the processor 601 runs the application programs stored in the memory 602 to realize various functions, as follows:
[0180] Determine the vehicle's current driving control mode and the corresponding voice broadcast mode; obtain the environmental information corresponding to the vehicle's current target location and the corresponding set of voice content; filter out the target voice content that matches the voice broadcast mode from the set of voice content, sort the target voice content by broadcast priority, and obtain the target voice content sequence; broadcast the target voice content in the target voice content sequence in sequence.
[0181] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0182] As can be seen from the above, the embodiments of this application can optimize the voice content that needs to be broadcast while the vehicle is in motion. Specifically, firstly, the current position and driving control mode of the vehicle are determined, and the voice broadcast mode under the current driving control mode is determined. At the same time, the set of voice content to be broadcast is determined based on the environmental information of the current position. Then, the voice content in the voice content set is filtered according to the voice broadcast mode, and the playback priority of the filtered voice content is sorted. Finally, the voice content is broadcast according to the sorted playback priority. In this way, the navigation voice can be broadcast in a neat and orderly manner, which is beneficial for the driver to obtain important voice information and avoids the impact of complicated navigation voice on the passenger experience.
[0183] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0184] Therefore, embodiments of this application provide a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the navigation voice broadcasting methods provided in embodiments of this application. For example, the instructions can execute the following steps:
[0185] Determine the vehicle's current driving control mode and the corresponding voice broadcast mode; obtain the environmental information corresponding to the vehicle's current target location and the corresponding set of voice content; filter out the target voice content that matches the voice broadcast mode from the set of voice content, sort the target voice content by broadcast priority, and obtain the target voice content sequence; broadcast the target voice content in the target voice content sequence in sequence.
[0186] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0187] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0188] This application also provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the navigation voice broadcasting method provided in the various optional implementations of the above embodiments.
[0189] Since the instructions stored in the computer-readable storage medium can execute the steps in any of the navigation voice broadcasting methods provided in the embodiments of this application, the beneficial effects that any of the navigation voice broadcasting methods provided in the embodiments of this application can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.
[0190] The above provides a detailed description of a navigation voice broadcasting method, apparatus, device, and computer-readable storage medium provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A navigation voice broadcasting method, characterized in that, include: Determine the current driving control mode of the vehicle, and determine the voice broadcast mode corresponding to the driving control mode; Obtain environmental information corresponding to the current target location of the vehicle, and obtain a set of voice content corresponding to the environmental information; Select target audio content that matches the audio playback mode from the audio content set, and determine the priority weight of each target audio content at the target position; Based on the priority weight of each target speech content, the playback priority among the target speech content is determined; the target speech content is sorted according to the playback priority to obtain a target speech content sequence; The target speech content in the target speech content sequence is broadcast sequentially; The target speech content includes road condition speech content and weather speech content. Determining the priority weight of each target speech content at the target location includes: Identify road and building information in the environmental information, and determine the first risk coefficient corresponding to the road and building information; Identify the climate information in the environmental information and determine the second risk coefficient corresponding to the climate information; When the first risk coefficient and the second risk coefficient are equal, the priority weight of the road condition voice content is determined by combining the category weight corresponding to the road construction information and the first risk coefficient, and the priority weight of the climate voice content is determined by combining the type weight corresponding to the climate information and the second risk coefficient.
2. The method according to claim 1, characterized in that, The acquisition of the voice content set corresponding to the environmental information includes: Multiple information categories are identified from the environmental information; Determine the voice attribute category to which each information category belongs; Based on the environmental information and voice attribute category, select the corresponding voice content to be processed, and construct the voice content set corresponding to the voice content to be processed.
3. The method according to claim 2, characterized in that, The environmental information includes road condition information, and the voice attribute category includes road condition voice category. Selecting the corresponding voice content based on the environmental information and voice attribute category includes: Identify the road structures and target building categories corresponding to the road condition information, and determine the target distance between the road structures and the target location; Query multiple traffic condition voice contents corresponding to the traffic condition voice category from the voice content library, and extract the target building voice contents corresponding to the target building category from the multiple traffic condition voice contents; Extract architectural speech templates from the target architectural speech content, and generate speech content to be processed based on the target distance and architectural speech templates.
4. The method according to claim 2, characterized in that, The environmental information includes road sign information, and the voice attribute category includes road sign voice categories. Selecting the corresponding voice content based on the environmental information and the voice attribute category includes: Identify the indicative semantics in the road sign information; Query the voice content library for multiple road sign voice contents corresponding to the road sign voice category; The target road sign voice content corresponding to the indication semantics is matched from the multiple road sign voice contents, and the target road sign voice content is determined as the voice content to be processed.
5. The method according to claim 1, characterized in that, The step of filtering target voice content that matches the voice broadcasting pattern from the voice content set includes: Determine the playable category associated with the voice broadcasting mode; Select target audio content that matches the playable category from the audio content set.
6. The method according to claim 1, characterized in that, Before sequentially broadcasting the target speech content in the target speech content sequence, the method further includes: Read the content of the played voice messages from the list of played voice messages; When the target speech content is identified to be consistent with the already played speech content, the duration of the playback interval of the already played speech content at the current time is determined; The target speech content sequence is updated according to the broadcast interval duration to obtain the updated target speech content sequence; The step of sequentially broadcasting the target speech content in the target speech content sequence includes: The target speech content in the updated target speech content sequence is broadcast sequentially.
7. A navigation voice broadcasting device, characterized in that, include: The determining unit is used to determine the current driving control mode of the vehicle and the voice broadcast mode corresponding to the driving control mode. The acquisition unit is used to acquire environmental information corresponding to the current target location of the vehicle, and to acquire a set of voice content corresponding to the environmental information; The sorting unit is used to filter out target audio content that conforms to the audio broadcasting mode from the audio content set, and determine the priority weight of each target audio content at the target position. Based on the priority weight of each target speech content, the playback priority among the target speech content is determined; the target speech content is sorted according to the playback priority to obtain a target speech content sequence; A broadcasting unit is used to broadcast the target speech content in the target speech content sequence in sequence. Specifically, the sorting unit is used to identify road and building information in the environmental information and determine a first risk coefficient corresponding to the road and building information; identify climate information in the environmental information and determine a second risk coefficient corresponding to the climate information; when the first risk coefficient and the second risk coefficient are equal, determine the priority weight of the road condition voice content by combining the category weight corresponding to the road and building information and the first risk coefficient, and determine the priority weight of the climate voice content by combining the type weight corresponding to the climate information and the second risk coefficient.
8. A computer device, characterized in that, It includes a processor and a memory, the memory storing a computer program, and the processor running the computer program in the memory to implement the steps of the navigation voice broadcasting method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium is computer-readable and stores a plurality of instructions adapted for loading by a processor to perform the steps of the navigation voice broadcasting method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Automatic traffic sign identification and prompting method and device and automobile data recorder
CN106204800A
Navigation broadcasting method, device and equipment and storage medium
CN114201567A
Vehicle control device provided in vehicle and vehicle control method
US20210139036A1