Reinforcement Learning Method, Electronic Device and Program Product for Speech Announcement Model
Through the reinforcement learning method of the voice broadcast model, the voice navigation system is trained to predict appropriate voice broadcast content, solving the problem that the existing technology cannot adapt to diversified navigation scenarios, and achieving more accurate and refined voice broadcasts.
Patent Information
- Application Number
- CN202210281103.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-21
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-03-21
AI Technical Summary
The existing voice navigation system cannot adapt to various broadcast scenarios encountered during navigation, and cannot enumerate all broadcast scenarios in advance, and cannot meet the user's refined voice broadcast needs.
The reinforcement learning method of the voice broadcast model is adopted. By obtaining navigation-related information and voice broadcast sample content in the sample navigation trajectory, the voice broadcast model is trained to predict the appropriate voice broadcast content, and the target reward value is calculated by matching the predicted content and the navigation elements in the sample content, and intensive training is carried out.
It realizes that the voice navigation system can adapt to different navigation scenarios, output more accurate and detailed voice broadcast content, and meet users' refined needs.
Smart Images

Figure CN114925180B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of location-based services, and particularly to a reinforcement learning method for a voice broadcast model, an electronic device, and a program product. Background Art
[0002] With the development of Internet technology, people's travel increasingly relies on location-based service systems. Location-based services include navigation, path planning, map rendering, etc. Navigation services are used to guide users during vehicle driving, such as by visually or audibly prompting users to perform corresponding navigation actions.
[0003] With the diversification of travel demands, users' requirements for in-trip navigation broadcast (TBT) are also more refined. In the existing technology, a voice navigation system usually pre-configures voice broadcast scenarios and configures corresponding voice broadcast scripts under the corresponding voice broadcast scenarios, so as to broadcast the corresponding voice broadcast scripts based on the pre-configured broadcast scenarios during navigation; however, pre-configuring voice broadcast scripts makes it impossible to adapt to various broadcast scenarios encountered during navigation, and it is also impossible to enumerate all broadcast scenarios in advance. Therefore, the voice navigation system in the existing technology cannot meet users' refined voice broadcast requirements. Summary of the Invention
[0004] Embodiments of the present disclosure provide a reinforcement learning method for a voice broadcast model, an electronic device, and a program product.
[0005] In a first aspect, an embodiment of the present disclosure provides a reinforcement learning method for a voice broadcast model, which includes:
[0006] Obtain sample data; the sample data includes navigation-related information at a sample position in a sample navigation trajectory and voice broadcast sample content output at the sample position;
[0007] Use the navigation-related information at the current sample position in the sample navigation trajectory as the current state and input it into the voice broadcast model to obtain action information in the current state; the action information includes predicted voice broadcast content at the current sample position;
[0008] Based on a matching result between the predicted voice broadcast content and the voice broadcast sample content output at the current sample position, calculate a target reward value; the matching result includes a matching result between predicted navigation elements in the predicted voice broadcast content and sample navigation elements in the voice broadcast sample content;
[0009] Perform reinforcement training on the voice broadcast model based on the target reward value.
[0010] Further, calculating a target reward value based on a matching result between the predicted content of the voice broadcast and the sample content of the voice broadcast output at the current sample position, includes:
[0011] Determine the predicted navigation elements included in the predicted content of the voice broadcast, and the broadcast order of the predicted navigation elements in the predicted content of the voice broadcast;
[0012] Perform element matching between the predicted navigation elements and the sample navigation elements in the sample content of the voice broadcast;
[0013] Match the broadcast order of the predicted navigation elements in the predicted content of the voice broadcast with the broadcast order of the sample navigation elements in the sample content of the voice broadcast;
[0014] Calculate the target reward value based on the element matching result and the order matching result.
[0015] Further, calculating the target reward value based on the element matching result and the order matching result, includes:
[0016] Determine a positive reward value based on the number of matching predicted navigation elements and sample navigation elements;
[0017] Determine a first negative reward value based on the number of non-matching predicted navigation elements and sample navigation elements;
[0018] Determine a second negative reward value based on the number of non-matching broadcast orders of the predicted navigation elements in the predicted content of the voice broadcast and the sample navigation elements in the sample content of the voice broadcast;
[0019] Determine the target reward value based on the positive reward value, the first negative reward value, and the second negative reward value.
[0020] Further, the obtaining of the sample data includes:
[0021] Obtain a sample navigation trajectory;
[0022] Based on a preset online policy, determine a sample position and a combination of sample navigation elements at the sample position for the sample navigation trajectory;
[0023] Use a value function to calculate the value of the sample navigation elements and the value of the reference navigation elements; wherein, the reference navigation elements are the navigation elements corresponding to not performing a voice broadcast;
[0024] Sort the sample navigation elements with a value greater than the value of the reference navigation elements in order of value magnitude to form the sample content of the voice broadcast at the sample position.
[0025] Further, the action information further includes a predicted voice announcement timing; the method further includes:
[0026] Updating the target reward value based on a matching result of whether a predicted sample position corresponding to the predicted voice announcement timing matches the current sample position.
[0027] Further, the action information further includes a voice announcement timing; the method further includes:
[0028] Determining state transition information based on a voice announcement duration of the predicted voice announcement content and the voice announcement timing; the state transition information includes a next sample position corresponding to a next state and navigation-related information corresponding to the next sample position;
[0029] Performing a next round of reinforcement training on the voice announcement model by using the state transition information.
[0030] In a second aspect, an embodiment of the present disclosure provides a navigation voice announcement method, which includes:
[0031] Obtaining navigation-related information at a current navigation position;
[0032] Inputting the navigation-related information into a pre-trained voice announcement model to obtain a navigation element for voice announcement and a voice announcement timing;
[0033] Assembling the navigation elements to generate a voice announcement script;
[0034] Outputting the voice announcement script at the voice announcement timing.
[0035] In a third aspect, an embodiment of the present disclosure provides a location-based service providing method, where the location-based service providing method trains a voice announcement model by using the method described in the first aspect, and provides a location-based service for a service recipient by using the trained voice announcement model, and the location-based service includes one or more of navigation, map rendering, and route planning.
[0036] In a fourth aspect, an embodiment of the present invention provides a reinforcement learning device for a voice announcement model, including:
[0037] An obtaining module, configured to obtain sample data; the sample data includes navigation-related information of a sample position in a sample navigation trajectory and a voice announcement sample content output at the sample position;
[0038] An input module, configured to input the navigation-related information of a current sample position in the sample navigation trajectory into the voice announcement model as a current state to obtain action information in the current state; the action information includes a predicted voice announcement content at the current sample position;
[0039] A calculation module, configured to calculate a target reward value based on a matching result between the predicted content of the voice broadcast and the content of the voice broadcast sample output at the current sample position; the matching result includes a matching result between a predicted navigation element in the predicted content of the voice broadcast and a sample navigation element in the content of the voice broadcast sample.
[0040] A first training module, configured to perform reinforcement training on the voice broadcast model based on the target reward value.
[0041] The above functions can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.
[0042] In a possible design, the structure of the above device includes a memory and a processor. The memory is used to store one or more computer instructions that support the above device to execute the corresponding method. The processor is configured to execute the computer instructions stored in the memory. The above device may further include a communication interface for the above device to communicate with other devices or communication networks.
[0043] In a fifth aspect, an embodiment of the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored on the memory. Wherein, the processor executes the computer program to implement the method described in any of the above aspects.
[0044] In a sixth aspect, an embodiment of the present disclosure provides a computer-readable storage medium for storing computer instructions used by any of the above devices. When the computer instructions are executed by a processor, they are used to implement the method described in any of the above aspects.
[0045] In a seventh aspect, an embodiment of the present disclosure provides a computer program product, which includes computer instructions. When the computer instructions are executed by a processor, they are used to implement the method described in any of the above aspects.
[0046] The technical solution provided by the embodiment of the present disclosure may include the following beneficial effects:
[0047] In the embodiments of the present disclosure, during the reinforcement learning process of the voice broadcast model, navigation-related information at the sample position in the sample navigation trajectory and the voice broadcast sample content output at the sample position are obtained and used as sample data. Furthermore, during the process of training the voice broadcast model based on the sample data, the navigation-related information at the current sample position is used as the state data of the current state and input into the voice broadcast model. The voice broadcast model outputs the action information to be executed in the current state based on the state data of the current state. The action information may include the predicted voice broadcast content to be output in the current state. The predicted voice broadcast content may include one or more predicted navigation elements, or the predicted voice broadcast content is empty, indicating that no voice broadcast is performed in the current state. By matching the predicted voice broadcast content with the voice broadcast sample content at the current sample position for the navigation elements, a matching result is obtained. Then, a target reward value is calculated based on the matching result, and the voice broadcast model is trained by reinforcement based on the target reward value. In the embodiments of the present disclosure, the voice broadcast content is split into a combination of navigation elements, and then the target reward value is calculated based on whether the combination of navigation elements is more matched with the true value (i.e., the navigation elements in the voice broadcast sample content). This fine-grained reinforcement learning method can adapt to different voice broadcast scenarios during navigation without the need to pre-configure the voice broadcast scenario and voice broadcast terms in advance. Therefore, the prediction result of the trained voice broadcast model can be more accurate.
[0048] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In conjunction with the accompanying drawings, through the following detailed description of non-limiting embodiments, other features, objects, and advantages of the present disclosure will become more apparent. In the drawings:
[0050] Figure 1 The flowchart of the reinforcement learning method for the voice broadcast model according to an embodiment of the present disclosure is shown;
[0051] FIG. 2(a) and FIG. 2(b) show the formation process of the matching element set and the schematic diagram of the calculation of the target reward value according to an embodiment of the present disclosure;
[0052] Figure 3 The flowchart of the navigation voice broadcast method according to an embodiment of the present disclosure is shown;
[0053] Figure 4 The schematic diagram of the voice navigation scenario according to an embodiment of the present disclosure is shown;
[0054] Figure 5A structural block diagram of a reinforcement learning device for a voice broadcast model according to an embodiment of the present disclosure is shown;
[0055] Figure 6 It is a schematic structural diagram of an electronic device suitable for implementing a navigation voice broadcast method according to an embodiment of the present disclosure. Specific embodiments
[0056] Hereinafter, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily implement them. In addition, for clarity, parts irrelevant to the description of the exemplary embodiments are omitted in the drawings.
[0057] In the present disclosure, it should be understood that terms such as "including" or "having" are intended to indicate the presence of features, numbers, steps, actions, components, parts, or combinations thereof disclosed in this specification, and do not exclude the possibility of the presence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0058] In addition, it should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0059] The details of the embodiments of the present disclosure will be introduced in detail below through specific embodiments.
[0060] Figure 1 A flowchart of a reinforcement learning method for a voice broadcast model according to an embodiment of the present disclosure is shown. As Figure 1 shown, the reinforcement learning method for the voice broadcast model includes the following steps:
[0061] In step S101, sample data is obtained; the sample data includes navigation-related information at a sample position in a sample navigation trajectory and voice broadcast sample content output at the sample position;
[0062] In step S102, the navigation-related information at the current sample position in the sample navigation trajectory is input as the current state into the voice broadcast model to obtain action information in the current state; the action information includes voice broadcast prediction content at the current sample position;
[0063] In step S103, based on a matching result between the voice broadcast prediction content and the voice broadcast sample content output at the current sample position, a target reward value is calculated; the matching result includes a matching result between a predicted navigation element in the voice broadcast prediction content and a sample navigation element in the voice broadcast sample content;
[0064] In step S104, the voice broadcast model is strengthened and trained based on the target reward value.
[0065] In this embodiment, the reinforcement learning method of the voice broadcast model can be executed on the server. During a navigation process, the navigation terminal can perform multiple navigation voice broadcasts, such as "There is a red light running and illegal photographing 100 meters ahead". During the process of strengthening and training the voice broadcast model, sample data can be collected first. The sample data can include navigation-related information at each sample position in the sample navigation trajectory and the voice broadcast sample content output at each sample position in the sample navigation trajectory. The sample navigation trajectory can be a historical navigation trajectory actually generated in the navigation server or a trajectory generated based on a certain strategy.
[0066] In some embodiments, the voice broadcast sample content can include one or more voice broadcast phrases. A voice broadcast phrase can be a complete navigation semantic phrase, such as "Slow down and turn left at the traffic light ahead". A voice broadcast phrase can include one or more navigation elements. A navigation element can be information that plays a key guiding role in driving actions during navigation. For example, the key information "100 meters", "traffic light intersection", "turn right" in the voice broadcast phrase "Please note to slow down and turn right at the traffic light intersection 100 meters ahead" can be understood as navigation elements.
[0067] In other embodiments, the voice broadcast sample content can also be a combination of one or more sample navigation elements.
[0068] Navigation elements can be divided into main actions and auxiliary information. The main action can be the turning action of the object being navigated at the intersection in the sample navigation trajectory, which can be the key turning information directly affecting the navigation trajectory. The auxiliary information can be navigation instructions and reminders other than the main action, such as "There is a speed measurement photograph ahead, please note to slow down".
[0069] The sample navigation trajectory can be composed of multiple temporally continuous trajectory points. The sample position can be understood as the position in the historical navigation trajectory where the voice broadcast sample content is output. The voice broadcast sample content can be obtained based on historical data or calculated based on a certain strategy and is relatively accurate voice broadcast content. During the reinforcement learning process of the voice broadcast model, the voice broadcast sample content can be understood as the true value of the voice broadcast content; and the voice broadcast prediction content output by the voice broadcast model can be understood as the predicted value of the voice broadcast content.
[0070] That is to say, the sample data collected in the embodiments of the present disclosure may include navigation-related information at the sample positions in the sample navigation trajectory and the voice broadcast sample content output at the sample positions. One sample data may include navigation-related information at multiple sample positions in one sample navigation trajectory and the corresponding voice broadcast sample content. To complete the training of the voice broadcast model, the embodiments of the present disclosure may collect multiple sample navigation trajectories and then process them to obtain sample data.
[0071] Reinforcement Learning (RL), also known as re-inforcement learning, evaluation learning, or enhancement learning, is one of the paradigms and methodologies of machine learning, used to describe and solve the problem that an agent maximizes the reward value or achieves a specific goal through learning strategies during the interaction with the environment.
[0072] During the reinforcement learning process, the reward value represents the reward value obtained after selecting a certain action in the current state. The reinforcement learning process is based on the assumed reward value; the action comes from the action space. Each time the agent is in a certain state, based on the reward value of the previous state, it determines which action to execute in the current state to obtain a higher reward value.
[0073] The purpose of using reinforcement learning to optimize the voice broadcast content is to select the appropriate combination of navigation elements at the appropriate spatio-temporal positions. However, the existing reinforcement learning methods are usually designed for the broadcast scenario, and the explicit benefits of the broadcast scenario are difficult to quantify, resulting in difficulties in designing the reward value; at the same time, the time position, space position, and the state of the navigated object itself are all closely related to the broadcast content, and the corresponding state space and state transition are also difficult to construct. Therefore, to solve the above technical problems, the embodiments of the present disclosure propose a reinforcement learning modeling method that unifies the reward value and the state space.
[0074] In the embodiments of the present disclosure, the reward value is decomposed into each navigation element (such as action information, lane information, electronic eye information, etc.) in the voice broadcast content.
[0075] In the training process of the voice broadcast model according to the embodiments of the present disclosure, the voice broadcast sample content in the collected sample data is regarded as a combination of multiple sample navigation elements (Elements). Therefore, all or part of the navigation elements can be arbitrarily combined to cover all possible broadcast expressions. Any combination of navigation elements can be regarded as an action (Action). In addition, a combination without any navigation elements (that is, without voice broadcast) can also be added as an action to the action space.
[0076] Therefore, in the embodiments of the present disclosure, the action (Action) can be designed to broadcast a set single navigation element (Element) or a combination of navigation elements (Elements). In addition, it also includes the action of not broadcasting. The reward value (Reward) is designed to be related to the number of navigation elements (Elements) appearing in the action (Action).
[0077] The state (State) is designed to be the navigation-related information of the current position, which may include but is not limited to the information that can be recalled in the subsequent road segment sequence to pass through (such as the type of recalled navigation element, the distance between the current position and the recalled navigation element), the speed information at the current position, the acceleration and deceleration information, the distance information between the voice broadcast point and a certain navigation element, the time interval information since the last voice broadcast, etc.
[0078] The state (State) transition is designed as follows: after selecting a certain action (Action), after delaying the broadcast timing of the voice broadcast expression corresponding to the selected action (Action) by a duration of t, the navigation-related information at the position reached by the navigated object is used as the next state; for the action (Action) of "not broadcasting", the navigation-related information at the position after delaying the broadcast timing by a duration of Δt can be used as the next state after the state transition. The duration of Δt can be a pre-set duration, such as 1 second.
[0079] Based on the above design, when performing reinforcement learning on the voice broadcast model, any sample position in the collected sample data can be used as the current sample position (starting from the first sample position in a sample navigation trajectory), and the navigation-related information at the current sample position is used as the current state and input into the voice broadcast model. The voice broadcast model outputs the action information to be selected in the current state. The action information may include but is not limited to the voice broadcast prediction content at the current sample position. The voice broadcast prediction content includes one or more predicted navigation elements. Of course, it can be understood that if there is no voice broadcast here, the action information output by the voice broadcast model can be empty. It should be noted that not considering the action information of not broadcasting does not affect the implementation of the solution of the embodiments of the present disclosure.
[0080] In some embodiments, the speech broadcast model can be a model with any structure. For example, it can be a model with a multi-layer perceptron (MLP) structure.
[0081] Match the speech broadcast prediction content output by the speech broadcast model with the speech broadcast sample content at the current sample position, and obtain the target reward value of the speech broadcast prediction content output by the speech broadcast model based on the matching result. As described above, in the embodiments of the present disclosure, the target reward value is set to be related to the navigation element. Therefore, matching the speech broadcast prediction content with the speech broadcast sample content can be understood as matching the predicted navigation element in the speech broadcast prediction content with the sample navigation element in the speech broadcast sample content. Therefore, the matching result is also the matching result between the predicted navigation element and the sample navigation element.
[0082] It can be understood that the higher the target reward value, the higher the accuracy of the result output by the speech broadcast model. If the target reward value is lower, the accuracy of the result output by the speech broadcast model is lower. Therefore, based on the target reward value, the speech broadcast model is strengthened and trained, so that the speech broadcast prediction content output by the speech broadcast model in the next training can be closer to the speech broadcast sample content, and thus a higher target reward value can be obtained. It can be understood that the specific process of strengthening and training the speech broadcast model based on the target reward value depends on the structure of the speech broadcast model. For example, when the speech broadcast model is a multi-layer perceptron (MLP), the target reward value can be used to construct a loss function, and then the parameters on the neural network elements in the multi-layer perceptron are adjusted based on the loss function, that is, the model parameters are adjusted.
[0083] In the embodiments of the present disclosure, during the reinforcement learning process of the voice broadcast model, navigation-related information at the sample position in the sample navigation trajectory and the voice broadcast sample content output at the sample position are obtained and used as sample data. Furthermore, during the process of training the voice broadcast model based on the sample data, the navigation-related information at the current sample position is input into the voice broadcast model as the state data of the current state. The voice broadcast model outputs the action information to be executed in the current state based on the state data of the current state. The action information may include the predicted voice broadcast content to be output in the current state. The predicted voice broadcast content may include one or more predicted navigation elements, or the predicted voice broadcast content is empty, indicating that no voice broadcast is performed in the current state. By matching the predicted voice broadcast content with the voice broadcast sample content at the current sample position for the navigation elements, a matching result is obtained. Then, based on the matching result, a target reward value is calculated, and the voice broadcast model is trained using the target reward value. In the embodiments of the present disclosure, the voice broadcast content is split into a combination of navigation elements, and then the target reward value is calculated based on whether the combination of navigation elements is more matched with the ground truth (i.e., the navigation elements in the voice broadcast sample content). This fine-grained reinforcement learning method can adapt to different voice broadcast scenarios during navigation without the need to pre-configure the voice broadcast scenarios and voice broadcast scripts in advance. Therefore, the prediction results of the trained voice broadcast model can be more accurate.
[0084] In an alternative implementation of this embodiment, step S103, that is, the step of calculating the target reward value based on the matching result between the predicted voice broadcast content and the voice broadcast sample content output at the current sample position, further includes the following steps:
[0085] Determine the predicted navigation elements included in the predicted voice broadcast content and the broadcast order of the predicted navigation elements in the predicted voice broadcast content;
[0086] Perform element matching between the predicted navigation elements and the sample navigation elements in the voice broadcast sample content;
[0087] Match the broadcast order of the predicted navigation elements in the predicted voice broadcast content with the broadcast order of the sample navigation elements in the voice broadcast sample content;
[0088] Calculate the target reward value based on the element matching result and the order matching result.
[0089] In this optional implementation, the voice announcement sample content can be a complete voice announcement script or a combination of sample navigation elements. When the voice announcement sample content is a complete voice announcement script, it can be pre-split into a combination form of sample navigation elements and arranged in the announcement order of the sample navigation elements in the voice announcement sample content to form a sample navigation element set.
[0090] In some embodiments, the voice announcement prediction content output by the voice announcement model can be prediction navigation elements arranged in the announcement order or a complete announcement script assembled from the prediction navigation elements according to a script template. Therefore, it is possible to first determine which prediction navigation elements are included in the voice announcement prediction content output by the voice announcement model and the announcement order of these prediction navigation elements.
[0091] After that, the prediction navigation elements are matched with the sample navigation elements, and the matching result can be determined based on the number of corresponding and consistent navigation elements. It should be noted that the matching between navigation elements can be independent of the announcement order.
[0092] In addition, the announcement orders of the matched prediction navigation elements and sample navigation elements can be matched in order, and the matching result can be determined based on whether the announcement orders are consistent.
[0093] Based on the two matching results, the target reward value can be comprehensively obtained. For example, if the number of element matches is large and the announcement order is consistent, the target reward value is higher; if the number of element matches is small and the announcement order is inconsistent, the target reward value is lower.
[0094] That is to say, the matching process between the voice announcement prediction content and the voice announcement sample content is divided into two parts: element matching and order matching. Element matching refers to matching the prediction navigation elements in the voice announcement prediction content with the sample navigation elements in the matching element set. The more the number of matched elements, the better the matching result. Order matching refers to matching the announcement order of the prediction navigation elements in the voice announcement prediction content with the order of the sample navigation elements matched in the matching element set. The more the order matches are consistent, the better the matching result.
[0095] Finally, the target reward value can be calculated based on the results of element matching and order matching.
[0096] In an optional implementation of this embodiment, the step of calculating the target reward value based on the element matching result and the order matching result further includes the following steps:
[0097] Determine a positive reward value based on the number of matches between the prediction navigation elements and the sample navigation elements;
[0098] Determine a first negative reward value based on the number of mismatches between the predicted navigation elements and the sample navigation elements;
[0099] Determine a second negative reward value based on the number of mismatches between the broadcast order of the predicted navigation elements in the predicted voice broadcast content and the broadcast order of the sample navigation elements in the sample voice broadcast content;
[0100] Determine the target reward value based on the positive reward value, the first negative reward value, and the second negative reward value.
[0101] In this optional implementation, the more the number of predicted navigation elements that match the sample navigation elements in the matching element set, the higher the target reward value can be. Therefore, a positive reward value can be obtained based on the number of matches. For example, if one of the predicted navigation elements matches a sample navigation element in the matching element set, the target reward value can be increased by 1 point.
[0102] And the more the number of predicted navigation elements that do not match the sample navigation elements in the matching element set, the lower the target reward value can be. Therefore, a first negative reward value can be obtained based on the number of mismatches. For example, if one of the predicted navigation elements does not match a sample navigation element in the matching element set, the target reward value can be decreased by 1 point.
[0103] The more the broadcast order of the predicted navigation elements in the predicted voice broadcast content is inconsistent with the order of the matching sample navigation elements in the matching element set, the lower the target reward value. Therefore, a second negative reward value can be obtained based on the number of order mismatches. For example, if the broadcast order of one of the predicted navigation elements is inconsistent with the order of the matching sample navigation elements in the matching element set, the target reward value can be decreased by 1 point.
[0104] Finally, the target reward value can be calculated through the above positive reward value, first negative reward value, and second negative reward value.
[0105] In an optional implementation of this embodiment, the step of obtaining sample data further includes the following steps:
[0106] Obtain a sample navigation trajectory;
[0107] Based on a preset online policy, determine a sample position and a combination of sample navigation elements at the sample position for the sample navigation trajectory;
[0108] Use a value function to calculate the value of the sample navigation elements and the value of the reference navigation elements; wherein, the reference navigation elements are the navigation elements corresponding to not performing a voice broadcast.
[0109] Sort the sample navigation elements whose values are greater than the value of the reference navigation element in descending order of value to form the voice broadcast sample content at the sample position.
[0110] In this optional implementation, considering that there are various difficulties in collecting historical navigation data, or the collected historical navigation data is limited, it is easy to result in an insufficient number of sample data. Therefore, the embodiments of the present disclosure also propose a method for obtaining sample data.
[0111] In the embodiments of the present disclosure, sample navigation trajectories can be generated for various address pairs (including start addresses and destinations) mined on the electronic map. The sample navigation trajectory includes a plurality of temporally continuous trajectory points on the navigation route from the start address to the destination address.
[0112] In the embodiments of the present disclosure, an online strategy can also be preset to determine the sample position and the voice broadcast sample content at the sample position based on the sample navigation trajectory. This online strategy can adopt the measurements used in traditional voice navigation systems. This online strategy can be the strategy of pre-configuring voice broadcast scenarios and voice broadcast scripts as mentioned in the background art.
[0113] For the sample navigation trajectory, navigation-related information can be assumed first during the navigation process of the sample navigation trajectory, and the online strategy is used to determine the sample position that needs voice navigation and the combination of sample navigation elements at the sample position based on the navigation-related information. It should be noted that the combination of sample navigation elements generated by this online strategy can include all possible sample navigation elements that can be selected for voice broadcast at the sample position. This combination of sample navigation elements also includes the element of not performing voice broadcast. In this embodiment, for the convenience of description, the element of not performing voice broadcast is referred to as the reference navigation element.
[0114] In this embodiment, a value function can also be predefined, and this value function is used to calculate the value of the sample navigation elements and the value of the reference navigation element in the voice broadcast sample content. This value function can calculate the value based on the importance of the sample navigation elements generated by the online strategy, and this value function can be predefined.
[0115] After calculating the value of the sample navigation elements, the sample navigation elements whose values are greater than the reference navigation element can be sorted in descending order of value to form the voice broadcast sample content.
[0116] The sample data collected in this way involves a limited variety of voice broadcast scenarios. However, after reinforcement learning, the voice broadcast model can adapt to voice broadcast scenarios not involved in the sample data, and finally can achieve the technical effect of outputting more accurate voice broadcast content in various voice broadcast scenarios.
[0117] In an alternative implementation of this embodiment, the action information further includes a prediction timing for voice broadcast; the method further includes the following steps:
[0118] Update the target reward value based on a matching result of whether a predicted sample position corresponding to the prediction timing for voice broadcast matches the current sample position.
[0119] In this alternative implementation, in addition to outputting the action information of the voice broadcast prediction content, the voice broadcast model can also output a prediction timing for the voice broadcast prediction content, that is, the action information can include the voice broadcast prediction content and the prediction timing for voice broadcast. In some embodiments, the prediction timing for voice broadcast can be given in the form of a trajectory point in the sample navigation trajectory. When calculating the target reward value, it can be calculated based on a matching result between the voice broadcast prediction content and the voice broadcast sample content, and whether a predicted sample position corresponding to the prediction timing for voice broadcast is consistent with the current sample position corresponding to the voice broadcast sample content. That is to say, there are two factors affecting the target reward value, namely, the matching degree of navigation elements between the voice broadcast prediction content and the voice broadcast sample content, and the matching degree of the voice broadcast timing between the voice broadcast prediction content and the voice broadcast sample content.
[0120] Therefore, when calculating the target reward value, a first target reward value can be calculated based on a matching result of navigation elements between the voice broadcast prediction content and the voice broadcast sample content, and then a second target reward value can be calculated based on a matching result of the voice broadcast timing between the voice broadcast prediction content and the voice broadcast sample content. The final target reward value can be obtained by adding the first target reward value and the second target reward value.
[0121] FIG. 2(a) and FIG. 2(b) show the formation process of a set of matching elements and a schematic diagram of the calculation of the target reward value according to an embodiment of the present disclosure. As shown in FIG. 2(a), the action space shown on the left is the sample navigation element combination output by the online policy in any state. The value of all sample navigation elements can be calculated using the value function, and all navigation elements with a value greater than "not announce" are extracted and added to the dynamic action space. The navigation elements are sorted based on the value magnitudes of different navigation elements in the dynamic action space. As shown in FIG. 2(a), the sample navigation element combination includes six navigation elements, namely "lane", "camera", "not announce", "main action", "slope", and "non-navigation". After these six navigation elements are calculated by the value function F(x), the corresponding values can be obtained, and the value magnitudes are: 500, 200, 300, 400, 100, and 10 respectively. The value of the element "not announce" is 300. Therefore, "lane" and "main action" with values greater than 300 can be added to the set of matching elements.
[0122] As shown in FIG. 2(b), assume that the model outputs four predicted voice announcement contents, and the predicted navigation element combinations in these four predicted voice announcement contents are: "lane->main action", "main action->lane", "lane", "lane->camera", where "->" indicates that the navigation element on the left is sorted before the navigation element on the right. Assume further that when the predicted navigation element revealed by the voice announcement model matches the sample navigation element in the dynamic action space, a reward value of 1 point is given, while if the voice announcement model reveals less, reveals additionally, or the revealed order is inconsistent with the order of the sample navigation elements in the dynamic action space, a penalty of -0.5 points is given, that is, a negative reward value is given; and if the predicted announcement timing does not match the current sample position, a penalty of -0.5 points is also given. Based on the above assumptions, according to the embodiment of the present disclosure, the target reward values shown in FIG. 2(b) can be obtained for the four predicted voice announcement contents output by the voice announcement model, which are: 2, 1.5, 0.5, and 0 points respectively.
[0123] In an alternative implementation manner of this embodiment, the action information further includes the predicted announcement timing; the method further includes the following steps:
[0124] Determine the state transition information based on the announcement duration of the predicted voice announcement content and the voice announcement timing; the state transition information includes the next sample position corresponding to the next state and the navigation-related information corresponding to the next sample position;
[0125] Use the state transition information to perform the next round of reinforcement training on the voice announcement model.
[0126] In this optional implementation, reinforcement learning is used to describe and solve the problem of an agent maximizing the reward value or achieving a specific goal through learning strategies during the interaction with the environment. That is, it is applied to the application scenario of the agent-environment interaction. To obtain the maximum reward value, the actions taken by the agent at each state are crucial. That is, based on the actions taken by the agent, the reward value is determined by the next state transferred by executing the action. If the action selection is accurate, a large reward value is obtained, while if the action selection is wrong, a small or even negative reward value is obtained.
[0127] As described above, in the embodiments of the present disclosure, the state is constructed as the navigation-related information when the object to be navigated reaches a certain position, and the action is constructed as the voice broadcast content to be output at this position. The voice broadcast content includes one or more navigation elements, that is, the voice broadcast content can be understood as a combination of navigation elements or multiple navigation elements. The next state in the embodiments of the present disclosure includes the navigation-related information at the next position where the voice broadcast content needs to be output. To determine the next state, it is necessary to first determine the next position where the voice broadcast content is output.
[0128] In the embodiments of the present disclosure, the next state is determined based on the predicted broadcast timing of the voice broadcast model and the broadcast duration of the predicted voice broadcast content. That is, the position obtained by delaying the predicted broadcast timing of the predicted voice broadcast content in the current state by the broadcast duration of the predicted voice broadcast content is used as the next sample position where the voice broadcast content is output, and the navigation-related information at the next sample position is determined as the next state after the state transition.
[0129] During the reinforcement learning process of the voice broadcast model, after a round of reinforcement training on the voice broadcast model using the navigation-related information at the current sample position and the voice broadcast sample content, the next sample position corresponding to the next state can be calculated based on the above principle. Then, based on the navigation-related information at the next sample position and the voice broadcast sample content, the next round of reinforcement training is performed on the voice broadcast model.
[0130] In this way, the voice broadcast model can learn the correct state transition process, that is, the voice broadcast model can learn when to output the voice broadcast content so as to output the correct voice broadcast timing based on the current state.
[0131] Figure 3 The flowchart of the navigation voice broadcast method according to an embodiment of the present disclosure is shown. As Figure 3 shown, the navigation voice broadcast method includes the following steps:
[0132] In step S301, obtain navigation-related information at the current navigation position;
[0133] In step S302, input the navigation-related information into a pre-trained voice broadcast model to obtain navigation elements for voice broadcast and the voice broadcast timing;
[0134] In step S303, assemble the navigation elements to generate a voice broadcast script;
[0135] In step S304, output the voice broadcast script at the voice broadcast timing.
[0136] In this embodiment, the navigation voice broadcast method can be executed on a navigation server or a navigation terminal. When executed on a navigation server, a pre-trained voice navigation model can be deployed on the navigation server. The voice navigation model can output navigation elements for voice broadcast and the voice broadcast timing based on navigation-related information at any position. After receiving a navigation request from a navigation terminal, the navigation server plans a navigation path based on the starting address and the destination address, and sends the navigation path to the navigation terminal. After the navigation starts, the navigation terminal feeds back the current position of the object being navigated to the navigation server in real time. The navigation server can use any received current position as the position in the current state, and collect navigation-related information, such as information that can be recalled in the subsequent road segment sequence (e.g., the type of recalled navigation element, the distance between the current position and the recalled navigation element), the speed information of the object being navigated at the current position, the acceleration and deceleration information, the distance information between the voice broadcast point and a certain navigation element, the time interval information since the last voice broadcast, etc.
[0137] The navigation server inputs the navigation-related information of the current position into the voice broadcast model. The voice broadcast model can output a navigation element combination and the voice broadcast timing based on the navigation-related information. The navigation server can assemble the navigation element combination into a voice broadcast script based on a pre-set script template, and send the voice broadcast script and the voice broadcast timing to the navigation terminal, so that the navigation terminal can output the voice broadcast script at the voice broadcast timing.
[0138] For the relevant details of the voice broadcast model, reference can be made to the reinforcement learning method of the voice broadcast model described above. In this reinforcement learning method, the voice broadcast prediction content output by the voice broadcast model based on sample data is split into a form of predicting a navigation element combination, and a target reward value is constructed based on the matching relationship between the predicted navigation element and the sample navigation element in the voice broadcast sample content. Then, the voice broadcast model is strengthened and trained using the target reward value.
[0139] It should be noted that the training method of the voice broadcast model in this embodiment is not limited to the reinforcement learning method described above, and other training methods can also be used, as long as the voice broadcast model can output the combination of navigation elements and the voice broadcast timing based on the navigation-related information at the current position.
[0140] It should also be noted that if voice broadcast is not required, the combination of navigation elements output by the voice navigation model can only include the element of "not broadcast", and the voice broadcast timing can be empty.
[0141] In the embodiments of the present disclosure, by inputting the navigation-related information at the current position into the voice broadcast model, the voice broadcast model outputs the combination of navigation elements and the voice broadcast timing based on the navigation-related information, and then generates the voice broadcast script based on the combination of navigation elements output by the voice broadcast model, so that the navigation terminal outputs the voice broadcast script at the voice broadcast timing. Through the embodiments of the present disclosure, the trained voice broadcast model can be used to adaptively output the navigation elements to be broadcast based on the information in the current scenario, so that the voice navigation can adapt to various voice broadcast scenarios, and the accuracy of the finally output voice navigation script is higher, which can better meet the refined requirements of users for voice navigation.
[0142] Figure 4 Show a schematic diagram of a voice navigation scenario according to an embodiment of the present disclosure. As Figure 4 shown, the model training server can collect a plurality of sample data, and each sample data can include the navigation-related information at the sample position in any sample navigation trajectory and the voice broadcast sample content at the sample position. Based on the plurality of sample data, reinforcement learning is performed on the voice broadcast model to obtain a trained voice broadcast model. The voice broadcast model can be sent to the navigation server for deployment on the navigation server. After receiving the navigation request of the navigation terminal, the navigation server pushes the planned route and related information to the navigation terminal. During the navigation process, the navigation terminal can send the current position of the object being navigated to the navigation server in real time. The navigation server obtains the navigation-related information at the current position and inputs it into the voice broadcast model. The voice broadcast model correspondingly outputs the combination of navigation elements for voice broadcast and the voice broadcast timing. The server assembles the navigation elements into a voice broadcast script and sends the voice broadcast script and the voice broadcast timing to the navigation terminal. During the navigation process, when the navigation terminal detects that the position of the object being navigated is at the position corresponding to the voice broadcast timing, it outputs the voice broadcast script to guide the object being navigated to perform the correct navigation action.
[0143] A method for providing location-based services according to an embodiment of the present disclosure. The method for providing location-based services trains a voice broadcast model using the reinforcement learning method of the above voice broadcast model, and uses the trained voice broadcast model to provide location-based services for the service recipient. The location-based services include one or more of navigation, map rendering, and route planning.
[0144] In this embodiment, the method for providing location-based services can be executed on a terminal, and the terminal is a mobile phone, iPad, computer, smart watch, vehicle, robot, etc. In the embodiments of the present disclosure, voice broadcast terms and voice broadcast timing can be obtained for the service recipient, and then, during the location-based service process, the voice broadcast terms and voice broadcast timing can be used to provide a more accurate voice navigation service for the service recipient. In addition, the content in the voice broadcast terms can also be graphically displayed when rendering the map, etc.
[0145] The following are the device embodiments of the present disclosure, which can be used to execute the method embodiments of the present disclosure.
[0146] Figure 5 The structural block diagram of a reinforcement learning device for a voice broadcast model according to an embodiment of the present disclosure is shown. The device can be implemented as part or all of an electronic device through software, hardware, or a combination of both. As Figure 5 shown, the reinforcement learning device for the voice broadcast model includes:
[0147] An acquisition module 501, configured to acquire sample data; the sample data includes navigation-related information of a sample position in a sample navigation trajectory and voice broadcast sample content output at the sample position;
[0148] An input module 502, configured to input the navigation-related information of the current sample position in the sample navigation trajectory as the current state into the voice broadcast model to obtain action information in the current state; the action information includes voice broadcast prediction content at the current sample position;
[0149] A calculation module 503, configured to calculate a target reward value based on a matching result between the voice broadcast prediction content and the voice broadcast sample content output at the current sample position; the matching result includes a matching result between a predicted navigation element in the voice broadcast prediction content and a sample navigation element in the voice broadcast sample content;
[0150] A first training module 504, configured to perform reinforcement training on the voice broadcast model based on the target reward value.
[0151] In an optional implementation manner of this embodiment, the calculation module includes:
[0152] The first determination sub-module is configured to determine the predicted navigation elements included in the predicted voice broadcast content and the broadcast order of the predicted navigation elements in the predicted voice broadcast content;
[0153] The first matching sub-module is configured to perform element matching between the predicted navigation elements and the sample navigation elements in the voice broadcast sample content;
[0154] The second matching sub-module is configured to match the broadcast order of the predicted navigation elements in the predicted voice broadcast content with the broadcast order of the sample navigation elements in the voice broadcast sample content;
[0155] The first calculation sub-module is configured to calculate the target reward value based on the element matching result and the order matching result.
[0156] In this optional implementation manner, the voice broadcast sample content can be a complete voice broadcast script or a combination of sample navigation elements. When the voice broadcast sample content is a complete voice broadcast script, it can be pre-split into a combination form of sample navigation elements and can be arranged in the broadcast order of the sample navigation elements in the voice broadcast sample content to form a sample navigation element set.
[0157] In some embodiments, the predicted voice broadcast content output by the voice broadcast model can be predicted navigation elements arranged in the broadcast order or a complete broadcast script assembled from the predicted navigation elements according to a script template. Therefore, it is possible to first determine which predicted navigation elements are included in the predicted voice broadcast content output by the voice broadcast model and the broadcast order of these predicted navigation elements.
[0158] After that, the predicted navigation elements are matched with the sample navigation elements, and the matching result can be determined based on the number of corresponding consistent navigation elements. It should be noted that the matching between navigation elements can be independent of the broadcast order.
[0159] In addition, the broadcast orders of the matched predicted navigation elements and sample navigation elements can be matched in order, and the matching result can be determined based on whether the broadcast orders are consistent.
[0160] Based on the two matching results, the target reward value can be comprehensively obtained. For example, if the number of element matches is large and the broadcast orders are consistent, the target reward value is higher; if the number of element matches is small and the broadcast orders are inconsistent, the target reward value is lower.
[0161] That is to say, the matching process between the predicted content of voice broadcast and the sample content of voice broadcast is divided into two parts: element matching and order matching. Element matching refers to matching the predicted navigation elements in the predicted content of voice broadcast with the sample navigation elements in the set of matching elements. The more the number of matching elements, the better the matching result. Order matching refers to matching the broadcast order of the predicted navigation elements in the predicted content of voice broadcast with the order of the sample navigation elements that match in the set of matching elements. The more consistent the order matching, the better the matching result.
[0162] Finally, the target reward value can be calculated based on the results of element matching and order matching.
[0163] In an alternative implementation of this embodiment, the calculation module includes:
[0164] The second determination sub-module is configured to determine a positive reward value based on the number of matching predicted navigation elements and sample navigation elements;
[0165] The third determination sub-module is configured to determine a first negative reward value based on the number of non-matching predicted navigation elements and sample navigation elements;
[0166] The fourth determination sub-module is configured to determine a second negative reward value based on the number of non-matching broadcast orders of the predicted navigation elements in the predicted content of voice broadcast and the broadcast orders of the sample navigation elements in the sample content of voice broadcast;
[0167] The fifth determination sub-module is configured to determine the target reward value based on the positive reward value, the first negative reward value, and the second negative reward value.
[0168] In this alternative implementation, the more the number of matching predicted navigation elements and sample navigation elements in the set of matching elements, the higher the target reward value. Therefore, the positive reward value can be obtained based on the number of matching elements. For example, if one of the predicted navigation elements matches a sample navigation element in the set of matching elements, the target reward value can be increased by 1 point.
[0169] And the more the number of non-matching predicted navigation elements and sample navigation elements in the set of matching elements, the lower the target reward value. Therefore, the first negative reward value can be obtained based on the number of non-matching elements. For example, if one of the predicted navigation elements does not match a sample navigation element in the set of matching elements, the target reward value can be decreased by 1 point.
[0170] The more the case where the broadcast order of the predicted navigation elements in the predicted content of the voice broadcast is inconsistent with the order of the sample navigation elements matching in the set of matching elements, the lower the target reward value. Therefore, the second negative reward value can be obtained based on the number of non-matching orders. For example, if the broadcast order of one of the predicted navigation elements is inconsistent with the order of the sample navigation elements matching in the set of matching elements, the target reward value can be reduced by 1 point.
[0171] Finally, the target reward value can be calculated through the above positive reward value, the first negative reward value, and the second negative reward value.
[0172] In an alternative implementation manner of this embodiment, the obtaining module includes:
[0173] The first obtaining sub-module is configured to obtain a sample navigation trajectory;
[0174] The sixth determining sub-module is configured to determine a sample position and a combination of sample navigation elements at the sample position based on a preset online policy for the sample navigation trajectory;
[0175] The second calculating sub-module is configured to calculate the value of the sample navigation elements and the value of the reference navigation elements by using a value function; wherein, the reference navigation element is the navigation element corresponding to not performing a voice broadcast;
[0176] The second obtaining sub-module is configured to sort the sample navigation elements with a value greater than the value of the reference navigation element in the order of value magnitude to form the voice broadcast sample content at the sample position.
[0177] In this alternative implementation manner, considering that there are various difficulties in collecting historical navigation data, or the collected historical navigation data is limited, it is easy to cause the insufficient quantity of sample data. Therefore, the present disclosure embodiment also proposes a method for obtaining sample data.
[0178] In the present disclosure embodiment, various address pairs (including the starting address and the destination) mined on the electronic map can be used to generate a sample navigation trajectory, and the sample navigation trajectory includes a plurality of temporally continuous trajectory points on the navigation route from the starting address to the destination address.
[0179] In the present disclosure embodiment, an online policy can also be preset in advance for determining a sample position and the voice broadcast sample content at the sample position based on the sample navigation trajectory. This online policy can adopt the measurements used in traditional voice navigation systems, and this online policy can be a policy of pre-configuring voice broadcast scenarios and voice broadcast words as mentioned in the background art.
[0180] For the sample navigation trajectory, it is possible to first assume the navigation-related information during the navigation process of the sample navigation trajectory, and use an online strategy to determine the sample positions that require voice navigation and the combination of sample navigation elements at the sample positions based on the navigation-related information. It should be noted that the combination of sample navigation elements generated by the online strategy may include all possible sample navigation elements that can be selected for voice broadcast at the sample position, and the combination of sample navigation elements also includes the element of not performing voice broadcast. In this embodiment, for the convenience of description, the element of not performing voice broadcast is referred to as the reference navigation element.
[0181] In this embodiment, a value function can also be predefined, and the value function is used to calculate the value of the sample navigation element and the value of the reference navigation element in the voice broadcast sample content. The value function can calculate the value based on the importance degree of the sample navigation elements generated by the online strategy, and the value function can be predefined.
[0182] After calculating the value of the sample navigation element, the sample navigation elements with a value greater than the reference navigation element can be sorted according to the value size to form the voice broadcast sample content.
[0183] The sample data collected in this way involves a limited variety of voice broadcast scenarios. However, after reinforcement learning, the voice broadcast model can adapt to voice broadcast scenarios not covered in the sample data, and finally can achieve the technical effect of outputting more accurate voice broadcast content in various voice broadcast scenarios.
[0184] In an alternative implementation of this embodiment, the action information further includes a prediction timing for the voice broadcast; the apparatus further includes:
[0185] An update module, configured to update the target reward value based on a matching result of whether the predicted sample position corresponding to the prediction timing for the voice broadcast matches the current sample position.
[0186] In this alternative implementation, in addition to outputting the action information of the voice broadcast prediction content, the voice broadcast model can also output the prediction timing for the voice broadcast prediction content, that is, the action information can include the voice broadcast prediction content and the prediction timing for the voice broadcast. In some embodiments, the prediction timing for the voice broadcast can be given in the form of a trajectory point in the sample navigation trajectory. When calculating the target reward value, it can be calculated based on the matching result between the voice broadcast prediction content and the voice broadcast sample content, and whether the predicted sample position corresponding to the prediction timing for the voice broadcast is consistent with the current sample position corresponding to the voice broadcast sample content. That is to say, there are two factors affecting the target reward value, namely, the matching degree of the navigation elements between the voice broadcast prediction content and the voice broadcast sample content, and the matching degree of the broadcast timing between the voice broadcast prediction content and the voice broadcast sample content.
[0187] Therefore, when calculating the target reward value, the first target reward value can be calculated based on the matching result of the navigation elements in the predicted voice broadcast content and the voice broadcast sample content. Then, the second target reward value can be calculated based on the matching result of the broadcast timing of the predicted voice broadcast content and the voice broadcast sample content. After superimposing the first target reward value and the second target reward value, the final target reward value can be obtained.
[0188] In an alternative implementation of this embodiment, the action information further includes the voice broadcast timing; the apparatus further includes:
[0189] A determination module, configured to determine state transition information based on the broadcast duration of the predicted voice broadcast content and the voice broadcast timing; the state transition information includes the next sample position corresponding to the next state and the navigation-related information corresponding to the next sample position;
[0190] A second training module, configured to perform the next round of reinforcement training on the voice broadcast model by using the state transition information.
[0191] In this alternative implementation, reinforcement learning is used to describe and solve the problem that an agent maximizes the reward value or achieves a specific goal through learning strategies during the interaction with the environment. That is, it is applied to the application scenario of the interaction between the agent and the environment. In order to obtain the maximum reward value, the actions taken by the agent in each state are crucial. That is, based on the actions taken by the agent, the size of the reward value is determined by the next state transferred by executing the action. If the action selection is accurate, a large reward value is obtained, while if the action selection is incorrect, a small reward value is obtained, or even negative.
[0192] As described above, in the embodiments of the present disclosure, the state is constructed as the navigation-related information when the object to be navigated reaches a certain position, and the action is constructed as the voice broadcast content to be output at this position. The voice broadcast content includes one or more navigation elements. That is, the voice broadcast content can be understood as a navigation element or a combination of multiple navigation elements. The next state in the embodiments of the present disclosure includes the navigation-related information at the next position where the voice broadcast content needs to be output. In order to determine the next state, it is necessary to first determine the next position where the voice broadcast content is output.
[0193] In the embodiments of the present disclosure, the next state is determined based on the predicted broadcast timing of the voice broadcast model output and the broadcast duration of the predicted voice broadcast content. That is, the position obtained by delaying the predicted broadcast timing of the predicted voice broadcast content in the current state by the broadcast duration of the predicted voice broadcast content is used as the next sample position for outputting the voice broadcast content, and the navigation-related information at the next sample position is determined as the next state after the state transition.
[0194] In the reinforcement learning process of the voice broadcast model, after performing a round of reinforcement training on the voice broadcast model by using the navigation-related information at the current sample position and the content of the voice broadcast sample, the next sample position corresponding to the next state can be calculated based on the above principle, and then the next round of reinforcement training can be performed on the voice broadcast model based on the navigation-related information and the content of the voice broadcast sample at the next sample position.
[0195] In this way, the voice broadcast model can learn the correct state transition process, that is, the voice broadcast model can learn when to output the voice broadcast content so as to output the correct voice broadcast timing based on the current state.
[0196] Figure 6 It is a schematic structural diagram of an electronic device suitable for implementing the navigation voice broadcast method according to the embodiments of the present disclosure.
[0197] As Figure 6 shown, the electronic device 600 includes a processing unit 601, which can be implemented as a processing unit such as a CPU, GPU, FPGA, NPU, etc. The processing unit 601 can execute various processes in the embodiments of any of the above methods of the present disclosure according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage section 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.
[0198] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that a computer program read from it can be installed into the storage section 608 as needed.
[0199] In particular, according to an embodiment of the present disclosure, any of the methods described above with reference to the embodiments of the present disclosure may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program tangibly embodied on a machine-readable medium, the computer program including program code for performing any of the methods in the embodiments of the present disclosure. In such an embodiment, the computer program may be downloaded and installed from a network via the communication section 609, and / or installed from the removable medium 611.
[0200] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that includes one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in an order different from that noted in the drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0201] The units or modules involved in the embodiments described in the present disclosure may be implemented in software or in hardware. The units or modules described may also be provided in a processor, and the names of these units or modules do not, in some cases, constitute a limitation on the units or modules themselves.
[0202] As another aspect, the present disclosure also provides a computer-readable storage medium, which may be the computer-readable storage medium included in the device described in the above embodiments; or it may exist separately and be a computer-readable storage medium not assembled into the device. The computer-readable storage medium stores one or more programs, and the one or more programs are used by one or more processors to execute the methods described in the present disclosure.
[0203] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present disclosure is not limited to the technical solution formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. For example, the technical solution formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the present disclosure that have similar functions.
Claims
1. A reinforcement learning method for a voice broadcast model, wherein, Including: Obtain sample data; the sample data includes navigation-related information at a sample position in a sample navigation trajectory and the voice broadcast sample content output at the sample position; Use the navigation-related information at the current sample position in the sample navigation trajectory as the current state and input it into a voice broadcast model to obtain action information in the current state; the action information includes the predicted voice broadcast content at the current sample position; Calculate a target reward value based on the matching result between the predicted voice broadcast content and the voice broadcast sample content output at the current sample position; the matching result includes the matching result between the predicted navigation elements in the predicted voice broadcast content and the sample navigation elements in the voice broadcast sample content; Perform reinforcement training on the voice broadcast model based on the target reward value.
2. The method according to claim 1, wherein, Calculating a target reward value based on the matching result between the predicted voice broadcast content and the voice broadcast sample content output at the current sample position includes: Determine the predicted navigation elements included in the predicted voice broadcast content and the broadcast order of the predicted navigation elements in the predicted voice broadcast content; Perform element matching between the predicted navigation elements and the sample navigation elements in the voice broadcast sample content; Match the broadcast order of the predicted navigation elements in the predicted voice broadcast content with the broadcast order of the sample navigation elements in the voice broadcast sample content; Calculate the target reward value based on the element matching result and the order matching result.
3. The method according to claim 2, wherein, Calculating the target reward value based on the element matching result and the order matching result includes: Determine a positive reward value based on the number of matches between the predicted navigation elements and the sample navigation elements; Determine a first negative reward value based on the number of mismatches between the predicted navigation elements and the sample navigation elements; Determine a second negative reward value based on the number of mismatches between the broadcast order of the predicted navigation elements in the predicted voice broadcast content and the broadcast order of the sample navigation elements in the voice broadcast sample content; Determine the target reward value based on the positive reward value, the first negative reward value, and the second negative reward value.
4. The method according to any one of claims 1-3, wherein The obtaining of the sample data includes: Obtain a sample navigation trajectory; Based on a preset online strategy, determine a sample position and a combination of sample navigation elements at the sample position for the sample navigation trajectory; Calculate the value of the sample navigation elements and the value of the reference navigation elements using a value function; wherein, the reference navigation elements are the navigation elements corresponding to not performing a voice broadcast; Sort the sample navigation elements with a value greater than the value of the reference navigation elements in order of value size to form the voice broadcast sample content at the sample position.
5. The method according to any one of claims 1-3, wherein, The action information further includes a predicted broadcast timing; the method further includes: Update the target reward value based on the matching result of whether the predicted sample position corresponding to the predicted broadcast timing matches the current sample position.
6. The method according to any one of claims 1-3, wherein, The action information further includes a voice broadcast timing; the method further includes: Determine state transition information based on the broadcast duration of the predicted content of the voice broadcast and the voice broadcast timing; the state transition information includes the next sample position corresponding to the next state and the navigation-related information corresponding to the next sample position; Use the state transition information to perform the next round of reinforcement training on the voice broadcast model.
7. A navigation voice broadcast method, wherein, including: Obtain navigation-related information at the current navigation position; Input the navigation-related information into a voice broadcast model pre-trained by the method according to any one of claims 1-6 to obtain the navigation elements for voice broadcast and the voice broadcast timing; Assemble the navigation elements to generate a voice broadcast script; Output the voice broadcast script at the voice broadcast timing.
8. A method for providing location-based services, wherein, The location-based service providing method trains a voice broadcast model by the method according to any one of claims 1-6, and uses the trained voice broadcast model to provide location-based services for the service recipient, and the location-based services include: one or more of navigation, map rendering, and route planning.
9. An electronic device, wherein, including a memory, a processor, and a computer program stored on the memory, wherein the processor executes the computer program to implement the method according to any one of claims 1-7.
10. A computer program product, comprising computer instructions, wherein, When the computer instruction is executed by the processor, it implements the method according to any one of claims 1-7.
Citation Information
Patent Citations
Voice broadcasting method and device, and electronic equipment
CN112735167A