A training sample generation method and device, computer equipment and a storage medium
By replacing style-expressing phrases and words with different tags in navigation broadcast statements and combining style splicing training samples, the problem of limited styles and low rewriting efficiency in existing technologies for navigation broadcast statements is solved, and the efficient generation of diverse style-expressing navigation broadcast statements is achieved.
Patent Information
- Application Number
- CN202211248296.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-12
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2042-10-12
AI Technical Summary
Existing technologies have limited styles for generating navigation announcements. Manual rewriting is time-consuming and labor-intensive, and methods based on artificial neural networks have difficulty effectively learning style and navigation-related knowledge, resulting in poor anthropomorphic rewriting effects.
By acquiring styled navigation broadcast statements, identifying style expression phrases and style expression words, and replacing them with different tags to generate first and second tags, training samples are generated. These samples are then combined with style to form feature values and labels for training the navigation broadcast statement generation network.
The navigation broadcast statement generation network has been made capable of learning style-related expressions, distinguishing different styles, and generating diverse navigation broadcast statements with the expected style, thereby improving the efficiency and effectiveness of anthropomorphic rewriting of navigation broadcast statements.
Smart Images

Figure CN115526262B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] One or more embodiments of the present specification relate to the technical field of navigation, and in particular to a training sample generation method and device, a computer device, and a storage medium. BACKGROUND
[0002] In the navigation function of a map application in the related art, the style of a navigation broadcast sentence can generally be selected, such as a soft style navigation broadcast sentence, a cute style navigation broadcast sentence, and the like.
[0003] In the related art, different styles of navigation broadcast sentences are generally obtained by manually rewriting navigation broadcast sentences without style. The manual rewriting manner causes the generation process of different styles of navigation broadcast sentences to consume more manpower. SUMMARY
[0004] Therefore, one or more embodiments of the present specification provide a training sample generation method and device, a computer device, and a storage medium.
[0005] According to a first aspect of one or more embodiments of the present specification, a training sample generation method is provided, including:
[0006] obtaining a navigation broadcast sentence with style, and determining a style expression phrase, a style expression word, and a style corresponding to the navigation broadcast sentence in the navigation broadcast sentence; the style expression phrase is a phrase that does not contain navigation broadcast information; the style expression word is a word related to style expression content;
[0007] replacing the style expression phrase with a first mark and replacing the style expression word with a second mark to obtain a processed navigation broadcast sentence; the first mark is used to represent a phrase to be restored by a to-be-trained network, and the second mark is used to represent a word to be restored by the to-be-trained network;
[0008] generating a training sample of the to-be-trained network, the training sample including a feature value and a label, the feature value of the training sample including the processed navigation broadcast sentence and the style corresponding to the navigation broadcast sentence, and the label of the training sample including the navigation broadcast sentence with style.
[0009] According to a second aspect of one or more embodiments of the present specification, a training method of a navigation broadcast sentence generation network is provided, including:
[0010] obtaining each training sample in a training sample set by using the training sample generation method described above;
[0011] The input information in the training sample is input as input of a pre-training language network, the language network is updated according to the input information and a label in the training sample, and a navigation broadcast sentence generation network is obtained; the navigation broadcast sentence generation network is used for generating a navigation broadcast sentence with a style based on a navigation broadcast sentence without the style and an expected style.
[0012] According to a third aspect of the embodiments of the present specification, a method for generating a navigation broadcast sentence is provided, comprising:
[0013] obtaining network input information, the network input information comprising: a first navigation broadcast sentence and a style of a navigation broadcast sentence expected to be generated, the first navigation broadcast sentence being a navigation broadcast sentence without style; the navigation broadcast sentence generation network being obtained by training according to the method for training a navigation broadcast sentence generation network described above;
[0014] inputting the network input information into the navigation broadcast sentence generation network to obtain a second navigation broadcast sentence output by the navigation broadcast sentence generation network, the second navigation broadcast sentence being a navigation broadcast sentence with style corresponding to the first navigation broadcast sentence.
[0015] According to a fourth aspect of the embodiments of the present specification, a training sample generation device is provided, comprising:
[0016] a sentence obtaining module configured to obtain a navigation broadcast sentence with style, and determine a style expression phrase, a style expression word and a style corresponding to the navigation broadcast sentence in the navigation broadcast sentence; the style expression phrase being a phrase without navigation broadcast information; the style expression word being a word related to style expression content;
[0017] a label replacing module configured to replace the style expression phrase with a first label, and replace the style expression word with a second label to obtain a processed navigation broadcast sentence; the first label being used to represent a phrase to be restored by a network to be trained, and the second label being used to represent a word to be restored by the network to be trained;
[0018] a sample generation module configured to generate a training sample of the network to be trained, the training sample comprising a feature value and a label, the feature value of the training sample comprising the processed navigation broadcast sentence and the style corresponding to the navigation broadcast sentence, and the label of the training sample comprising the navigation broadcast sentence with style.
[0019] According to a fifth aspect of the embodiments of the present specification, a training device of a navigation broadcast sentence generation network is provided, comprising:
[0020] a sample obtaining module configured to obtain each training sample in a training sample set by using the method for generating a training sample described above;
[0021] a network training module configured to take the input information in the training sample as input of a pre-trained language network, update the language network according to the input information and a label in the training sample, and obtain a navigation broadcast sentence generation network; the navigation broadcast sentence generation network is configured to generate a navigation broadcast sentence with a style based on a navigation broadcast sentence without the style and an expected style.
[0022] According to a sixth aspect of the embodiments of the present specification, a navigation broadcast sentence generation apparatus is provided, including:
[0023] an input information acquisition module configured to acquire network input information, the network input information including a first navigation broadcast sentence and a style of a navigation broadcast sentence expected to be generated, the first navigation broadcast sentence being a navigation broadcast sentence without a style;
[0024] a sentence generation module configured to input the network input information into the navigation broadcast sentence generation network, and obtain a second navigation broadcast sentence output by the navigation broadcast sentence generation network, the second navigation broadcast sentence being a navigation broadcast sentence with the style corresponding to the first navigation broadcast sentence.
[0025] According to a seventh aspect of the embodiments of the present specification, a computer readable storage medium is provided, and the computer readable storage medium stores computer instructions, and the computer instructions are executed by a processor to implement the training sample generation method, the navigation broadcast sentence generation method, or the training method of the navigation broadcast sentence generation network.
[0026] According to an eighth aspect of the embodiments of the present specification, a computer device is provided, and the computer device includes:
[0027] a processor;
[0028] a memory configured to store processor executable instructions;
[0029] the processor is configured to run the executable instructions to implement the training sample generation method, the navigation broadcast sentence generation method, or the training method of the navigation broadcast sentence generation network.
[0030] The training sample generation method, device, computer equipment and storage medium provided in the specification obtain a navigation broadcast sentence with a style, and determine a style expression short sentence, a style expression word and a style corresponding to the navigation broadcast sentence in the navigation broadcast sentence. The style expression short sentence is a short sentence that does not contain navigation broadcast information. The style expression word is a word related to style expression content. The style expression short sentence is replaced by a first mark, and the style expression word is replaced by a second mark to obtain a processed navigation broadcast sentence. The first mark is used to represent a short sentence that needs to be restored by a to-be-trained network, and the second mark is used to represent a word that needs to be restored by the to-be-trained network. A training sample of the to-be-trained network is generated, the training sample includes a feature value and a label, the feature value of the training sample includes the processed navigation broadcast sentence and the style corresponding to the navigation broadcast sentence, and the label of the training sample includes the navigation broadcast sentence with the style.
[0031] By replacing the style expression short sentence and the style expression word with different marks, the network can learn the use of the short sentence expression related to the style and the style expression word in the navigation broadcast sentence. And by splicing the processed navigation broadcast sentence and the style, the network can correspond the learned style related expression to the matching style, so that the model has the ability to generate the navigation broadcast sentence with the style based on the input style and sentence.
[0032] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the specification. BRIEF DESCRIPTION OF DRAWINGS
[0033] The accompanying drawings incorporated in the specification and forming a part of the specification illustrate embodiments consistent with the specification and serve to explain the principles of the specification.
[0034] Figure 1 FIG. 1 is a flowchart of a training sample generation method according to an exemplary embodiment of the specification.
[0035] Figure 2 FIG. 2 is a flowchart of a training method of a navigation broadcast sentence generation network according to an exemplary embodiment of the specification.
[0036] Figure 3 FIG. 3 is a flowchart of a navigation broadcast sentence generation method according to an exemplary embodiment of the specification.
[0037] Figure 4 FIG. 4 is a block diagram of a training sample generation device according to an exemplary embodiment of the specification.
[0038] Figure 5is a block diagram of a training device of a navigation announcement sentence generation network according to an example embodiment of the present specification.
[0039] Figure 6 is a block diagram of a navigation announcement sentence generation device according to an example embodiment of the present specification.
[0040] Figure 7 is a hardware structure diagram of an electronic device in which a training sample generation device or a navigation announcement sentence generation device or a training device of a navigation announcement sentence generation network according to an example embodiment of the present specification. DETAILED DESCRIPTION
[0041] The example embodiments will be described in detail herein with reference to the accompanying drawings. In the following description, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following example embodiments are not meant to represent all implementations consistent with one or more embodiments of the present specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of one or more embodiments of the present specification as detailed in the appended claims.
[0042] It should be noted that the steps of the corresponding methods are not necessarily performed in the order shown and described in the present specification in other embodiments. In some other embodiments, the steps included in the methods can be more or less than those described in the present specification. In addition, a single step described in the present specification can be divided into multiple steps for description in other embodiments, and multiple steps described in the present specification can be combined into a single step for description in other embodiments.
[0043] In order to assist the user to walk according to the route recommended by the navigation function of the map application, the navigation function is generally equipped with the corresponding navigation announcement sentence, and the map application generally plays the navigation announcement sentence to remind the user when it needs to pay attention to do something. For example, the route recommended by the navigation function shows that the user needs to turn left at the next intersection, and the map application needs to play the following navigation announcement sentence: "Slow down, turn left at the red and green light intersection ahead"; the navigation function finds that the user is speeding, and the application needs to play the corresponding navigation announcement sentence to remind the user not to speed.
[0044] In order to improve the user experience, the map application generally provides the user with a variety of personification styles (hereinafter referred to as style) of navigation broadcast sentences, such as star A style navigation broadcast sentence, cartoon character B style navigation broadcast sentence and cute style navigation broadcast sentence and so on. Taking the cute style as an example, the original navigation broadcast sentence without style is "there is a sharp turn ahead", and the broadcast sentence after rewriting in a soft and cute style can be "there is a sharp turn ahead, slow down to safely turn".
[0045] The existing navigation broadcast sentence style is limited, and in order to provide better user experience, the map application needs to provide more styles of navigation broadcast sentences to the user.
[0046] There are several ways to generate navigation broadcast sentences in related technologies:
[0047] First, the navigation broadcast sentence is rewritten by artificial means. This method usually requires a professional artificial team to rewrite the navigation broadcast sentence without style according to the required style, but this rewriting requires the cooperation of professionals with professional navigation knowledge, and is time-consuming, requiring high labor and time costs. And generally, one style will only rewrite one navigation broadcast sentence for one navigation broadcast sentence, and the navigation broadcast sentence is difficult to have diversity. Moreover, it is impossible to generate a large number of different styles quickly.
[0048] Second, based on artificial neural networks to rewrite.
[0049] First, a regularization module can be added after the network for generating navigation broadcast sentences with style, and a pre-trained language network is used to change the navigation broadcast sentence without style into a navigation broadcast sentence with the expected style. The regularization module is used to judge whether the content completed by the pre-trained language network meets the requirements of style, navigation, etc., and outputs the sentences that meet the requirements of the regularization module, and discards the sentences that do not meet the requirements.
[0050] Second, in the adjustment stage, parameters can be added in the pre-trained language network (i.e. adding a custom conditional coding before the word embedding layer) to enable the network to learn the content related to the style expression. However, due to the increase in network parameters, overfitting is easy to occur, which makes it impossible to train the desired network.
[0051] As can be seen from the above method, the first method cannot learn the style and navigation related knowledge, cannot effectively play the performance of the pre-trained language network, and because the network cannot learn the style and navigation related knowledge, the regularization module may have difficulty outputting more diverse expressions for the same meaning, resulting in poor training results.
[0052] The network in the second method can theoretically realize the personification rewriting of the navigation broadcast sentence, but the network cannot be trained and the demand for personification rewriting of the navigation broadcast sentence cannot be realized in practice.
[0053] To solve the above problems, in the method shown in the specification, the content contained in each part of the training sample is changed, so that the network can learn the style-related content.
[0054] To better illustrate the method shown in the specification, the pre-trained language network will be described first. The pre-trained language network (PLM, Pre-trained Language Model) is a pre-trained language model trained on a large corpus, and the pre-training method is unsupervised learning. The pre-training sample is to randomly mask (mask, that is, replace the words in the sentence with a mask mark) the words in the sentence, and learn through unsupervised learning to make the network learn the association between sentences.
[0055] The input of the pre-trained language network is to build a template, such as a training sample: [CLS] I like the Disney films very much. [SEP] It was [MASK]. [SEP], wherein [CLS] is followed by the original text, [SEP] is used to splice two context-related sentences, and [MASK] is the [MASK] mark mentioned above. Here, fun in the original sentence is actually replaced with [MASK]. In this way, the network can complete the pre-training of the language network through such text.
[0056] In related technologies, after the pre-training of the pre-trained language network is completed, a small batch of samples are often used to fine-tune the network so that the network can be more suitable for the application scenario. The input of the fine-tuning training sample is the same as that of the pre-training stage, that is, the words in the sentence are randomly masked. The difference is that the fine-tuning process adopts a supervised learning process. The language network after fine-tuning in related technologies is generally used for sentence classification, and the content of the mask restored by the final network is the classification result.
[0057] In the specification, first of all, the mask processing in related technologies is randomly selected to mask the words in the sentence, which actually takes all the words as the learning goal of the network, and the fine-tuning efficiency is low. In the specification, the mask processing is performed on the expressions related to the style (that is, replaced with a special mark), so that the network can learn the expressions related to the style by changing the input of the network.
[0058] Secondly, considering the style-related expressions, in some cases, the whole sentence is a style-related expression, such as "The speed limit of the road ahead is 80, slow down, I will worry about you", which is actually a personification rewriting of "The speed limit of the road ahead is 80", and "slow down, I will worry about you" in the sentence are style-related expressions (this rewriting method is called hanging generation). In some cases, only some adverbs in the sentence are style-related expressions, such as rewriting "The road ahead merges into the main road, be careful of the oncoming car on the left" into "The road ahead merges into the main road, be careful of the oncoming car on the left", and "la" and "oh" in the rewritten sentence are style-related expressions (this rewriting method is called sentence rewriting).
[0059] In the two rewriting methods, if the style-related short sentences and the style-related words are replaced by one mark, the replacement effect may not be good, and the network cannot distinguish between the two expressions and cannot determine whether a sentence or a word needs to be generated. In order to solve the above problem, in the specification, the style-related short sentences and the style-related words are replaced by different marks, so that the network can learn the difference between the two expressions, and the network can generate appropriate words / sentences at appropriate positions.
[0060] In the above case, the network only learns the style-related expressions, but cannot distinguish between different style expressions. In order to enable the final generation of navigation broadcast sentences with the expected style, the network needs to have the ability to distinguish between different style expressions. Considering that in related technologies, multiple sentences are generally spliced in training samples to enable the network to infer the content of the mask position through context information, thereby completing the restoration of the sentence, in the specification, it is considered that multiple sentences can not be spliced in the training samples, but the sentence to be restored and the corresponding style (such as gentle, lovely, humorous, etc.) are spliced, so that the network can correspond different style expressions according to the style included in the training sample.
[0061] In other words, the specification provides a training sample generation method, device, computer equipment and storage medium, obtains a navigation broadcast sentence with a style, and determines a style expression short sentence, a style expression word in the navigation broadcast sentence, and a style corresponding to the navigation broadcast sentence. The style expression short sentence is a short sentence that does not contain navigation broadcast information. The style expression short sentence is replaced by a first mark, and the style expression word is replaced by a second mark to obtain a processed navigation broadcast sentence. The processed navigation broadcast sentence and the style corresponding to the navigation broadcast sentence are spliced to obtain input information of a network to be trained. The navigation broadcast sentence with the style is used as a label, and the label and the input information are used as training samples of the network to be trained.
[0062] By replacing the style expression short sentences (i.e. the style expression related sentences above) and the style expression words (i.e. the style expression related words above) with different markers, the network can learn the use of the style related short sentence expressions and the style expression words in the navigation broadcast sentences. And by splicing the processed navigation broadcast sentences with the styles, the network can correspond the learned style related expressions with the matched styles, so that the model has the ability to generate navigation broadcast sentences with styles based on the input styles and sentences.
[0063] The present specification first shows a training sample generation method. After obtaining the training samples through the method, the training of the navigation broadcast sentence generation network can be completed through the training method of the navigation broadcast sentence generation network shown in the present specification. Further, the network obtained by training can be used to realize the personification rewriting of the navigation broadcast sentence through the navigation broadcast sentence generation method shown in the present specification.
[0064] Next, first, a training sample generation method shown in the present specification will be introduced.
[0065] As shown in Figure 1 , Figure 1 is a flowchart of a training sample generation method according to an exemplary embodiment shown in the present specification, which includes:
[0066] Step 101, obtaining a navigation broadcast sentence with style, and determining the style expression short sentences, the style expression words in the navigation broadcast sentence and the style corresponding to the navigation broadcast sentence.
[0067] Among them, the style expression short sentence is a short sentence that does not contain navigation broadcast information. The style expression word is a word related to the content of the style expression.
[0068] Specifically, by obtaining a navigation broadcast sentence with style, and labeling the required information in the navigation broadcast sentence, it is convenient to construct the training sample of the network according to the preset method in the subsequent process. Among them, the style expression short sentence and the style expression word are used to replace the preset marker in the subsequent process, so that the network can learn the style related expression, and the determined style is used to assist the network to distinguish the characteristics between the expressions of different styles, so as to generate the navigation broadcast sentence corresponding to the expected style according to the input of the expected style.
[0069] For the execution subject of the method shown in the present specification, it can be any computing device to execute, such as the server of the map application. The present specification does not limit the execution subject of the method in the present specification.
[0070] It should be noted that the method in the specification is proposed for the problem in the navigation scene, but the method can not only be used for the personification style rewriting of the navigation broadcast sentence, but also be applied to the personification style rewriting of other scenes, such as the style rewriting of teaching sentences, the personification style rewriting of arrival prompt sound and the like. It should be noted that when applied to the above scenes, the content related to the navigation field needs to be modified accordingly, such as modifying the navigation field words in the following text to the corresponding style field words, replacing the keywords in the following text with keywords related to other scenes, and the like.
[0071] Next, each noun involved in step 101 will be described.
[0072] The navigation broadcast sentence is the broadcast sentence mentioned in the preceding text for assisting the user in driving, reminding the user to walk according to the route recommended by the navigation function or reminding the user to avoid problems in the navigation process, which can be, for example, “100 meters ahead, there are red light violations and illegal photography”, “current speed 87, speed limit 80, please slow down”, “500 meters ahead of Zhuang, please straight ahead at the red and green light, walk on the left side of three lanes”.
[0073] The navigation broadcast sentence with style is the personified navigation broadcast sentence, and the style carried is the personified direction, which can be, for example, a navigation broadcast sentence with a gentle style, a navigation broadcast sentence with a cute style, a navigation broadcast sentence with a humorous style, and the like.
[0074] The short sentence is divided by punctuation marks, and the content between two punctuation marks is considered a short sentence. For example, “the speed limit of the road ahead is 80, slow down, I will worry about you” contains three short sentences: “the speed limit of the road ahead is 80”, “slow down” and “I will worry about you”.
[0075] The style expression short sentence is a short sentence that does not contain navigation broadcast content but is related to style, as described in the preceding text, such as the navigation broadcast sentence with style “the speed limit of the road ahead is 80, slow down, I will worry about you”, and the style expression short sentence is “slow down” and “I will worry about you”.
[0076] The style expression word generally refers to mood words, conjunction words (such as ah, ne, le, kuo, etc.) and some special non-mood words: such as fixed expressions, similar to baby. For example, the style expression word can be a stop word. The reason for choosing to replace the style expression word with the second mark is that different style expression words are used in different styles, and different style expression words may express different moods, so it is necessary to determine the style expression word in the navigation broadcast sentence.
[0077] It also needs to be explained that, since part or all of the style expression short sentences will be replaced by the first mark in step 103, in order to reduce the workload, the following processing can be carried out: considering that part or all of the style expression words in the style expression short sentences will not be replaced by the second mark in the subsequent text, only the style expression words in the broadcast short sentences except the style expression short sentences replaced by the first mark can be determined, that is, only the style expression words in the broadcast short sentences used to broadcast the navigation broadcast information can be determined.
[0078] The style corresponding to the navigation broadcast sentence is the personification style type to which the navigation broadcast sentence belongs, such as gentle, lovely, humorous, and the like.
[0079] After the nouns in step 101 are explained, the specific implementation method of step 101 will be exemplarily explained.
[0080] The process of obtaining the navigation broadcast sentence with style can be: obtaining the artificially rewritten navigation broadcast sentence with style; or first obtaining a certain number of artificially rewritten navigation broadcast sentences with style, and then performing data enhancement on the obtained sentences through synonym replacement and the like, so as to obtain more navigation broadcast sentences with style, thereby providing more samples for training. Of course, the process of obtaining the navigation broadcast sentence with style can also be other processes, and the present specification does not limit the process.
[0081] The way to determine the style expression words in the navigation broadcast sentence with style can be: since the style expression word is a general word expression, the style expression word can be pre-stored, and the style expression word is compared with each word in the navigation broadcast sentence, so as to determine the style expression word in the navigation broadcast sentence with style. Of course, the style expression word can also be labeled through other ways, such as through the pre-trained neural network for determining the style expression word to complete the labeling of the style expression word, and the present specification does not limit the specific determination method of the style expression word.
[0082] The way to determine the style expression short sentence in the navigation broadcast sentence with style can be:
[0083] First, it is considered that the style expression short sentence does not contain navigation broadcast information, so the style expression short sentence can be determined by the proportion of navigation scene field words. Specifically, the navigation broadcast sentence with style is first split into multiple broadcast short sentences according to punctuation, and the short sentences are processed by word segmentation (that is, the short sentences are split into multiple individual words, of course, word segmentation can also not be performed, and the navigation field words are directly searched in the short sentences, so as to determine the proportion of navigation field words), so as to determine the proportion of navigation field words in the short sentences. The short sentence with small navigation field word proportion is the style expression short sentence.
[0084] In other words, the determining the style expression short sentence in the navigation broadcast sentence includes: obtaining a navigation field word; based on a punctuation symbol, the navigation broadcast sentence with style is split into a plurality of broadcast short sentences; determining the proportion of the navigation field word in each broadcast short sentence, and taking the broadcast short sentence with a navigation field word proportion less than a preset proportion threshold as the style expression short sentence.
[0085] The navigation field word is a professional word that appears in a navigation scene.
[0086] Secondly, considering that the proportion of the word related to style (such as a style expression word) in the style expression short sentence will be high, whether it is a style expression short sentence can be determined based on the proportion of the style expression word in each broadcast short sentence. Specifically, each broadcast short sentence is first processed by word segmentation, and then the proportion of the style expression word in each short sentence is determined, and the broadcast short sentence with a high proportion of the style expression word is taken as the style expression short sentence.
[0087] Thirdly, the annotation of the style expression short sentence can also be performed based on a pre-trained neural network for determining whether it is a style expression short sentence.
[0088] Of course, the manner of determining the style expression short sentence is not limited to the above examples.
[0089] Finally, the manner of determining the style corresponding to the navigation broadcast sentence with style can be: obtaining the style of each navigation broadcast sentence annotated by a person, such as "there is a sharp turn ahead, slow down to pass the turn smoothly" corresponding to the style "gentle, lovely"; of course, other methods can also be used to determine the style corresponding to the navigation broadcast sentence, such as through a pre-trained neural network for judging the style type to annotate the style of the navigation broadcast sentence, and the manner of determining the style corresponding to the navigation broadcast sentence is not limited in the method shown in the specification.
[0090] In step 103, the style expression short sentence is replaced with a first mark, and the style expression word is replaced with a second mark to obtain a processed navigation broadcast sentence.
[0091] The first mark is used to represent a short sentence that needs to be restored by the network to be trained, and the second mark is used to represent a word that needs to be restored by the network to be trained.
[0092] As described above, in order to enable the network to learn the difference between the two different types of style-related expressions, different marks are needed to replace the style expression short sentence and the style expression word. Here, the replacement into the mark is actually a mask processing, and the goal of network training is to restore the words or sentences at the mask processing position.
[0093] The to-be-trained network is also a pre-trained language network, and the task to be performed in the training process of the network (the pre-training and post-pre-training fine-tuning process) is to restore the content replaced by the mark, and the relationship between words is learned in the restoration process in the pre-training process. The method shown in the specification changes the way of randomly replacing the mark in the related art, and replaces the words with specific content, so that the fine-tuning process (i.e., the process of updating the pre-trained language network) can learn the style-related expressions in the restoration process.
[0094] The first mark and the second mark can be any mark, but they must be different marks. For example, the first mark can be [MASK], and the second mark can be a space; the first mark can be [MASK1], and the second mark can be [MASK2], and so on. The form of the first mark and the second mark is not limited in the specification.
[0095] It should be noted that in order to improve the network training efficiency and enable the network to quickly learn the style-related expressions according to the context information of the training sample input, the style expression short sentence style expression words in the navigation broadcast short sentence with style can be selectively replaced, i.e., the style expression short sentence style expression words are replaced with the corresponding marks with a certain probability. For example, there are 5 style expression words in the original sentence, and only two of them are replaced with the second mark. There are 3 style expression short sentences in the original sentence, and only two of them are replaced with the first mark. In this way, the amount of context information in the training sample can be increased, and the network can quickly infer the word or sentence at the mark position according to the input, so that the network training effect is better.
[0096] In other words, step 103 includes: selecting at least one style expression short sentence from the style expression short sentences included in the navigation broadcast sentence, and replacing the selected style expression short sentence with the first mark; selecting at least one style expression word from the style expression words included in the navigation broadcast sentence, and replacing the selected style expression word with the second mark.
[0097] Still taking “There is a sharp turn ahead, slow down to safely pass the turn” as an example, in the case where the first mark is [MASK] and the second mark is a space, the processed navigation broadcast sentence processed by the above method can be “There is a sharp turn ahead, slow down to safely pass the turn”.
[0098] In addition, to increase the number of training samples, after obtaining the processed navigation announcement statements, some words in the navigation announcement statements can be replaced with synonyms. For example, "steadily" in the statement "There is a sharp turn ahead. Slow down to pass the turn steadily." can be replaced with "smoothly". This can greatly increase the number of training samples, thereby making the fine-tuning effect better, and the network can more easily learn similar expressions, making the diversity of the finally generated navigation announcement statements with styles stronger (that is, expressing the same meaning in multiple different ways).
[0099] Step 105: Generate the training samples for the to-be-trained network.
[0100] The training samples include feature values and labels. The feature values of the training samples include the processed navigation announcement statements and the styles corresponding to the navigation announcement statements. The labels of the training samples include the navigation announcement statements with styles.
[0101] As described above, in order to enable the network to distinguish the differences in expressions between different styles, after obtaining the processed navigation announcement statements, the statement needs to be concatenated with the style to obtain the input (i.e., the feature value) of the training sample, and the original sentence of the navigation announcement statement with style is used as the label to complete the generation process of the training samples for supervised learning.
[0102] The process of obtaining the feature values can be achieved by concatenating the processed navigation announcement statements and the above styles. The concatenation here adopts the method of concatenating different sentences in the training samples of the pre-trained language network in the related technology, that is, the processed navigation announcement statement and the style are connected by [SEP]. Of course, the present specification does not limit the specific concatenation method, as long as the network can know the style corresponding to the statement without increasing parameters.
[0103] Still taking "There is a sharp turn ahead. Slow down to pass the turn steadily." as an example, when the first marker is [MASK], the second marker is a space, and the concatenation is completed by [SEP], the processed navigation announcement statement processed in the above manner can be "Remind of a sharp turn ahead. Slow down [MASK][SEP] Copywriting style: gentle, cute". Thus, the input of the training sample is obtained.
[0104] The input of the training sample, that is, the feature of the training sample, is opposite to the label and is used for training the supervised network.
[0105] In addition, it also needs to be explained that the artificially rewritten navigation broadcast sentence with style in the related art is generally a navigation broadcast sentence associated with a certain IP, such as a navigation broadcast sentence corresponding to star A or a navigation broadcast sentence corresponding to cartoon character B. In this case, the artificially rewritten navigation broadcast sentence (i.e., the navigation broadcast sentence obtained in step 101) generally contains information of a special broadcast entity (i.e., a broadcast subject or an entity associated with the broadcast subject), such as “Star A reminds you that there is a sharp turn ahead, slow down to safely negotiate the turn”. In the navigation broadcast sentence corresponding to cartoon character B, the village where cartoon character B lives or the name of the family member of cartoon character B may be mentioned, and the village and the name of the family member of cartoon character B can also be regarded as a broadcast entity.
[0106] In the above case, in order to enable the network to distinguish between the entity and the content related to the style expression, a broadcast entity related label can also be added to the input of the constructed training sample, such as replacing the broadcast entity with a third label, or inputting the information of the broadcast entity through splicing, so that the network can distinguish between the broadcast entity and the style related expression, and fully utilize the performance of the pre-trained language network.
[0107] In other words, the method further comprises: determining a broadcast entity in the navigation broadcast sentence; the broadcast entity includes a broadcast subject and entity information (which can also be understood as content that the broadcast subject will mention and other IPs or the broadcast subject will not mention) associated with the broadcast subject; and the process of obtaining the feature value of the training sample comprises: splicing the processed navigation broadcast sentence, the style corresponding to the navigation broadcast sentence, and the determined broadcast entity to obtain the feature value of the network to be trained.
[0108] Still using the example of “Star A reminds you that there is a sharp turn ahead, slow down to safely negotiate the turn” in the above, in the case where the first label is [MASK], the second label is a space, the splicing is completed through [SEP], and the broadcast entity is marked through [SEP], the input of the processed training sample can be “Star A reminds you that there is a sharp turn ahead, slow down to safely negotiate the turn [MASK] [SEP] Style: Gentle, lovely [SEP] Broadcast entity: Star A”.
[0109] In addition, it also needs to be explained that in the navigation scenario, in order to enable the network to learn the knowledge related to navigation, information related to navigation such as a center sentence and a keyword can also be added to the input of the training sample. And the corresponding navigation sentence is marked in the training sample through the replacement method or the splicing method.
[0110] Of course, considering that replacing all the content that needs to be labeled with different labels might render the contextual information in the training sample input almost non-existent, potentially hindering network training, we have already replaced the style expression phrases and style expression words with the first and second labels respectively. Furthermore, to provide clues for the content the network can reconstruct, we can use a concatenation method to label the corresponding navigation statements in the training samples.
[0111] The central sentence is the unstyled navigation voiceover corresponding to the styled navigation voiceover statement, that is, one or more short sentences containing the most navigation voiceover information. For example, the central sentence of "There is a sharp turn ahead, slow down so you can pass the turn safely" is "There is a sharp turn ahead".
[0112] Keywords are words related to the navigation broadcast scenario, that is, the content that the sentence most wants to remind the user of. For example, the keywords for "There is a sharp turn ahead, slow down so you can pass the turn steadily" could be "steadily" and "remind".
[0113] The following will use the example of tagging keywords through concatenation to illustrate the above process. The tagging method for the central sentence is similar and will not be repeated here. In other words, the method also includes: determining keywords in the navigation broadcast statement; the keywords are words in the navigation broadcast statement that are related to the navigation broadcast scenario; the process of obtaining the feature values of the training samples includes: concatenating the style corresponding to the processed navigation broadcast statement with the determined keywords to obtain the feature values of the network to be trained.
[0114] Using the example of "Celebrity A reminds you that there is a sharp turn ahead, so slow down to make it through smoothly," with the first marker being [MASK], the second marker being a space, and the concatenation being done through [SEP], and the broadcast entity, central sentence, and keywords all being marked with [SEP], the input of the processed training sample could be "Celebrity A reminds you to slow down on the sharp turn ahead [MASK][SEP] Text style: gentle, cute [SEP] Broadcast entity: Celebrity A [SEP] Central sentence: There is a sharp turn ahead [SEP] Keywords: smooth, reminder".
[0115] The above method yields the training samples needed to train the navigation broadcast statement generation network. Repeating the above process will produce the training sample set.
[0116] The training method for a navigation broadcast statement generation network, as illustrated in the embodiments of this specification, will be described next.
[0117] like Figure 2 As shown, Figure 2is a flowchart of a training method of a navigation broadcast sentence generation network according to an exemplary embodiment of the present specification, comprising:
[0118] In step 201, each training sample in the training sample set is obtained through the aforementioned training sample generation method.
[0119] In step 203, the input information in the training sample is taken as the input of the pre-trained language network, and the language network is updated according to the input information and the label in the training sample, to obtain a navigation broadcast sentence generation network.
[0120] The navigation broadcast sentence generation network is used to generate a navigation broadcast sentence with a style based on a navigation broadcast sentence without a style and an expected style.
[0121] The same name in steps 201 and 203 as in the foregoing method will not be repeated, and the specific meaning is described in the foregoing.
[0122] The pre-trained language network is the network mentioned in the foregoing, and its training process is generally first unsupervised learning through a large amount of corpus, so that the network can learn semantic and grammatical knowledge, and then fine-tuning is completed through a labeled training sample. The present specification does not change the training process of the pre-trained language network in the related art, but changes the meaning represented by each part of the training sample in the fine-tuning process. Steps 201 and 203 are the process of fine-tuning the training sample in the foregoing.
[0123] For example, it can be BART (Bidirectional and Auto-Regressive Transformers) in PLMs, and of course, GPT (Generative Pre-Training), Palm (Scaling Language Modeling with Pathways), and T5 (Transfer Text-to-Text Transformer) can also be used. The language network used in the present specification is not limited.
[0124] In addition, it also needs to be explained that in the input information of the training sample, not only the corresponding style information is included, but also the broadcast entity, the center sentence and / or the keyword, in order to enable the network to better learn the above-mentioned content, the training can be performed multiple times, and each time one or two contents are taken as the training target to perform training, so that the network can gradually learn different knowledge step by step.
[0125] For example, in the first stage, the center sentence and the keywords in the training sample can be removed first, and the network can only learn the entity and the style-related expression. In the second stage, the keywords in the training sample can be removed, and on the basis of the network having learned the entity and the style-related expression, further learning and knowledge related to the center sentence can be performed to enable the network to learn the personification rewriting of the center sentence, so that the network knows what content needs to be retained in the rewriting process. In the third stage, on the basis of the obtained completed training sample and on the basis of the network having learned the entity, the style-related expression and the center sentence, further learning and knowledge related to the keywords can be performed to determine the general direction of the content that needs to be completed by the network, so that the sentence generated by the network is more in line with the expectation.
[0126] Through the above method, the training of the navigation broadcast sentence generation network is completed.
[0127] Next, a navigation broadcast sentence generation method shown in an embodiment of the present specification will be described.
[0128] As shown in Figure 3 , Figure 3 is a flowchart of a navigation broadcast sentence generation method according to an exemplary embodiment shown in the present specification, which includes:
[0129] Step 301, obtaining network input information.
[0130] The network input information includes a first navigation broadcast sentence and a style of a navigation broadcast sentence expected to be generated, and the first navigation broadcast sentence is a styleless navigation broadcast sentence from which the style expression words are removed.
[0131] Step 303, inputting the network input information into the navigation broadcast sentence generation network to obtain a second navigation broadcast sentence output by the navigation broadcast sentence generation network.
[0132] The second navigation broadcast sentence is a style sentence corresponding to the first navigation broadcast sentence. The navigation broadcast sentence generation network can be a network trained by the method described above.
[0133] It should be noted that in addition to the above content, the network input information can also include a first mark and a second mark, so that the network can know where to generate the style expression short sentence and where to generate the style expression word.
[0134] In the case of generating a navigation broadcast sentence associated with a certain IP, a broadcast entity corresponding to the IP can also be added to the input information, so that the network can automatically generate a navigation broadcast sentence associated with a certain IP to realize the automatic generation of the navigation broadcast sentence.
[0135] In addition, the center sentence and the keyword can be added in the input information to assist in generating the navigation broadcast sentence.
[0136] By the above method, the style expression short sentence and the style expression word are replaced by different marks, so that the network can distinguish the style expression word and the style expression short sentence, and meet the rewriting demand of the two different forms of the whole sentence rewriting and the hanging generation in the personification navigation broadcast sentence. By introducing the broadcast entity, the style, the center sentence and the keyword layer by layer, the prompt matched with the target of each training stage is constructed, so that the network gradually improves the learning ability, and better learning effect is achieved. Through the prediction stage prompt rhetoric theme, special entity, style, [MASK] position and other rewriting methods, the generated content has the characteristics of style specification, reasonable content, diverse rhetoric, controllable content (that is, the style and navigation information of the generated content can be controlled), and better meets the landing of the business.
[0137] Next, the method of sample generation and network training will be described through a specific embodiment.
[0138] Step 1: training data
[0139] 1. Construct a training data pair, <original rhetoric (i.e. navigation broadcast sentence without style), personification rewriting rhetoric>, such as: <there is a sharp turn ahead, star A reminds you, there is a sharp turn ahead, slow down, and you can safely turn the corner>
[0140] The personification rewriting rhetoric can be the corresponding personification rewriting rhetoric (i.e. navigation broadcast sentence with style) obtained by artificial rewriting.
[0141] 2. Mark the broadcast entity, style and center sentence in the personification rewriting rhetoric; the style (type) can be marked by part of artificial marking, or marked by a trained style classification model. For example, in "star A, there is a sharp turn ahead, slow down, and you can safely turn the corner", the broadcast entity is "star A", the style is "gentle and lovely", and the center sentence is "there is a sharp turn ahead" (i.e. navigation broadcast sentence without style).
[0142] Step 2: non-navigation text segment identification
[0143] 1. Construct a navigation field word library by performing word segmentation and word frequency statistics on the original rhetoric.
[0144] 2. Divide the personification rewriting rhetoric by punctuation, such as splitting the above example into ['star A reminds you', 'there is a sharp turn ahead','slow down', 'you can safely turn the corner']
[0145] 3. Through the navigation field lexicon constructed in 1, according to the proportion of navigation high-frequency words in each sentence, it is judged whether each sentence belongs to a navigation text segment or a style expression sentence. For example, "Star A reminds you", 'Slow down', 'Can smoothly turn the corner' belong to style expression sentences, and 'There is a sharp turn ahead' belongs to navigation text segments.
[0146] Step 3: Style rhetoric mask
[0147] 1. For the broadcast sentence in step 2, perform sentence-by-sentence style rhetoric mask processing.
[0148] 2. If the broadcast sentence belongs to a style expression sentence, replace the entire sentence with a special mark [MASK] with a certain probability; the rest, after being segmented, are processed to remove style-related words such as adverbs and conjunctions, and the remaining words are separated by spaces. For example, 'Can smoothly turn the corner' is replaced with [MASK] with a certain probability, or processed into "Can turn the corner".
[0149] 3. If the sub-sentence data is a navigation text segment, first perform segmentation, remove style-related words such as adverbs and conjunctions, and only keep navigation-related words separated by spaces, such as "There is a sharp turn ahead" becomes "There is a sharp turn ahead".
[0150] 4. The processed sentences are connected by spaces, and broadcast entity prompts and style prompts are added, connected by [SEP]. For example, "Star A reminds you, there is a sharp turn ahead, slow down, can smoothly turn the corner" is processed as follows: "Star A reminds you, there is a sharp turn ahead, slow down, can smoothly turn the corner" is processed as follows: "Star A reminds you, there is a sharp turn ahead, slow down, can smoothly turn the corner".
[0151] Step 4: Text enhancement
[0152] 1. Through the prompt input of PLMs constructed in step 3, the corresponding label rhetoric is constructed.
[0153] 2. The corresponding personification rhetoric rewriting text is segmented, non-navigation field words are replaced with a certain probability of synonym replacement, and data enhancement is performed, such as "Star A reminds you, there is a sharp turn ahead, slow down, can smoothly turn the corner" is replaced with "Star A reminds you, there is a sharp turn ahead, slow down, can smoothly turn the corner".
[0154] 3. The final model input is < "Star A reminds you, there is a sharp turn ahead, slow down, can smoothly turn the corner [MASK] [SEP] copy style: gentle, lovely [SEP] special entity: Star A", "Star A reminds you, there is a sharp turn ahead, slow down, can smoothly turn the corner" >.
[0155] 4. The prompt construction method constructed in this scheme learns stylistic expressions such as conjunctions, modal particles, and titles in the anthropomorphic speech rewriting and sentence replacement form through spaces; learns anthropomorphic speech rewriting of the attached content through [MASK]; and enables the model to distinguish speech styles and identify special entities in the speech through "[SEP] copywriting style" and "[SEP] special entities".
[0156] Step 5: Construct the presentation layer by layer, including entities, style, central sentence, and keywords.
[0157] 1. This scheme adopts the idea of training and learning in stages and layers. By broadcasting entities, styles, central sentences and keywords in the prompt layer by layer, different content is emphasized in different stages.
[0158] 2. The training data pairs constructed in step 4 are used to train the first-stage model. The first-stage model focuses on learning the style of anthropomorphic rewriting and the broadcasting entity.
[0159] 3. Based on the prompt built in step 4, add a central sentence prompt, such as "Celebrity A reminds you to slow down at a sharp turn ahead [MASK][SEP] Copywriting style: gentle, cute [SEP] Special entity: Celebrity A [SEP] Central sentence: There is a sharp turn ahead", for training the second-stage model. The second-stage model focuses on learning the central sentence rewritten with anthropomorphic language.
[0160] 4. Based on the above, add the keyword prompt, such as "Celebrity A reminds you to slow down at a sharp turn ahead [MASK][SEP] Copywriting style: gentle, cute [SEP] Special entity: Celebrity A [SEP] Central sentence: There is a sharp turn ahead [SEP] Keywords: smooth, reminder", for training the third-stage model. The third-stage model focuses on learning the theme of rewriting anthropomorphic language.
[0161] It should be noted that in this manual, "Celebrity A" does not mean that the actual navigation broadcast statement contains the three characters "Celebrity A", but is used to represent the name of any celebrity.
[0162] This solution, through steps 1-5, ultimately obtains a text generation model for TBT anthropomorphic broadcast scenarios through phased training. In the prediction phase, controllable and diverse text generation is achieved by modifying various aspects of the prompt, such as the topic, broadcast entity, style, and [MASK] position. For example, "Sharp turn ahead [MASK] [SEP] Copywriting style: Gentle [SEP] Special entity: Celebrity B [SEP] Central sentence: Sharp turn ahead [SEP] Keywords: Concern."
[0163] Corresponding to the embodiments of the foregoing method, the specification also provides embodiments of an apparatus and a terminal to which the apparatus is applied.
[0164] As Figure 4 shown, Figure 4 is a block diagram of a training sample generation apparatus according to an exemplary embodiment, the apparatus comprising:
[0165] The sentence acquisition module 410 is configured to acquire a navigation broadcast sentence with a style, and determine a style expression phrase, a style expression word, and a style corresponding to the navigation broadcast sentence in the navigation broadcast sentence. The style expression phrase is a phrase that does not contain navigation broadcast information. The style expression word is a word related to style expression content.
[0166] The label replacement module 420 is configured to replace the style expression phrase with a first label, and replace the style expression word with a second label, to obtain a processed navigation broadcast sentence. The first label is used to represent a phrase that needs to be restored by a to-be-trained network. The second label is used to represent a word that needs to be restored by the to-be-trained network.
[0167] The sample generation module 430 is configured to generate a training sample of the to-be-trained network. The training sample includes a feature value and a label. The feature value of the training sample includes the processed navigation broadcast sentence and the style corresponding to the navigation broadcast sentence. The label of the training sample includes the navigation broadcast sentence with the style.
[0168] In an optional embodiment, the sentence acquisition module 410 is configured to acquire a navigation broadcast sentence with a style, and acquire a navigation field word. Based on a punctuation symbol, the navigation broadcast sentence with the style is split into a plurality of broadcast phrases. The proportion of the navigation field word in each broadcast phrase is determined. A broadcast phrase in which the proportion of the navigation field word is less than a preset proportion threshold is determined as a style expression phrase. The style expression word in the navigation broadcast sentence and the style corresponding to the navigation broadcast sentence are determined.
[0169] In an optional embodiment, the apparatus further includes a broadcast entity determination module 340 (not shown in the figure) configured to determine a broadcast entity in the navigation broadcast sentence. The broadcast entity includes a broadcast subject and entity information associated with the broadcast subject. The sample generation module 430 is configured to generate a training sample of the to-be-trained network. The training sample includes a feature value and a label. The feature value of the training sample includes the processed navigation broadcast sentence, the determined broadcast entity, and the style corresponding to the navigation broadcast sentence. The label of the training sample includes the navigation broadcast sentence with the style.
[0170] In an optional embodiment, the apparatus further comprises a keyword determining module 340 (not shown in the figure) configured to determine a keyword in the navigation broadcast sentence; the keyword is a word in the navigation broadcast sentence that is related to the navigation broadcast scene; and a sample generating module 430 configured to generate a training sample of the network to be trained, the training sample comprising a feature value and a label, the feature value of the training sample comprising the processed navigation broadcast sentence, the determined keyword, and a style corresponding to the navigation broadcast sentence, and the label of the training sample comprising the navigation broadcast sentence with the style.
[0171] In an optional embodiment, a mark replacing module 420 is configured to filter at least one style expression phrase from the style expression phrases included in the navigation broadcast sentence, and replace the filtered style expression phrase with a first mark; and filter at least one style expression word from the style expression words included in the navigation broadcast sentence, and replace the filtered style expression word with a second mark.
[0172] As shown in Figure 5 , Figure 5 is a block diagram of a training apparatus of a navigation broadcast sentence generation network according to an exemplary embodiment of the present specification, the apparatus comprising:
[0173] a sample obtaining module 510 configured to obtain each training sample in a training sample set by using the foregoing training sample generation method;
[0174] a network training module 520 configured to use the input information in the training sample as an input of a pre-trained language network, update the language network according to the input information and the label in the training sample, and obtain a navigation broadcast sentence generation network; the navigation broadcast sentence generation network is configured to generate a navigation broadcast sentence with a style based on a navigation broadcast sentence without the style and an expected style.
[0175] As shown in Figure 6 , Figure 6 is a block diagram of a generation apparatus of a navigation broadcast sentence according to an exemplary embodiment of the present specification, the apparatus comprising:
[0176] an input information obtaining module 610 configured to obtain network input information, the network input information comprising a first navigation broadcast sentence and a style of a navigation broadcast sentence expected to be generated, the first navigation broadcast sentence being a navigation broadcast sentence without style, and the navigation broadcast sentence generation network being obtained by training using the foregoing training method of the navigation broadcast sentence generation network;
[0177] The sentence generation module 620 is configured to input the network input information into a navigation broadcast sentence generation network to obtain a second navigation broadcast sentence output by the navigation broadcast sentence generation network, the second navigation broadcast sentence being a style sentence corresponding to the first navigation broadcast sentence.
[0178] The implementation process of the functions and roles of each module in the above device is specifically described in the implementation process of the corresponding steps in the above method, which will not be repeated here.
[0179] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the part of the method embodiment. The device embodiments described above are only schematic, and the modules described as separate components can or can not be physically separated, and the components displayed as modules can or can not be physical modules, that is, they can be located in one place, or distributed on multiple network modules. According to the actual needs, some or all of the modules can be selected to achieve the purpose of the scheme of the present specification. Those skilled in the art can understand and implement it without creative labor.
[0180] As shown in Figure 7 , Figure 7 A hardware structure diagram of a computer device in which the embodiment training sample generation device or the navigation broadcast sentence generation device or the training device of the navigation broadcast sentence generation network is shown. The device can include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040 and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030 and the communication interface 1040 are connected to each other through the bus 1050 for communication connection within the device.
[0181] The processor 1010 can be implemented in the form of a general-purpose CPU (Central Processing Unit, central processor), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, etc., for executing related programs to implement the technical solutions provided by the embodiments of the present specification.
[0182] The memory 1020 can be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 can store an operating system and other application programs, and when the technical solutions provided by the embodiments of the present specification are implemented by software or firmware, the related program codes are saved in the memory 1020 and are called and executed by the processor 1010.
[0183] The input / output interface 1030 is configured to connect an input / output module to realize information input and output. The input / output module can be configured in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. The input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.
[0184] The communication interface 1040 is configured to connect a communication module (not shown in the figure) to realize the communication interaction between the device and other devices. The communication module can realize communication through a wired manner (such as USB, network cable, etc.) or through a wireless manner (such as mobile network, WIFI, Bluetooth, etc.).
[0185] The bus 1050 includes a channel to transmit information between various components (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040) of the device.
[0186] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device can also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device can also only include the components necessary to implement the embodiments of the present specification, and does not have to include all the components shown in the figure.
[0187] The embodiments of the present specification also provide a computer readable storage medium having a computer program stored thereon, which is executed by a processor to implement the above-mentioned training sample generation method, or the navigation broadcast sentence generation method, or the training method of the navigation broadcast sentence generation network.
[0188] Computer-readable media includes permanent and non-permanent, moveable and non- moveable media that can be implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other memory technology, compact disc read only memory (CD-ROM), digital versatile disks (DVDs) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device. According to the definitions provided herein, computer readable media does not include transitory computer readable medium, such as a modulated data signal and a carrier wave.
[0189] It is also important to note that the terms "comprises", "comprising", or other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0190] The above description of certain examples of the application has been presented for the purposes of illustration and description. Other examples are within the scope and range of equivalents of the claims. In some cases, acts or steps can be performed in an order different from that of the examples, and / or at least partially concurrently, and still accomplish the desired results. Also, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous.
Claims
1. A training sample generation method, comprising: obtaining a navigation broadcast sentence with a style, and determining a style expression phrase, a style expression word in the navigation broadcast sentence, and a style corresponding to the navigation broadcast sentence; the style expression phrase is a phrase that does not contain navigation broadcast information, and the style expression phrase is determined by the proportion of navigation scene field words; the style expression word is a word related to style expression content; replacing the style expression phrase with a first mark and the style expression word with a second mark to obtain a processed navigation broadcast sentence; the first mark is used to represent a phrase to be restored by a to-be-trained network, and the second mark is used to represent a word to be restored by the to-be-trained network; generating a training sample of the to-be-trained network, the training sample comprising a feature value and a label, the feature value of the training sample comprising the processed navigation broadcast sentence and the style corresponding to the navigation broadcast sentence, and the label of the training sample comprising the navigation broadcast sentence with the style.
2. The method of claim 1, wherein the determination of the style expression phrase in the navigation broadcast sentence comprises: obtaining navigation field words; based on punctuation, the navigation broadcast sentence with the style is split into multiple broadcast phrases; determining the proportion of navigation field words in each broadcast phrase, and taking a broadcast phrase in which the proportion of navigation field words is less than a preset proportion threshold as a style expression phrase.
3. The method of claim 1, The method further comprises: determining a broadcast entity in the navigation broadcast sentence; the broadcast entity comprises a broadcast subject and entity information associated with the broadcast subject; the feature value of the training sample is obtained by splicing the processed navigation broadcast sentence, the style corresponding to the navigation broadcast sentence, and the determined broadcast entity to obtain the feature value of the to-be-trained network.
4. The method of claim 1, The method further comprises: determining a keyword in the navigation broadcast sentence; the keyword is a word related to the navigation broadcast scene in the navigation broadcast sentence; the feature value of the training sample is obtained by splicing the processed navigation broadcast sentence, the style corresponding to the navigation broadcast sentence, and the determined keyword to obtain the feature value of the to-be-trained network.
5. The method of claim 1, wherein the replacement of the style expression phrase with the first mark and the replacement of the style expression word with the second mark comprise: selecting at least one style expression phrase from the style expression phrases included in the navigation broadcast sentence, and replacing the selected style expression phrase with the first mark; selecting at least one style expression word from the style expression words included in the navigation broadcast sentence, and replacing the selected style expression word with the second mark.
6. A navigation broadcast sentence generation network training method, comprising: obtaining each training sample in a training sample set by the method of any one of claims 1-5; using the feature value in the training sample as an input of a pre-trained language network, updating the language network according to the feature value and the label in the training sample to obtain a navigation broadcast sentence generation network. The navigation broadcast sentence generation network is used for generating a navigation broadcast sentence with a style based on a navigation broadcast sentence without a style and an expected style.
7. A method for generating a navigation broadcast sentence, comprising: obtaining network input information, the network input information comprising: a first navigation broadcast sentence, and a style of a navigation broadcast sentence expected to be generated, the first navigation broadcast sentence being a navigation broadcast sentence without a style by removing style expression words; the navigation broadcast sentence generation network being obtained by training through the method of claim 6; inputting the network input information into the navigation broadcast sentence generation network to obtain a second navigation broadcast sentence output by the navigation broadcast sentence generation network, the second navigation broadcast sentence being a navigation broadcast sentence with a style corresponding to the first navigation broadcast sentence.
8. An apparatus for generating training samples, comprising: a sentence obtaining module configured to obtain a navigation broadcast sentence with a style, and determine a style expression phrase, a style expression word, and a style corresponding to the navigation broadcast sentence in the navigation broadcast sentence; the style expression phrase being a phrase without navigation broadcast information, the style expression phrase being determined by a proportion of navigation scene field words; the style expression word being a word related to style expression content; a label replacing module configured to replace the style expression phrase with a first label, and replace the style expression word with a second label to obtain a processed navigation broadcast sentence; the first label being used to represent a phrase to be restored by a to-be-trained network, and the second label being used to represent a word to be restored by the to-be-trained network; a sample generating module configured to generate a training sample of the to-be-trained network, the training sample comprising a feature value and a label, the feature value of the training sample comprising the processed navigation broadcast sentence and the style corresponding to the navigation broadcast sentence, and the label of the training sample comprising the navigation broadcast sentence with the style.
9. An apparatus for training a navigation broadcast sentence generation network, comprising: a sample obtaining module configured to obtain each training sample in a training sample set through the method of any one of claims 1-5; a network training module configured to take input information in the training sample as input of a pre-trained language network, update the language network according to the input information and a label in the training sample, and obtain a navigation broadcast sentence generation network; the navigation broadcast sentence generation network is used for generating a navigation broadcast sentence with a style based on a navigation broadcast sentence without a style and an expected style.
10. An apparatus for generating a navigation broadcast sentence, comprising: an input information obtaining module configured to obtain network input information, the network input information comprising: a first navigation broadcast sentence, and a style of a navigation broadcast sentence expected to be generated, the first navigation broadcast sentence being a navigation broadcast sentence without a style by removing style expression words; the navigation broadcast sentence generation network being obtained by training through the method of claim 6; The sentence generation module is configured to input the network input information into a navigation broadcast sentence generation network to obtain a second navigation broadcast sentence output by the navigation broadcast sentence generation network, the second navigation broadcast sentence being a style sentence corresponding to the first navigation broadcast sentence.
11. A computer device comprising: a processor; a memory for storing processor-executable instructions; wherein the processor, by running the executable instructions, implements the method of any one of claims 1-7.
12. A computer-readable storage medium having stored thereon computer instructions, the computer instructions, when executed by a processor, implementing the method of any one of claims 1-7.
Citation Information
Patent Citations
Text conversion method and device and storage medium
CN110287461A