Role annotation and speech synthesis method, device, system and storage medium

The method enhances role tagging in multi-role speech synthesis by integrating candidate role interactions and context through vector mapping and classification, improving efficiency and accuracy.

CN115062585BActive Publication Date: 2025-07-15DATABAKER (BEIJING) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210389200.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-13
Publication Date
2025-07-15
Estimated Expiration
2042-04-13

AI Technical Summary

Technical Problem

In the prior art, role labeling is inefficient, computational resources are consumed, and the model is not adaptable enough, so it is impossible to accurately handle the interaction between diversified text and candidate roles.

Method used

Map the label text, target dialogue and candidate role list as feature vectors, and fuse each vector relationship through the encoder module to pool and classify to determine the role labeling results, support the input of multiple candidate roles at once and improve efficiency.

Benefits of technology

It improves the accuracy and efficiency of role labeling, enhances the adaptability of the model, reduces computing resource consumption, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115062585B_ABST
    Figure CN115062585B_ABST
Patent Text Reader

Abstract

The present invention provides a method, apparatus, system and storage medium for role annotation and speech synthesis, including: obtaining an annotated text, a target dialogue and a list of candidate roles; mapping the annotated text character by character to a feature vector one by one; mapping the target dialogue character by character to a feature vector one by one; pooling the vectors in the dialogue vector sequence; for each candidate role name in the list of candidate roles, mapping the candidate role name character by character to a feature vector one by one; pooling the vectors in the name vector sequence; inputting the text vector sequence, the dialogue vector and the role vector sequence into an encoder module; classifying the encoding results corresponding one by one to at least one candidate role name belonging to each candidate role group through a classification module; and determining the role annotation result of the target dialogue based on the classification results of at least one candidate role group. It can save computing resources, has high annotation efficiency and good user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular, to a method, device, system and storage medium for role annotation, and a method, device, system and storage medium for speech synthesis. Background Art

[0002] With the increasing demand of users for resources such as audiobooks, the practice of relying on manual recording of audiobook corpus can no longer meet the demand. Therefore, it is particularly important to develop technologies / tools / systems that support multi-role and multi-emotion speech synthesis.

[0003] Text role annotation is a technology that automatically determines the role (speaker) to which the dialogue in a text such as a novel belongs based only on limited text information, and is an important part of multi-role and multi-emotion speech synthesis. Text role annotation requires the target dialogue and its context as input information, and in addition, a list of candidate roles needs to be provided, including formal names and the corresponding relationships with aliases, etc.

[0004] In the prior art, role annotation is usually regarded as a task of scoring candidate roles or other equivalent tasks. For example, Chen Y et al. constructed a candidate role scoring network based on the pre-trained language representation model BERT (Bidirectional Encoder Representation from Transformers). On the basis of obtaining the representations of the context, dialogue, and candidate roles, it calculates the scores of each candidate role in the candidate role list in sequence, and finally takes the candidate role with the highest score as the final role. This solution has the following problems: 1) Only one candidate role can be input for scoring each time in the model. The more candidate roles there are, the more scoring times are required, with low efficiency and high computational resource consumption; 2) For diverse texts, the method of selecting the final candidate role name using the "proximity principle" is obviously not very accurate; 3) The interaction between the context, target dialogue, and each candidate role is not sufficient, especially the interaction between each candidate role is hardly considered, which affects the adaptability of the model to a certain extent. Summary of the Invention

[0005] In order to at least partially solve the above problems existing in the prior art, a method, device, system and storage medium for role annotation, and a method, device, system and storage medium for speech synthesis are provided.

[0006] According to one aspect of the present invention, there is provided a role annotation method, including: obtaining an annotated text, a target dialogue, and a candidate role list, where the annotated text includes the target dialogue and the context of the target dialogue, the candidate role list includes one or more candidate role names appearing in the annotated text, and the one or more candidate role names are divided into at least one candidate role group; mapping the annotated text character by character into feature vectors one by one to obtain a text vector sequence; mapping the target dialogue character by character into feature vectors one by one to obtain a dialogue vector sequence; performing pooling on the vectors in the dialogue vector sequence to obtain a dialogue vector; for each candidate role name in the candidate role list, mapping the candidate role name character by character into feature vectors one by one to obtain a name vector sequence corresponding to the candidate role name; performing pooling on the vectors in the name vector sequence to obtain a role vector corresponding to the candidate role name; inputting the text vector sequence, the dialogue vector, and the role vector sequence into an encoder module to obtain an encoded result sequence, where the encoded result sequence includes encoded results corresponding one by one to the vectors in the text vector sequence, the dialogue vector, and the role vector sequence, the encoded result of any vector includes the correlation information between the vector and other vectors, and the role vector sequence includes role vectors corresponding one by one to all candidate role names in the candidate role list; for each candidate role group in the at least one candidate role group, classifying the encoded results corresponding one by one to at least one candidate role name belonging to the candidate role group through a classification module to obtain the classification result of the candidate role group, where the classification result of any role is used to indicate whether the speaker of the target dialogue belongs to the candidate role group; determining the role annotation result of the target dialogue based on the classification results of the at least one candidate role group.

[0007] Exemplarily, for each candidate role group in the at least one candidate role group, classifying the encoded results corresponding one by one to at least one candidate role name belonging to the candidate role group through a classification module to obtain the classification result of the candidate role group includes: for each candidate role group in the at least one candidate role group, performing pooling on the encoded results corresponding one by one to at least one candidate role name belonging to the candidate role group to obtain a pooled encoded result; inputting the pooled encoded result into the classification module to obtain the classification result of the candidate role group.

[0008] Exemplarily, before pooling the encoding results corresponding one-to-one to at least one candidate role name belonging to each candidate role group in at least one candidate role group to obtain the pooled encoding results, the method further includes: for each candidate role group in at least one candidate role group, multiplying the encoding result sequence by the mask vector corresponding to the candidate role group to obtain the encoding results corresponding one-to-one to at least one candidate role name belonging to the candidate role group; wherein, the mask vector includes elements corresponding one-to-one to each character in the annotation text, the target dialogue, and the candidate role names in the candidate role list, and each element is used to indicate whether the corresponding character or the corresponding target dialogue or the corresponding candidate role name belongs to the candidate role group corresponding to the mask vector.

[0009] According to another aspect of the present invention, there is provided a speech synthesis method, including: obtaining the text to be synthesized; extracting the annotation text, the target dialogue, and the candidate role list from the text to be synthesized; using the annotation text, the target dialogue, and the candidate role list to perform role annotation through the above-mentioned role annotation method to obtain the role annotation result of the target dialogue; obtaining the acoustic model corresponding to the role in the role annotation result; using the acoustic model to perform speech synthesis on the target dialogue to obtain the target speech.

[0010] According to another aspect of the present invention, there is also provided a role annotation device, including: an acquisition module, configured to acquire an annotation text, a target dialogue, and a candidate role list, where the annotation text includes the target dialogue and the context of the target dialogue, the candidate role list includes one or more candidate role names appearing in the annotation text, and the one or more candidate role names are divided into at least one candidate role group; a first mapping module, configured to map the annotation text character by character into feature vectors one by one to obtain a text vector sequence; a second mapping module, configured to map the target dialogue character by character into feature vectors one by one to obtain a dialogue vector sequence; a first pooling module, configured to perform pooling on the vectors in the dialogue vector sequence to obtain a dialogue vector; a third mapping module, configured to, for each candidate role name in the candidate role list, map the candidate role name character by character into feature vectors one by one to obtain a name vector sequence corresponding to the candidate role name; a second pooling module, configured to, for each candidate role name in the candidate role list, perform pooling on the vectors in the name vector sequence to obtain a role vector corresponding to the candidate role name; an input module, configured to input the text vector sequence, the dialogue vector, and the role vector sequence into an encoder module to obtain an encoded result sequence, where the encoded result sequence includes encoded results corresponding one by one to the vectors in the text vector sequence, the dialogue vector, and the role vector sequence, the encoded result of any vector includes the correlation information between the vector and other vectors, and the role vector sequence includes role vectors corresponding one by one to all candidate role names in the candidate role list; a classification module, configured to, for each candidate role group in the at least one candidate role group, classify the encoded results corresponding one by one to at least one candidate role name belonging to the candidate role group through the classification module to obtain a classification result of the candidate role group, where the classification result of any role is used to indicate whether the speaker of the target dialogue belongs to the candidate role group; a determination module, configured to determine the role annotation result of the target dialogue based on the classification results of the at least one candidate role group.

[0011] According to another aspect of the present invention, there is also provided a speech synthesis device, including: a first acquisition module, configured to acquire a text to be synthesized; an extraction module, configured to extract an annotation text, a target dialogue, and a candidate role list from the text to be synthesized; an annotation module, configured to use the annotation text, the target dialogue, and the candidate role list to perform role annotation through the above-mentioned role annotation method to obtain a role annotation result of the target dialogue; a second acquisition module, configured to acquire an acoustic model corresponding to the role in the role annotation result; a synthesis module, configured to use the acoustic model to perform speech synthesis on the target dialogue to obtain a target speech.

[0012] According to another aspect of the present invention, there is provided a role annotation system, including a processor and a memory, wherein computer program instructions are stored in the memory and are used to execute the above-mentioned role annotation method when run by the processor.

[0013] According to another aspect of the present invention, there is provided a speech synthesis system, including a processor and a memory, wherein computer program instructions are stored in the memory and are used to execute the above-mentioned speech synthesis method when run by the processor.

[0014] According to another aspect of the present invention, there is provided a storage medium on which program instructions are stored, and the program instructions are used to execute the above-mentioned role annotation method when running.

[0015] According to another aspect of the present invention, there is also provided a storage medium on which program instructions are stored, and the program instructions are used to execute the above-mentioned speech synthesis method when running.

[0016] For the role annotation method, device, system and storage medium, and the speech synthesis method, device, system and storage medium according to the embodiments of the present invention, the obtained annotation text, target dialogue and candidate role list are used as input information and input into the model at one time. The candidate roles in the annotation text, target dialogue and candidate role list are mapped and converted into feature vectors. By encoding each feature vector, the interaction relationships between each candidate role and between each candidate role and the annotation text and target dialogue are fully integrated. Further, the obtained encoded result is classified to obtain the classification result of each candidate role group. Finally, the role annotation result of the target dialogue is determined based on the classification result of the candidate role group. Different from the prior art solution that needs to score each candidate role in turn, the above technical solution can input the candidate role list at one time, can obtain the classification results of multiple candidate roles at one time and determine the final annotation result, which can not only save computing resources, but also has higher annotation efficiency and better user experience. In addition, different from the focus of the existing model, the present invention regards the context as a character sequence, while regarding the target dialogue and candidate roles as a whole, and performs a pooling operation on the basis of obtaining the character sequences for the two, so as to facilitate representing the relationships between each candidate role and the target dialogue and context. And, the context, target dialogue and candidate roles can be fully interacted, and the relationships between each candidate role are also fully considered. Therefore, the obtained annotation result is more accurate, reasonable, and the model adaptability is also stronger, thus greatly improving the user experience.

[0017] A series of simplified concepts are introduced in the summary of the invention, which will be further described in detail in the detailed implementation part. The summary of the invention does not mean to attempt to define the key features and essential technical features of the claimed technical solution, nor does it mean to attempt to determine the protection scope of the claimed technical solution.

[0018] The advantages and features of the present invention will be described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The following drawings of the present invention are hereby incorporated as part of the present invention for understanding the present invention. The embodiments of the present invention shown in the drawings and their descriptions are used to explain the principles of the present invention. In the drawings,

[0020] Figure 1 A schematic flowchart showing a role annotation method according to an embodiment of the present invention;

[0021] Figure 2 A schematic diagram showing a role annotation model according to an embodiment of the present invention;

[0022] Figure 3 A schematic flowchart showing a speech synthesis method according to an embodiment of the present invention;

[0023] Figure 4 A schematic block diagram showing a role annotation apparatus according to an embodiment of the present invention;

[0024] Figure 5 A schematic block diagram showing a speech synthesis apparatus according to an embodiment of the present invention;

[0025] Figure 6 A schematic block diagram showing a role annotation system according to an embodiment of the present invention;

[0026] Figure 7 A schematic block diagram showing a speech synthesis system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] In the following description, a large number of details are provided to enable a thorough understanding of the present invention. However, those skilled in the art can understand that the following description only exemplarily shows the preferred embodiments of the present invention, and the present invention can be implemented without one or more of such details. In addition, to avoid confusion with the present invention, some well-known technical features in the art are not described in detail.

[0028] To at least partially solve the above technical problems, an embodiment of the present invention provides a text role annotation method. In the present invention, the role annotation task is regarded as a sequence annotation task, and the context, the target dialogue, and the semantic relationship between each candidate role are fully utilized, and on this basis, whether each candidate role is the target role is automatically determined. Compared with the prior art, the above solution has higher accuracy and efficiency for the role annotation task, and stronger adaptability to various texts, thereby greatly improving the user experience of the natural language processing system.

[0029] Figure 1 A schematic flowchart showing a role annotation method 1000 according to an embodiment of the present invention is shown. As Figure 1 shown, the role annotation method 1000 includes step S1100, step S1200, step S1300, step S1400, step S1500, step S1600, step S1700, step S1800, and step S1900.

[0030] Step S1100: Obtain an annotated text, a target dialogue, and a list of candidate roles. Among them, the annotated text includes the target dialogue and the context of the target dialogue. The list of candidate roles includes one or more candidate role names that appear in the annotated text, and the one or more candidate role names are divided into at least one candidate role group.

[0031] It should be noted that duplicate occurrences of one or more candidate role names included in the list of candidate roles involved in step S1100 are allowed. For example, if the candidate role name "Zhang San" appears three times in the annotated text, it can also appear three times in the list of candidate roles.

[0032] Exemplarily, the role annotation method in the embodiments of the present invention can perform annotation based on existing text data. For example, the role annotation method 1000 may further include: obtaining an initial text. Step S1100 may include: extracting the target dialogue, the context of the target dialogue, and the list of candidate roles located in the initial text. Exemplarily, the initial text may be in electronic form, such as text data in txt format. The initial text may be all the text data of a novel to be annotated with roles, or it may be one chapter, one page, or one or more paragraphs of the text data therein. It can be understood that the initial text includes the target dialogue, the context, and all the candidate roles involved in the text. The candidate roles are, for example, all or some of the character roles involved in the novel. The target dialogue may be one or more sentences spoken by one or more of the character roles. The determination of the context may be based on the position of the target dialogue in the text, and the above and below texts of the target dialogue are extracted according to preset rules.

[0033] Taking the famous contemporary Chinese novel "Ordinary World" as an example, the initial text obtained can be the entire electronic text of the tenth chapter of the third part. Exemplarily, a character dialogue "I... am embarrassed" can be used as the target dialogue, and the context can be extracted according to the preset rules. Exemplarily and not restrictively, the target dialogue can be used as a benchmark to search for sentences with candidate character names in the previous and next texts in the order from back to front and from front to back, respectively, such as the first previous sentence of the target dialogue "Shao Ping did not understand what An Suozi meant". It can be understood that in the context of the target dialogue, there may also be sentences without candidate character names, such as "Why don't you come to the master's house?".

[0034] Step S1100 includes: according to a preset rule, extracting a preset number of sentences located before and after the target dialogue, and determining them as the context of the target dialogue. In the embodiment of the present invention, both the context and the context of the determined target dialogue contain at least one sentence with the candidate character name. That is, the number of sentences containing the candidate character name in the context / context can be one or more, and even each sentence in the context can contain the candidate character name. The following is the determined target dialogue "I... was too shy." A specific example of the context:

[0035] "Then why don't you come to your master's house?" Shao Ping didn't understand what An Suozi meant. "I... didn't have good intentions." An Suozi stammered. "I brought the flashlight to light the way for you, because it was dark and you might make a mistake..." Oh my God, that's how it was! Shao Ping really wanted to slap him in the face for his "Lei Feng spirit"!

[0036] In the above specific example of context, the number of sentences containing candidate character names in the target dialogue above is 1, and the number of sentences containing candidate character names in the target dialogue below is 2. In the embodiment of the present invention, the user can set a preset rule to determine the context according to actual needs, which is not limited here.

[0037] The annotated text obtained in step S1100 consists of the target dialogue and the context of the target dialogue extracted according to preset rules. The obtained candidate role list includes one or more candidate roles appearing in the annotated text. The candidate role can be a character portrayed in the novel appearing in the annotated text. The candidate role list can be a list of character names, for example, a list consisting of characters such as Shao Ping, Shao An, Run Ye, and Run Sheng appearing in the annotated text. It can be understood that when only one character is involved in the annotated text, the candidate role list only includes the character name of the character. When multiple characters are involved in the annotated text, the candidate role list consists of the character names of the multiple characters. That is, the candidate role list corresponds to the names of all candidate characters involved in the annotated text.

[0038] Exemplarily, each candidate role name in the candidate role list may appear multiple times. For example, in the specific example of the above context, there are two candidate role names, "Shaoping" and "An Suozi", and both of these two candidate role names appear twice in the context. Therefore, the candidate role list for this context can be expressed as "Shaoping, An Suozi, An Suozi, Shaoping".

[0039] Exemplarily, the candidate role list can be obtained by the user in advance, for example, by means of manual extraction. In another example, the candidate role list can also be obtained through a candidate role recognition and extraction module. This candidate role recognition and extraction module can sequentially recognize and extract the names of the candidate roles involved in the annotated text according to existing natural language processing techniques. Exemplarily, first, the annotated text can be segmented according to certain segmentation rules, and the person names can be annotated, for example, annotated as nr. Secondly, the text after segmentation and part-of-speech annotation is further extracted, and the words with the part-of-speech of nr are extracted as the names of the character roles, so as to construct the candidate role list. It should be particularly noted that any existing and future methods that can obtain the candidate role list in the annotated text can fall within the protection scope of the present invention and will not be limited here.

[0040] After obtaining the candidate role list in step S1100, one or more candidate role names can also be divided into at least one candidate role group. Exemplarily but not restrictively, in at least one candidate role group, the candidate role names that are exactly the same are grouped together; or, in at least one candidate role group, the candidate role names belonging to the same role are grouped together.

[0041] In the annotated text, the same candidate role name may appear multiple times. In one example, the candidate role names that are exactly the same can be grouped together. For example, the candidate role names with the candidate role name of "Shaoping" in the annotated text are divided into a candidate role group.

[0042] In addition, in a novel, according to the setting of the novel plot, each character role can have a formal name, or an alias or appellation name. Even in some novels, a specific character role may have multiple aliases or appellation names. In another example, the candidate role names belonging to the same role can be grouped together. For example, "Tian Er" and "Tian Fushun" belong to the same character role, and the candidate role names "Tian Er" and "Tian Fushun" can be divided into a candidate role group; "Shao'an" and "Sun Shao'an", "An'an" belong to the same character role and can be divided into another candidate role group; "Director Tian" and "Tian Fujun" also belong to the same character role and can be divided into yet another candidate role group.

[0043] According to the above technical solution, users can select an appropriate method for dividing candidate role groups according to actual needs to construct each candidate role group involved in the candidate annotation text, so as to quickly establish the interaction relationship between candidate role groups and improve the efficiency and accuracy of role annotation.

[0044] Step S1200: Map the annotation text character by character into feature vectors one by one to obtain a text vector sequence.

[0045] Exemplarily, according to the preset function f: X→Y, each character in the annotation text can be mapped into a multi-dimensional dense feature vector one by one. It can be understood that when the annotation text contains m characters, m multi-dimensional dense feature vectors can be obtained, and thus a multi-dimensional dense feature vector sequence, that is, a text vector sequence, can be formed. Exemplarily, step S1200 can also be implemented through, for example, the Embedding layer in any neural network model. Each feature vector in the above text vector sequence can be an embedding vector. Those skilled in the art can understand the extraction method and meaning of the Embedding vector, which will not be elaborated in this article.

[0046] Step S1300: Map the target dialogue character by character into feature vectors one by one to obtain a dialogue vector sequence.

[0047] Similar to step S1200, the same method can be used to obtain the feature vector sequence representing the target dialogue, that is, the target dialogue vector sequence, which will not be elaborated here.

[0048] Step S1400: Perform pooling on the vectors in the dialogue vector sequence to obtain a dialogue vector.

[0049] Exemplarily, a first pooling operation can be performed on the multi-dimensional dense vector representing the target dialogue obtained in step S1300. The method of the first pooling can be max pooling or average pooling, etc. Thus, the target dialogue is further transformed into a vector representation.

[0050] Step S1500: For each candidate role name in the candidate role list, map the candidate role name character by character into feature vectors one by one to obtain a name vector sequence corresponding to the candidate role name.

[0051] Step S1600: For each candidate role name in the candidate role list, perform pooling on the vectors in the name vector sequence to obtain a role vector corresponding to the candidate role name.

[0052] Step S1500 and step S1600 can respectively adopt methods similar to the above-mentioned step S1300 and step S1400 to map each candidate role name in the candidate role list at the character granularity to obtain their respective name vector sequences, and obtain the role vectors corresponding to each candidate role name after the second pooling operation. It can be understood that the number of role vectors obtained in step S1600 is the same as the number of candidate role names in the candidate role list. For example, according to the specific example of the candidate list in the foregoing step S1100, the 4 candidate role names therein can be respectively mapped at the character granularity to obtain multi-dimensional dense feature vector sequences, and the corresponding role vectors can be obtained respectively after the second pooling operation, and finally 4 role vectors representing the 4 candidate role names in the candidate role list are obtained.

[0053] Exemplarily, the method of the second pooling can be max pooling or average pooling, etc.

[0054] Exemplarily, step S1200 to step S1600 can be executed sequentially in order, or step S1200, step S1300 and step S1500 can be executed simultaneously first, and then step S1400 and step S1600 can be executed simultaneously. For example, the labeled text and the target dialogue obtained in step S1100 can be input into the embedding layer of the neural network model together with the candidate role list, and then the obtained dialogue vector sequence and multiple name vector sequences can be respectively input into the first pooling layer and multiple second pooling layers of the neural network model to respectively obtain the pooled dialogue vector and multiple role vectors. In the embodiments of the present invention, the execution order of step S1200 to step S1600 can include various suitable orders, which are not limited herein.

[0055] Step S1700, input the text vector sequence, the dialogue vector and the role vector sequence into the encoder module to obtain an encoded result sequence, where the encoded result sequence includes encoded results corresponding one by one to the vectors in the text vector sequence, the dialogue vector and the role vector sequence, and the encoded result of any vector includes the correlation information between this vector and other vectors, and the role vector sequence includes role vectors corresponding one by one to all candidate role names in the candidate role list.

[0056] In the above step S1600, the role vectors representing each candidate role name in the candidate role list can be obtained. When there are multiple role names, a role vector sequence can be obtained. This role vector sequence includes role vectors corresponding one by one to all candidate role names in the candidate role list.

[0057] Exemplarily, the sequence of role vectors obtained in the above steps, the sequence of text vectors obtained in steps S1200, S1400, and S1600, and the dialogue variables can be input into the encoder module at once, and the relationships between the vectors can be fully integrated in the encoder module. The sequence of encoding results obtained at the output layer of the encoder module includes encoding results corresponding one by one to the vectors in the sequence of text vectors, the dialogue vectors, and the sequence of role vectors.

[0058] Moreover, the encoding result of any vector includes the correlation information between this vector and other vectors. The finally obtained encoding results can not only represent the interaction relationships between each candidate role and the labeled text and the target dialogue respectively, but also represent the relationships between the names of the candidate roles in the candidate role list.

[0059] In one example, the encoder module may include a Transformer module or a Bidirectional Encoder Representations from Transformers (BERT) module.

[0060] In another example, the encoder module may further include one or more of a convolutional neural network, a recurrent neural network, and a long short-term memory network, or the encoder module may include one or more of a convolutional neural network, a recurrent neural network, and a long short-term memory network and an attention module.

[0061] Step S1800: For each candidate role group in at least one candidate role group, classify the encoding results corresponding one by one to at least one candidate role name belonging to this candidate role group through a classification module to obtain the classification result of this candidate role group, where the classification result of any role group is used to indicate whether the speaker of the target dialogue belongs to this candidate role group.

[0062] The role vectors obtained in the foregoing steps are in one-to-one correspondence with each candidate role name in the candidate role list. Therefore, through the encoder module in step S1700, encoding results corresponding one by one to each candidate role name in the candidate role list can be obtained. Exemplarily, the encoding results corresponding one by one to each candidate role name can be represented by vectors, so vectors of multiple encoding results for multiple candidate role names can be obtained.

[0063] Since in step S1100, on the basis of obtaining the candidate role list, one or more candidate role names can be divided into at least one candidate role group. Therefore, the encoding results corresponding to each candidate role name in the obtained candidate role list can be classified according to the foregoing division of the role groups to obtain the classification result for each candidate role group.

[0064] Exemplarily, the classification result can be used to indicate the category to which the current candidate role group belongs. In the embodiments of the present invention, the classification result for each candidate role group can be used to indicate whether the speaker of the target dialogue belongs to this candidate role group. According to the foregoing steps, since the methods for dividing candidate role groups are different, the corresponding classification results for each candidate role group can also be different. For candidate role groups divided in the manner that the candidate role names are exactly the same, the classification result correspondingly indicates that the speaker of the target dialogue belongs to / does not belong to this candidate role name; for candidate role groups divided according to the same role, the classification result correspondingly indicates that the speaker of the target dialogue belongs to / does not belong to this role.

[0065] Exemplarily, the classification result can be directly represented by "Yes" or "No", or can also be represented by "1" or "0". In another example, the classification result can also be represented by the probability that the speaker of the target dialogue belongs to the current candidate role group or the probability of not belonging to the current candidate role group.

[0066] Exemplarily, the classification module can be a supervised or semi-supervised Softmax classifier, logistic classifier, or non-linear kernel SVM classifier, etc., or can also be an unsupervised classification clustering operation, such as the Expectation-Maximization (EM) algorithm and the Fuzzy C-Means clustering algorithm, etc.

[0067] Step S1900, determine the role annotation result of the target dialogue based on the classification results of at least one candidate role group.

[0068] Since in the above step S1800, the classification results of each candidate role group in the candidate role list can be obtained through the classifier. Exemplarily, when the number of candidate role groups in the candidate role list is only one, the role annotation result of the target dialogue can be directly determined according to the classification result of this candidate role group. When the number of candidate role groups in the candidate role list is multiple, the classification results of multiple candidate role groups can be integrated to determine the role annotation result of the target dialogue.

[0069] According to the foregoing examples, due to different methods for dividing candidate role groups, different processing operations need to be performed on the classification results for different situations to determine the annotation results of the target dialogue. In one example, for candidate role groups where only those with exactly the same role names are grouped together, the judgment on whether the candidate role groups belong to the same role can be skipped first. Instead, the classification results of multiple candidate role groups can be directly compared, and a candidate role group can be determined according to the preset result determination rules, and the role represented by this candidate role group can be labeled as the speaker of the target dialogue. It is also possible to first further comprehensively calculate the classification results of multiple candidate role groups belonging to the same role and then compare them with the remaining candidate role groups, and determine the final annotation result of the target dialogue according to the preset result determination rules. For example, the classification result of the candidate role group "Shaoping" and the classification result of the candidate role group "Sun Shaoping" can be weighted and summed based on the preset weight setting rules to obtain the first classification result for this role, and then the annotation result of the target dialogue can be determined based on the first classification results of multiple roles. In another example, for the case of candidate role groups directly grouped by role, the annotation result of the target dialogue can be directly determined based on the classification results of the candidate role groups.

[0070] Exemplarily, step S1900 can be implemented by a result determination module. Exemplarily, the result determination module can include a classifier and / or a pooling layer.

[0071] The role annotation method 1000 according to an embodiment of the present invention can be executed in one or successively in multiple role annotation models. Exemplarily and non - restrictively, the role annotation model can include a feature mapping module (such as an Embedding module), an encoder module, a classification module, etc. When a role annotation model includes the above - mentioned modules, end - to - end text role annotation can be performed in this model. It can be understood that by simply inputting the annotated text, the target dialogue, and the candidate role list into the input layer of the given role annotation model, the role annotation result for this target dialogue can be obtained at the output layer.

[0072] According to the above technical solution, the obtained labeled text, target dialogue, and candidate character list can be input into the model at one time as input information. The candidate characters in the labeled text, target dialogue, and candidate character list are mapped and converted into feature vectors. By encoding each feature vector, the interaction relationships between the candidate characters and between the candidate characters and the labeled text and target dialogue are fully integrated. Further, the obtained encoded result is classified to obtain the classification result of each candidate character group. Finally, based on the classification result of the candidate character group, the role annotation result of the target dialogue is determined. Different from the prior art solution that needs to score each candidate character in turn, the above technical solution can input the candidate character list at one time, obtain the classification results of multiple candidate characters at one time, and determine the final annotation result, which can not only save computing resources, but also has higher annotation efficiency and better user experience. In addition, different from the focus of the existing model, the present invention regards the context as a character sequence, and regards the target dialogue and candidate characters as a whole respectively. Based on the obtained character sequences of the two, a pooling operation is performed, so as to facilitate the characterization of the relationship between each candidate character and the target dialogue and context. Moreover, the context, target dialogue, and candidate characters can be fully interacted, and the relationship between the candidate characters is also fully considered. Therefore, the obtained annotation result is more accurate, reasonable, and the model adaptability is also stronger, thus greatly improving the user experience.

[0073] Exemplarily, step S1800 further includes step S1810 and step S1820.

[0074] In step S1810, for each candidate character group in at least one candidate character group, pooling is performed on the encoding results corresponding one by one to at least one candidate character name belonging to the candidate character group to obtain a pooled encoding result.

[0075] Exemplarily, the candidate character names belonging to a candidate character group may appear multiple times in the labeled text. For example, the candidate character name "Shaoping" in the candidate character list appears 5 times. Therefore, through the encoder module, 5 encoding results for "Shaoping" can be obtained. Exemplarily, the encoding result can be represented by a vector. Exemplarily, all 5 vectors of the encoding result for "Shaoping" can be input into one or more pooling layers to perform a third pooling operation to obtain a pooled encoding result. Exemplarily but not restrictively, the pooled encoding result can be a lower-dimensional vector.

[0076] Exemplarily, the method of the third pooling may include max pooling and average pooling. It should be noted that the third pooling may be the same as or different from the foregoing first pooling and second pooling.

[0077] Step S1820: For each candidate role group in at least one candidate role group, input the pooled encoding result into the classification module to obtain the classification result of this candidate role group.

[0078] Exemplarily, the pooled encoding result vector obtained in the above example for the candidate role name "Shaoping" can be input into a classifier, and a classification result indicating that "Shaoping" is the speaker of the target dialogue can be obtained.

[0079] Exemplarily, this classifier can be a softmax classifier based on binary classification. The classification result output by the classifier can be represented by "0" and "1", where "0" represents that "Shaoping" is not the speaker of this target dialogue, and "1" represents that "Shaoping" is the speaker of this target dialogue.

[0080] According to the above technical solution, on the basis of pooling the encoding results corresponding to the candidate role names belonging to each candidate role group, the pooled encoding results are classified. This method is simple and easy to implement, has high efficiency, and the obtained results are more accurate.

[0081] Exemplarily, before step S1810, the role annotation method 1000 further includes step S1805: For each candidate role group in at least one candidate role group, multiply the encoding result sequence by the mask vector corresponding to this candidate role group to obtain encoding results corresponding one by one to at least one candidate role name belonging to this candidate role group. Among them, the mask vector includes elements corresponding one by one to each character in the annotation text, the target dialogue, and the candidate role names in the candidate role list, and each element is used to indicate whether the corresponding character or the corresponding target dialogue or the corresponding candidate role name belongs to the candidate role group corresponding to the mask vector.

[0082] Exemplarily, the mask vectors of each candidate role group can form a mask vector matrix (or mask matrix), and the encoding result sequence can be directly multiplied by this mask vector matrix to obtain a vector of the encoding results corresponding to this candidate role group.

[0083] Exemplarily, in the mask vector, the elements corresponding to the candidate role names belonging to this candidate role group can be represented by 1, and the elements corresponding to the candidate role names not belonging to this candidate role group can be represented by 0. For example, in a candidate role group, the mask vector of the candidate role name "Shaoping" includes elements corresponding one by one to each character in the annotation text, the target dialogue, and the candidate role names in the candidate role list, and the element of the candidate role name "Shaoping" can be represented by "1", and the remaining elements can all be represented by "0".

[0084] Through the above technical solution, the encoding results that do not belong to the candidate role group can be filtered out in the form of a mask vector, and the encoding results that belong to the candidate role group can be extracted, facilitating subsequent analysis for different candidate role groups.

[0085] According to the description in the foregoing step S1100, the initial text containing the target dialogue can be obtained first, and then the annotated text including the context and the target dialogue can be extracted from the initial text based on a preset rule. Exemplarily, the preset rule can be based on the number of sentences in the context of the preset target dialogue.

[0086] In one example, step S1100 may include: obtaining the initial text containing the target dialogue; in the initial text, determining a first number of preceding sentences located before the target dialogue and a second number of following sentences located after the target dialogue, where, when there are a first predetermined number of preceding sentences before the target dialogue, the first number of preceding sentences is the first predetermined number of preceding sentences, and when there are no first predetermined number of preceding sentences before the target dialogue, the first number of preceding sentences includes all preceding sentences before the target dialogue, and, when there are a second predetermined number of following sentences after the target dialogue, the second number of following sentences is the second predetermined number of following sentences, and when there are no second predetermined number of following sentences after the target dialogue, the second number of following sentences includes all following sentences after the target dialogue; finding a first position where the candidate role name first appears from the first number of preceding sentences; finding a second position where the candidate role name last appears from the second number of following sentences; determining the characters from the first position to the second position as the annotated text.

[0087] For example, in the initial text, taking the target dialogue as a reference, the first predetermined number of sentences counted forward are determined as the preceding sentences, and the second predetermined number of sentences counted backward are determined as the following sentences. The first predetermined number and the second predetermined number can be any suitable numbers, and the two can be the same or different. In a specific example, both the first predetermined number and the second predetermined number can be set to 5, that is, the 5 sentences closest to the target dialogue before and after are set as the context. The settings of the first predetermined number and the second predetermined number can be adaptively adjusted according to actual needs to ensure that at least one candidate role name is included in the selected context.

[0088] In addition, when the first preset number / second predetermined number of sentences cannot be found in the preceding / following text of the target dialogue in the initial text, the maximum number of sentences included in the corresponding position of the initial text can be set as the corresponding preceding / following text sentences.

[0089] After determining the corresponding preceding sentence and succeeding sentence, the characters from the position where the first candidate role name appears to the position where the last candidate role name appears can be determined as the annotated text in the order of appearance. It should be particularly noted that in this example, the candidate role name is any candidate role name, not a specific candidate role name.

[0090] According to the above technical solution, the annotated text can be directly determined based on the number of sentences in the context of the target dialogue. This solution can be implemented by loading a simple mathematical algorithm, which can save computing resources to a certain extent.

[0091] Exemplarily, the annotated text including the context and the target dialogue is extracted from the initial text based on a preset rule, and the preset rule can also be based on the number of times the candidate role name appears in the text.

[0092] In another example, step S1100 may further include: obtaining an initial text including the target dialogue; in the initial text, searching forward from the target dialogue for a third number of candidate role names, and taking the earliest position where the third number of candidate role names are found as the third position, where, when there are a third predetermined number of candidate role names before the target dialogue, the third number of candidate role names is the third predetermined number of candidate role names, and when there are no third predetermined number of candidate role names before the target dialogue, the third number of candidate role names includes all candidate role names before the target dialogue; searching backward from the target dialogue for a fourth number of candidate role names, and taking the last position where the fourth number of candidate role names are found as the fourth position, where, when there are a fourth predetermined number of candidate role names after the target dialogue, the fourth number of candidate role names is the fourth predetermined number of candidate role names, and when there are no fourth predetermined number of candidate role names after the target dialogue, the fourth number of candidate role names includes all candidate role names after the target dialogue; determining the characters from the third position to the fourth position as the annotated text.

[0093] For example, taking the position of the target dialogue in the initial text as a reference, the third number of candidate role names can be searched forward and the fourth number of candidate role names can be searched backward respectively, and the annotated text can be determined according to the positions of the first and last candidate role names. Similar to the previous example, when the initial text does not meet the above settings, adaptive adjustments can be made according to the actual situation. Those of ordinary skill in the art can easily understand the above solutions and will not be elaborated here.

[0094] The solution of determining the annotated text based on the number of candidate role names before and after the target dialogue relatively comprehensively considers more candidate roles, so the accuracy of role annotation is higher.

[0095] Exemplarily, the classification result of any candidate role group obtained through step S1800 includes the probability that the speaker of the target dialogue belongs to the candidate role group. Step S1900 includes step S1910 and step S1920. In step S1910, the candidate role group with the highest probability is selected from at least one candidate role group. In step S1920, it is determined that the role corresponding to the selected candidate role group is the role to which the speaker of the target dialogue belongs.

[0096] Exemplarily, in the embodiments of the present invention, the classification result for each candidate role group may include the probability that the speaker of the target dialogue belongs to / does not belong to the candidate role group. For example, for a specific example where the candidate role list is "Shaoping, An Suozi, An Suozi, Shaoping", two candidate role groups "Shaoping" and "An Suozi" can be obtained. By inputting the encoded result sequence obtained in step S1700 into the classifier, the probability P1 that the speaker of the target dialogue belongs to "Shaoping" and the probability P2 that the speaker of the target dialogue belongs to "An Suozi" can be obtained respectively, and the magnitudes of P1 and P2 are compared. When P1 > P2, it can be determined that "Shaoping" is the speaker of the target dialogue.

[0097] In addition, in an example where the classification result of each candidate role group includes the probability that the speaker of the target dialogue belongs to the candidate role group, the candidate role groups can also be initially screened according to a preset probability threshold, and then from the candidate role groups that meet the threshold, the candidate role group with the largest probability value is selected to determine the role annotation result of the target dialogue. The preset probability threshold is, for example, 0.6, and the probability that the speaker of the target dialogue belongs to the candidate role group is compared with the preset probability threshold of 0.6. For the case where the candidate role list includes only one candidate role group, it can be set that the role annotation result of the target dialogue is set as the role represented by the candidate role group only when the probability is greater than 0.6, otherwise it can be annotated as "to be determined"; for the case where the candidate role list includes multiple candidate role groups, the candidate role groups with probability values greater than 0.6 can be first screened out, and then the corresponding probability values of the screened multiple candidate role groups are sorted from largest to smallest, and finally the role annotation result of the target dialogue is determined as the role represented by the candidate role group with the largest probability value.

[0098] For the case of grouping according to candidate role names in the foregoing example, multiple candidate role groups belonging to the same role can be comprehensively calculated first to obtain the first probability that the speaker of the target dialogue belongs to this role. Then, the first probabilities corresponding to each role in the candidate role list are sorted, and finally, the role annotation result of the target dialogue is determined to be the role with the largest first probability value. For example, for the case where the candidate role list includes the candidate role group "Shaoping" and the candidate role group "Sun Shaoping", the probability that the speaker of the target dialogue belongs to "Shaoping" and the probability that the speaker of the target dialogue belongs to "Sun Shaoping" can be weighted and summed based on a preset weight setting rule to obtain the first probability that the speaker of the target dialogue belongs to the role Sun Shaoping. The preset weight rule can be, for example, based on the number of times each candidate role group appears in the candidate role list. The greater the number of times, the higher its weight; it can also be based on the comprehensive distance between each candidate role group and the target dialogue. The closer the distance, the higher the weight.

[0099] According to the above technical solution, the classification result is represented by the probability that the speaker of the target dialogue belongs to this candidate role group, and the candidate role group with the largest probability value is selected to determine the final role annotation result. This solution is more intuitive and has a high accuracy rate.

[0100] Exemplarily, before step S1700, the role annotation method 1000 further includes: obtaining initial position information, where the initial position information includes the start position information and end position information of each character in the annotation text, the target dialogue, and each candidate role name in the candidate role list; determining position encoding information based on the initial position information; step 1700 includes: inputting the text vector sequence, the dialogue vector, the role vector sequence, and the position encoding information into the encoder module together to obtain an encoded result sequence.

[0101] Exemplarily, for the feature vector representing each character in the annotation text, its corresponding start position information and end position information are the same; for the dialogue vector representing the target dialogue and the role vector representing each candidate role name, their corresponding start position information and end position information are usually different. However, when the candidate role name contains only a single character, the start position information and end position information corresponding to its role vector can also be the same.

[0102] Based on the start position information and end position information corresponding to the above vectors, the position encoding information can be determined. The position encoding information can be absolute position encoding information or relative position encoding information. Exemplarily but not restrictively, the start position information and end position information can be mapped to dense vectors based on a preset position encoding rule. The preset position encoding rule can include that positive integers greater than 0 can be used to represent position coordinates. For example, for the first character in the labeled text, both its start position coordinate and end position coordinate can be represented by 1, and for the nth character in the labeled text, its start position coordinate and end position coordinate can be represented by n accordingly; correspondingly, the start position coordinate of the target dialogue can be represented by m1, and the end position coordinate can be represented by m2, that is, the first character of the target dialogue can be the m1th character in the labeled text, and the last character of the target dialogue can be the m2th character in the labeled text; correspondingly, the start position coordinates and end position coordinates can be set respectively for encoding according to the position relationship of each candidate character name in the labeled text. Thus, the absolute position encoding information related to the absolute positions corresponding to the feature vectors, dialogue vectors, and character vectors in the text vector sequence can be obtained. Optionally, the text vector sequence, dialogue vector, character vector sequence, and the absolute position encoding information can be input into the encoder module for processing together to obtain an encoded result sequence.

[0103] Exemplarily, a relative position matrix between any two vectors in the text vector sequence, dialogue vector, and character vector sequence can also be calculated based on the preset position encoding rule. The relative position matrix includes a first relative distance matrix between the start positions of the two vectors, a second relative distance matrix between the end positions, a third relative distance matrix between the start position of the first vector and the end position of the second vector, and a fourth relative distance matrix between the end position of the first vector and the start position of the second vector.

[0104] Exemplarily, the four relative distance matrices of any two vectors can also be concatenated and transformed to obtain a position encoding vector between any two vectors, and thus a position vector sequence can be obtained. This position vector sequence is the above-mentioned relative position encoding information. Step 1700 can also include: inputting the text vector sequence, dialogue vector, character vector sequence, and position vector sequence into the encoder module together to obtain an encoded result sequence.

[0105] According to the above technical solution, position encoding can be performed based on the positions of each character in the annotated text and the positional relationships between the target dialogue and each candidate character name in the annotated text, and the encoded results are input into the encoder module for full integration to achieve the positional interaction between any two of the annotated text, the target dialogue, and each candidate character name. The character annotation scheme is more reasonable, which helps to further enhance the accuracy of the annotation results and the adaptability of the model.

[0106] Exemplarily, the character annotation method 1000 further includes steps S1001 to S1010. Step S1001, obtain a sample text, a sample dialogue, a sample character list, and label data, where the sample text includes the sample dialogue and the context of the sample dialogue, the sample character list includes one or more sample character names that appear in the sample text, the one or more sample character names are divided into at least one sample character group, and the label data includes the actual character annotation result of the sample dialogue; Step S1002, map the annotated text to feature vectors one by one at the granularity of a single character to obtain a sample text vector sequence; Step S1003, map the sample dialogue to feature vectors one by one at the granularity of a single character to obtain a sample dialogue vector sequence; Step S1004, perform pooling on the vectors in the sample dialogue vector sequence to obtain a sample dialogue vector; Step S1005, for each sample character name in the sample character list, map the sample character name to feature vectors one by one at the granularity of a single character to obtain a sample name vector sequence corresponding to the sample character name; Step S1006, perform pooling on the vectors in the sample name vector sequence to obtain a sample character vector corresponding to the sample character name; Step S1007, input the sample text vector sequence, the sample dialogue vector, and the sample character vector sequence into the encoder module to obtain a sample encoding result sequence, the sample encoding result sequence includes encoding results corresponding one by one to the vectors in the sample text vector sequence, the sample dialogue vector, and the sample character vector sequence, and the sample character vector sequence includes sample character vectors corresponding one by one to all sample character names in the sample character list; Step S1008, for each sample character group in at least one sample character group, classify the encoding results corresponding one by one to at least one sample character name belonging to the sample character group through a classification module to obtain a classification result of the sample character group; Step S1009, determine the sample character annotation result of the sample dialogue based on the classification results of at least one sample character group; Step S1010, calculate the loss value of the loss function based on the sample character annotation result and the actual character annotation result; Step S1011, optimize the encoder module and the classification module through the calculated loss value.

[0107] Exemplarily, the solutions of steps S1001 to S1009 can be understood by referring to the foregoing steps S1010 to S1090 one by one, and will not be elaborated here. It should be noted that, different from the foregoing steps, step S1001 can be regarded as obtaining sample data, and the sample data includes sample text, sample dialogue, sample role list, and label data, where the label data is used to indicate the real role that matches the sample dialogue.

[0108] Step S1010 can also compare the annotation result of the sample role obtained in step S1009 with the real result annotated in the label data, and calculate the loss value of the loss function. Exemplarily, the loss function can include mean square error loss function, mean absolute error loss function, cross entropy loss function, etc. Step S1011 can then optimize the encoder module and the classification module based on the obtained loss value. In another example, steps S1002 to S1006 can be executed in the feature extraction module, and step S1011 can also include: optimizing the feature extraction module through the loss value calculated in step S1010.

[0109] Exemplarily, the role annotation method 1000 can also include obtaining a sample data set. The sample data set includes at least one sample data group, where each sample data group includes corresponding sample text, sample dialogue, sample role list, and label data. Exemplarily, after the above step 1011, it further includes: for the next sample data group in the sample data set, executing the above steps S1001 to S1011. Exemplarily, the sample data set can include positive sample data groups and / or negative sample data groups. The feature extraction module, encoder module, and classification module can be optimized based on the positive sample data groups and negative sample data groups, thereby realizing the performance optimization of each of the modules, and further improving the accuracy of the role annotation method.

[0110] Exemplarily, the character annotation method 1000 may further include steps S1020 to S1029. Step S1020, obtain a sample text, sample dialogue, sample character list, and label data, where the sample text includes the sample dialogue and the context of the sample dialogue, the sample character list includes one or more sample character names appearing in the sample text, the one or more sample character names are divided into at least one sample character group, and the label data includes the actual character annotation result of the sample dialogue; Step S1021, map the annotation text to feature vectors one by one at the single-character granularity to obtain a sample text vector sequence; Step S1022, map the sample dialogue to feature vectors one by one at the single-character granularity to obtain a sample dialogue vector sequence; Step S1023, perform pooling on the vectors in the sample dialogue vector sequence to obtain a sample dialogue vector; Step S1024, for each sample character name in the sample character list, map the sample character name to feature vectors one by one at the single-character granularity to obtain a sample name vector sequence corresponding to the sample character name; Step S1025, perform pooling on the vectors in the sample name vector sequence to obtain a sample character vector corresponding to the sample character name; Step S1026, input the sample text vector sequence, sample dialogue vector, and sample character vector sequence into an encoder module to obtain a sample encoding result sequence, the sample encoding result sequence includes encoding results corresponding one by one to the vectors in the sample text vector sequence, sample dialogue vector, and sample character vector sequence, and the sample character vector sequence includes sample character vectors corresponding one by one to all sample character names in the sample character list; Step S1027, for each sample character group in the at least one sample character group, classify the encoding results corresponding one by one to at least one sample character name belonging to the sample character group through a classification module to obtain a classification result of the sample character group; Step S1028, calculate the loss value of the loss function based on the classification results of the at least one sample character group and the actual character annotation result; Step S1029, optimize the encoder module and the classification module through the calculated loss value.

[0111] The technical solutions of steps S1020 to S1027 in this example are similar to those of steps S1001 to S1008 described above and can be understood by referring to the foregoing steps, which will not be elaborated here. It should be noted that the label data obtained in step S1020 in this example may include the setting of label values for each sample character group in the sample character list. For example, the label value of the sample character group belonging to the true character of the target dialogue may be set to 1, and vice versa to 0. Step S1028 may calculate the loss value of the loss function based on the comparison between the classification result of each sample character group and the corresponding value in the label data. Step S1029 may then optimize each module according to the calculated loss value.

[0112] According to the above technical solution, using the actual result as the target value, the classification results of each role group are supervised to optimize each module for role annotation, which can further improve the performance of each module and the accuracy of role annotation.

[0113] Exemplarily, according to the foregoing statement, the above role annotation method can be implemented by a role annotation model. Exemplarily and non - restrictively, referring to Figure 2 , Figure 2 FIG. shows a schematic diagram of a role annotation model 200 according to an embodiment of the present invention. Exemplarily, the processing flow of the role annotation model 200 shown in the figure is as follows.

[0114] At the input end of the model, an annotated text including the target dialogue "I... was too shy" and its context is input. Each character of the annotated text is mapped through an Embedding layer (abbreviated as "Emb" in the figure) to obtain a text vector sequence. At the same time, each character of the target dialogue is also mapped through the Embedding layer to obtain a dialogue vector sequence, and then a first pooling operation is performed on the dialogue vector sequence to obtain a dialogue vector. Similarly, each character of each candidate role name in the candidate role list is sent into the Embedding layer to obtain a corresponding name vector sequence, and then a second pooling operation is performed on each name vector sequence to obtain a corresponding role vector. In Figure 2 the candidate role list shown, there are 10 candidate role names, so 10 corresponding role vectors can be obtained. Then, based on the initial position information including the start position information and end position information of each character in the annotated text, the target dialogue, and each candidate role name in the candidate role list, position encoding information is determined. The position encoding information, together with the above - obtained text vector sequence, dialogue vector, and role vector sequence, is input into the encoder module to obtain an encoded result sequence. In Figure 2 the candidate role list shown, there are two candidate role groups, "Shaoping" and "An Suozi". The role annotation model 200 also obtains mask vectors corresponding to the candidate role groups "Shaoping" and "An Suozi" respectively. Then, the mask vector corresponding to each candidate role group is multiplied by the encoded result sequence, and the encoded results corresponding to 5 candidate role names for each of the candidate role groups "Shaoping" and "An Suozi" can be obtained respectively. The encoded results are respectively input into the classification module in the role annotation model 200, and a third pooling operation is performed on the encoded results of each candidate role group to obtain classification results corresponding one - to - one to the candidate role groups "Shaoping" and "An Suozi". Exemplarily, a binary classification of yes / No is performed on the classification results, and finally, the result of whether the target dialogue belongs to "Shaoping" or "An Suozi" can be obtained.

[0115] Thus, only by inputting the annotated text, the target dialogue, and the list of candidate character names into Figure 2 the character annotation model 200 shown in

[0116] According to another aspect of the present invention, there is also provided a speech synthesis method 300. Figure 3 The schematic flowchart of the speech synthesis method 300 is shown. As Figure 3 shown, the speech synthesis method 300 includes: step S310, obtaining the text to be synthesized; step S320, extracting the annotated text, the target dialogue, and the candidate character list from the text to be synthesized; step S330, using the annotated text, the target dialogue, and the candidate character list to perform character annotation through the above-mentioned character annotation method 1000 to obtain the character annotation result of the target dialogue; step S340, obtaining the acoustic model corresponding to the character in the character annotation result; step S350, using the acoustic model to perform speech synthesis on the target dialogue to obtain the target speech.

[0117] Exemplarily, the acoustic model can be a feature conversion model that generates acoustic features based on text features. Exemplarily, the speech synthesis method 300 further includes establishing an acoustic model library, which includes acoustic models corresponding to each character in the candidate character list. Exemplarily, each character corresponds to at least one acoustic model. Step S340 includes obtaining the acoustic model corresponding to the character in the character annotation result from the acoustic model library. For example, the acoustic model library includes acoustic models corresponding to the candidate character names "Shaoping", "An Suozi", and "Shao'an" in the foregoing example. When the character annotation result of the target dialogue obtained by step S330 through the foregoing character annotation method 1000 indicates that the speaker of the target dialogue is "Shaoping", the acoustic model corresponding to "Shaoping" is selected from the acoustic model library, and the target dialogue is synthesized using this acoustic model to obtain the target speech.

[0118] According to the acoustic synthesis method of the embodiments of the present invention, an acoustic model that is more matched to the character of the target dialogue can be obtained. Based on this, speech synthesis is performed on the target dialogue, and acoustic information that is more matched to the target dialogue can be generated, which can effectively improve the accuracy of the generated acoustic information and can significantly improve the user experience.

[0119] According to another aspect of the present invention, there is also provided a character annotation device 400, Figure 4 The schematic block diagram of the character annotation device 400 is shown. Exemplarily, the character annotation device 400 includes:

[0120] An acquisition module 410, configured to acquire an annotated text, a target dialogue, and a list of candidate roles, where the annotated text includes the target dialogue and the context of the target dialogue, the list of candidate roles includes one or more candidate role names appearing in the annotated text, and the one or more candidate role names are divided into at least one candidate role group;

[0121] A first mapping module 420, configured to map the annotated text into feature vectors one by one at the granularity of single characters to obtain a text vector sequence;

[0122] A second mapping module 430, configured to map the target dialogue into feature vectors one by one at the granularity of single characters to obtain a dialogue vector sequence;

[0123] A first pooling module 440, configured to perform pooling on the vectors in the dialogue vector sequence to obtain a dialogue vector;

[0124] A third mapping module 450, configured to, for each candidate role name in the list of candidate roles, map the candidate role name into feature vectors one by one at the granularity of single characters to obtain a name vector sequence corresponding to the candidate role name;

[0125] A second pooling module 460, configured to, for each candidate role name in the list of candidate roles, perform pooling on the vectors in the name vector sequence to obtain a role vector corresponding to the candidate role name;

[0126] An input module 470, configured to input the text vector sequence, the dialogue vector, and the role vector sequence into an encoder module to obtain an encoded result sequence, where the encoded result sequence includes encoded results corresponding one by one to the vectors in the text vector sequence, the dialogue vector, and the role vector sequence, the encoded result of any vector includes the correlation information between the vector and other vectors, and the role vector sequence includes role vectors corresponding one by one to all candidate role names in the list of candidate roles;

[0127] A result classification module 480, configured to, for each candidate role group in at least one candidate role group, classify the encoded results corresponding one by one to at least one candidate role name belonging to the candidate role group through a classification module to obtain a classification result of the candidate role group, where the classification result of any role is used to indicate whether the speaker of the target dialogue belongs to the candidate role group;

[0128] A determination module 490, configured to determine a role annotation result of the target dialogue based on the classification results of at least one candidate role group.

[0129] According to another aspect of the present invention, there is also provided a speech synthesis device 500, Figure 5The schematic block diagram of the speech synthesis device 500 is shown. Exemplarily, the speech synthesis device 500 includes:

[0130] A first acquisition module 510, configured to acquire the text to be synthesized;

[0131] An extraction module 520, configured to extract the annotation text, the target dialogue, and the candidate role list from the text to be synthesized;

[0132] An annotation module 530, configured to perform role annotation on the target dialogue by using the annotation text, the target dialogue, and the candidate role list through the above-mentioned role annotation method 1000 to obtain the role annotation result of the target dialogue;

[0133] A second acquisition module 540, configured to acquire the acoustic model corresponding to the role in the role annotation result;

[0134] A synthesis module 550, configured to perform speech synthesis on the target dialogue by using the acoustic model to obtain the target speech.

[0135] According to another aspect of the present invention, a role annotation system 600 is further provided. Figure 6 The schematic block diagram of the role annotation system 600 is shown. The role annotation system 600 includes a processor 610 and a memory 620. Among them, computer program instructions are stored in the memory 620, and when the computer program instructions are run by the processor 610, they are used to execute the above-mentioned role annotation method 1000.

[0136] According to another aspect of the present invention, a speech synthesis system 700 is further provided. Figure 7 The schematic block diagram of the speech synthesis system 700 is shown. The speech synthesis system 700 includes a processor 710 and a memory 720. Among them, computer program instructions are stored in the memory 720, and when the computer program instructions are run by the processor 710, they are used to execute the above-mentioned speech synthesis method 300.

[0137] According to another aspect of the present invention, a storage medium is further provided, on which program instructions are stored, and when the program instructions are run, they are used to execute the above-mentioned role annotation method 1000.

[0138] According to another aspect of the present invention, a storage medium is further provided, on which program instructions are stored, and when the program instructions are run, they are used to execute the above-mentioned speech synthesis method 300.

[0139] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0140] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed.

[0141] Similarly, it should be understood that, in order to streamline the present invention and help understand one or more of the various inventive aspects, in the description of the exemplary embodiments of the present invention, the various features of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. However, the method of the present invention should not be construed as reflecting the intention that the claimed invention requires more features than those expressly recited in each claim. More precisely, as reflected by the corresponding claims, the inventive point lies in that the corresponding technical problems can be solved with features less than all the features of a single disclosed embodiment. Therefore, the claims following the specific implementation manner are hereby expressly incorporated into the specific implementation manner, where each claim itself serves as a separate embodiment of the present invention.

[0142] Those skilled in the art can understand that, except for features that are mutually exclusive, any combination can be used to combine all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature that provides the same, equivalent, or similar purpose.

[0143] Each component embodiment of the present invention can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some modules in the acoustic model training system or the speech synthesis system according to the embodiments of the present invention. The present invention can also be implemented as a device program (for example, a computer program and a computer program product) for executing part or all of the methods described herein. Such a program for implementing the present invention can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0144] It should be noted that the above embodiments illustrate the present invention rather than limit the present invention, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In the unit claims listing several devices, several of these devices can be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names.

[0145] As described above, it is only the specific implementation manner of the present invention or the description of the specific implementation manner, and the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should be covered by the protection scope of the present invention. The protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A role annotation method, comprising: Obtaining an annotation text, a target dialogue, and a candidate role list, wherein the annotation text includes the target dialogue and the context of the target dialogue, the candidate role list includes one or more candidate role names that appear in the annotation text, and the one or more candidate role names are divided into at least one candidate role group; Mapping the annotation text character by character to obtain a text vector sequence; Mapping the target dialogue character by character to obtain a dialogue vector sequence; Pooling the vectors in the dialogue vector sequence to obtain a dialogue vector; For each candidate role name in the candidate role list, Mapping the candidate role name character by character to obtain a name vector sequence corresponding to the candidate role name; Pooling the vectors in the name vector sequence to obtain a role vector corresponding to the candidate role name; Inputting the text vector sequence, the dialogue vector, and the role vector sequence into an encoder module to obtain an encoded result sequence, wherein the encoded result sequence includes encoded results corresponding one by one to the vectors in the text vector sequence, the dialogue vector, and the role vector sequence, and the encoded result of any vector includes the correlation information between this vector and other vectors, and the role vector sequence includes role vectors corresponding one by one to all candidate role names in the candidate role list; For each candidate role group in the at least one candidate role group, classifying the encoded results corresponding one by one to at least one candidate role name belonging to this candidate role group through a classification module to obtain the classification result of this candidate role group, wherein the classification result of any candidate role group is used to indicate whether the speaker of the target dialogue belongs to this candidate role group; Determining the role annotation result of the target dialogue based on the classification results of the at least one candidate role group.

2. The method according to claim 1, wherein The step of, for each candidate role group in the at least one candidate role group, classifying the encoded results corresponding one by one to at least one candidate role name belonging to this candidate role group through a classification module to obtain the classification result of this candidate role group includes: For each candidate role group in the at least one candidate role group, Pooling the encoded results corresponding one by one to at least one candidate role name belonging to this candidate role group to obtain a pooled encoded result; Inputting the pooled encoded result into the classification module to obtain the classification result of this candidate role group.

3. The method according to claim 2, wherein, Before, for each candidate role group in the at least one candidate role group, pooling the encoded results corresponding one by one to at least one candidate role name belonging to this candidate role group to obtain a pooled encoded result, the method further includes: For each candidate role group in the at least one candidate role group, multiplying the encoded result sequence by the mask vector corresponding to this candidate role group to obtain the encoded results corresponding one by one to at least one candidate role name belonging to this candidate role group; Among them, the mask vector includes elements corresponding one-to-one to each character in the labeled text, the target dialogue, and the candidate character names in the candidate character list, and each element is used to indicate whether the corresponding character, the corresponding target dialogue, or the corresponding candidate character name belongs to the candidate character group corresponding to the mask vector.

4. The method according to any one of claims 1 to 3, wherein The pooling is average pooling or max pooling.

5. The method according to any one of claims 1 to 4, wherein, The obtaining of the labeled text, the target dialogue, and the candidate character list includes: Obtaining an initial text containing the target dialogue; In the initial text, Determining a first number of preceding sentences before the target dialogue and a second number of following sentences after the target dialogue, where, when there are a first predetermined number of preceding sentences before the target dialogue, the first number of preceding sentences is the first predetermined number of preceding sentences, and when there are not the first predetermined number of preceding sentences before the target dialogue, the first number of preceding sentences includes all preceding sentences before the target dialogue, and, when there are a second predetermined number of following sentences after the target dialogue, the second number of following sentences is the second predetermined number of following sentences, and when there are not the second predetermined number of following sentences after the target dialogue, the second number of following sentences includes all following sentences after the target dialogue; Finding a first position where a candidate character name first appears from the first number of preceding sentences; Finding a second position where a candidate character name last appears from the second number of following sentences; Determining the characters from the first position to the second position as the labeled text.

6. The method according to any one of claims 1 to 4, wherein, The obtaining of the labeled text, the target dialogue, and the candidate character list includes: Obtaining an initial text containing the target dialogue; In the initial text, Searching forward from the target dialogue for a third number of candidate character names, and taking the earliest position where the third number of candidate character names appears as a third position, where, when there are a third predetermined number of candidate character names before the target dialogue, the third number of candidate character names is the third predetermined number of candidate character names, and when there are not the third predetermined number of candidate character names before the target dialogue, the third number of candidate character names includes all candidate character names before the target dialogue; Searching backward from the target dialogue for a fourth number of candidate character names, and taking the last position where the fourth number of candidate character names appears as a fourth position, where, when there are a fourth predetermined number of candidate character names after the target dialogue, the fourth number of candidate character names is the fourth predetermined number of candidate character names, and when there are not the fourth predetermined number of candidate character names after the target dialogue, the fourth number of candidate character names includes all candidate character names after the target dialogue; Determining the characters from the third position to the fourth position as the labeled text.

7. The method according to any one of claims 1 to 4, wherein The classification result of any candidate role group includes the probability that the speaker of the target dialogue belongs to this candidate role group. The determination of the role annotation result of the target dialogue based on the classification results of the at least one candidate role group includes: Selecting the candidate role group with the highest probability from the at least one candidate role group; Determining that the role corresponding to the selected candidate role group is the role to which the speaker of the target dialogue belongs.

8. The method according to any one of claims 1 to 4, wherein, Before inputting the text vector sequence, the dialogue vector, and the role vector sequence into the encoder module to obtain an encoded result sequence, the method further includes: Obtaining initial position information, where the initial position information includes the start position information and end position information of each character in the annotated text, the target dialogue, and each candidate role name in the candidate role list; Determining position encoding information based on the initial position information; The inputting of the text vector sequence, the dialogue vector, and the role vector sequence into the encoder module to obtain an encoded result sequence includes: Inputting the text vector sequence, the dialogue vector, the role vector sequence, and the position encoding information into the encoder module together to obtain an encoded result sequence.

9. The method according to any one of claims 1 to 4, wherein In the at least one candidate role group, candidate role names that are exactly the same are grouped together; or, in the at least one candidate role group, candidate role names belonging to the same role are grouped together.

10. The method according to any one of claims 1 to 4, wherein, The method further includes: Obtaining a sample text, a sample dialogue, a sample role list, and label data, where the sample text includes the sample dialogue and the context of the sample dialogue, the sample role list includes one or more sample role names that appear in the sample text, the one or more sample role names are divided into at least one sample role group, and the label data includes the actual role annotation result of the sample dialogue; Mapping the annotated text character by character to feature vectors one by one to obtain a sample text vector sequence; Mapping the sample dialogue character by character to feature vectors one by one to obtain a sample dialogue vector sequence; Performing pooling on the vectors in the sample dialogue vector sequence to obtain a sample dialogue vector; For each sample role name in the sample role list, Mapping the sample role name character by character to feature vectors one by one to obtain a sample name vector sequence corresponding to the sample role name; Performing pooling on the vectors in the sample name vector sequence to obtain a sample role vector corresponding to the sample role name; Inputting the sample text vector sequence, the sample dialogue vector, and the sample role vector sequence into the encoder module to obtain a sample encoded result sequence, where the sample encoded result sequence includes encoded results corresponding one by one to the vectors in the sample text vector sequence, the sample dialogue vector, and the sample role vector sequence, and the sample role vector sequence includes sample role vectors corresponding one by one to all sample role names in the sample role list; For each of the at least one sample role group, classify the encoding results corresponding one-to-one to at least one sample role name belonging to this sample role group through the classification module to obtain the classification result of this sample role group; Determine the sample role annotation result of the sample dialogue based on the classification results of the at least one sample role group; Calculate the loss value of the loss function based on the sample role annotation result and the actual role annotation result; Optimize the encoder module and the classification module through the calculated loss value.

11. The method according to any one of claims 1 to 4, wherein, The method further includes: Obtain sample text, sample dialogue, a sample role list, and label data, where the sample text includes the sample dialogue and the context of the sample dialogue, the sample role list includes one or more sample role names appearing in the sample text, the one or more sample role names are divided into at least one sample role group, and the label data includes the actual role annotation result of the sample dialogue; Map the annotation text to feature vectors one by one at the granularity of a single character to obtain a sample text vector sequence; Map the sample dialogue to feature vectors one by one at the granularity of a single character to obtain a sample dialogue vector sequence; Perform pooling on the vectors in the sample dialogue vector sequence to obtain a sample dialogue vector; For each sample role name in the sample role list, Map this sample role name to feature vectors one by one at the granularity of a single character to obtain a sample name vector sequence corresponding to this sample role name; Perform pooling on the vectors in the sample name vector sequence to obtain a sample role vector corresponding to this sample role name; Input the sample text vector sequence, the sample dialogue vector, and the sample role vector sequence into the encoder module to obtain a sample encoding result sequence, the sample encoding result sequence includes encoding results corresponding one-to-one to the vectors in the sample text vector sequence, the sample dialogue vector, and the sample role vector sequence, and the sample role vector sequence includes sample role vectors corresponding one-to-one to all sample role names in the sample role list; For each of the at least one sample role group, classify the encoding results corresponding one-to-one to at least one sample role name belonging to this sample role group through the classification module to obtain the classification result of this sample role group; Calculate the loss value of the loss function based on the classification results of the at least one sample role group and the actual role annotation result; Optimize the encoder module and the classification module through the calculated loss value.

12. A speech synthesis method, including: Obtain the text to be synthesized; Extract annotation text, target dialogue, and a candidate role list from the text to be synthesized; Use the annotation text, the target dialogue, and the candidate role list to perform role annotation through the role annotation method according to any one of claims 1 to 11 to obtain the role annotation result of the target dialogue; Obtain an acoustic model corresponding to the role in the role annotation result; Use the acoustic model to perform speech synthesis on the target dialogue to obtain target speech.

13. A role annotation device, comprising: An acquisition module, configured to acquire an annotation text, a target dialogue, and a candidate role list, where the annotation text includes the target dialogue and the context of the target dialogue, the candidate role list includes one or more candidate role names that appear in the annotation text, and the one or more candidate role names are divided into at least one candidate role group; A first mapping module, configured to map the annotation text character by character to obtain a sequence of text vectors; A second mapping module, configured to map the target dialogue character by character to obtain a sequence of dialogue vectors; A first pooling module, configured to perform pooling on the vectors in the sequence of dialogue vectors to obtain a dialogue vector; A third mapping module, configured to, for each candidate role name in the candidate role list, map the candidate role name character by character to obtain a sequence of name vectors corresponding to the candidate role name; A second pooling module, configured to, for each candidate role name in the candidate role list, perform pooling on the vectors in the sequence of name vectors to obtain a role vector corresponding to the candidate role name; An input module, configured to input the sequence of text vectors, the dialogue vector, and the sequence of role vectors into an encoder module to obtain a sequence of encoding results, where the sequence of encoding results includes encoding results corresponding one by one to the vectors in the sequence of text vectors, the dialogue vector, and the sequence of role vectors, and the encoding result of any vector includes the correlation information between the vector and other vectors, and the sequence of role vectors includes role vectors corresponding one by one to all candidate role names in the candidate role list; A result classification module, configured to, for each candidate role group in the at least one candidate role group, classify the encoding results corresponding one by one to at least one candidate role name belonging to the candidate role group through a classification module to obtain a classification result of the candidate role group, where the classification result of any role is used to indicate whether the speaker of the target dialogue belongs to the candidate role group; A determination module, configured to determine the role annotation result of the target dialogue based on the classification results of the at least one candidate role group.

14. A speech synthesis device, comprising: A first acquisition module, configured to acquire text to be synthesized; An extraction module, configured to extract an annotation text, a target dialogue, and a candidate role list from the text to be synthesized; An annotation module, configured to perform role annotation on the target dialogue by using the annotation text, the target dialogue, and the candidate role list through the role annotation method according to any one of claims 1 to 11 to obtain a role annotation result of the target dialogue; A second acquisition module, configured to acquire an acoustic model corresponding to the role in the role annotation result; A synthesis module, configured to use the acoustic model to perform speech synthesis on the target dialogue to obtain target speech.

15. A role annotation system, including a processor and a memory, wherein, Computer program instructions are stored in the memory, and when the computer program instructions are run by the processor, they are used to execute the role annotation method according to any one of claims 1 to 11.

16. A speech synthesis system, comprising a processor and a memory, wherein, Computer program instructions are stored in the memory, and when the computer program instructions are run by the processor, they are used to execute the speech synthesis method according to claim 12.

17. A storage medium, on which program instructions are stored, and when the program instructions are run, they are used to execute the role annotation method according to any one of claims 1 to 11.

18. A storage medium, on which program instructions are stored, and when the program instructions are run, they are used to execute the speech synthesis method according to claim 12.

Citation Information

Patent Citations

  • Data processing method and device

    CN106155998A

  • Role labeling method and device, electronic equipment and storage medium

    CN112270167A