Method and apparatus for identifying a speaker of text and training a speaker recognition model
By identifying the speaker-related sentences of dialogue and adjacent sentences in the target text, and using the MRC model for identification, the problem of inaccurate speaker recognition caused by insufficient sample articles is solved, and the recognition accuracy of the model is improved.
Patent Information
- Application Number
- CN202210865694.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-21
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-07-21
AI Technical Summary
The prior art cannot guarantee the accuracy of the predictive results of the speaker's identification model when the number of sample articles is small.
By obtaining the dialogue sentences and their adjacent sentences in the target text for speaker name recognition, identifying the speaker-related sentences, and inputting them into the trained speaker recognition model, and using the machine reading comprehension MRC model for identification.
This improves the recognition accuracy of the speaker recognition model in the case of small number of sample articles and reduces the need for training samples.
Smart Images

Figure CN115312066B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a method and device for identifying speakers in text and training a speaker recognition model. Background Art
[0002] With the continuous development of speech synthesis technology, the reading of audiobooks has widely emerged. When reading an audiobook, adopting a multi-speaker dubbing reading mode can greatly improve the user experience. To implement the multi-speaker dubbing reading mode, it is necessary to determine the speakers of the dialogues in the text. Currently, most often through machine learning, the complete article text is input into the model to obtain the speaker corresponding to each dialogue sentence therein.
[0003] However, if accurate prediction results of the model are desired, a large number of complete sample articles are required to train the model. When the number of sample articles is small, the accuracy of the prediction results of the trained model cannot be guaranteed. Summary of the Invention
[0004] Embodiments of this application provide a method and device for identifying speakers in text and training a speaker recognition model to solve the problems of related technologies. The technical solutions are as follows:
[0005] In a first aspect, a method for identifying a speaker in text is provided. The method includes:
[0006] Obtain a target dialogue sentence in the target text, where the target dialogue sentence is any dialogue sentence in the target text;
[0007] Based on the speaker list corresponding to the target text, perform speaker name recognition on the adjacent sentences of the target dialogue sentence, where the speaker list includes at least one speaker name;
[0008] Based on the speaker name recognition result of the adjacent sentences, determine the speaker-related sentence of the target dialogue sentence in the adjacent sentences;
[0009] If the speaker-related sentence of the target dialogue sentence is determined, input the target dialogue sentence and the speaker-related sentence into a trained speaker recognition model to obtain the speaker information of the target dialogue sentence, where the speaker information is the speaker name of the target dialogue sentence or an indication information for indicating that there is no corresponding speaker name for the target dialogue sentence.
[0010] In a possible implementation, before obtaining the target dialogue sentence in the target text, the method further includes:
[0011] Based on the end punctuation of the sentence, perform sentence splitting on the target text to obtain multiple sentences;
[0012] For a sentence that contains quotation marks among the multiple sentences, the content inside the quotation marks and the content outside the quotation marks are divided into different sentences;
[0013] Among all the sentences obtained by clause-splitting the target text, the content inside the quotation marks is determined as the dialogue sentence in the target text.
[0014] In a possible implementation manner, the obtaining of the target dialogue sentence in the target text includes:
[0015] Among all the dialogue sentences in the target text, the target dialogue sentence is obtained in the order from the front to the back according to the position in the target text.
[0016] In a possible implementation manner, the identifying of the speaker name for the adjacent sentence of the target dialogue sentence based on the speaker list corresponding to the target text includes:
[0017] In the adjacent sentences of the target dialogue sentence, the speaker name in the speaker list corresponding to the target text is searched for, and the adjacent sentence containing the speaker name is determined as the adjacent sentence of the target dialogue sentence.
[0018] In a possible implementation manner, the determining of the speaker-related sentence of the target dialogue sentence in the adjacent sentence based on the identification result of the speaker name of the adjacent sentence includes:
[0019] If there is only one adjacent sentence among the adjacent sentences of the target dialogue sentence that contains the speaker name, and the adjacent sentence containing the speaker name has not been determined as the speaker-related sentence corresponding to other dialogue sentences outside the target dialogue sentence, then the adjacent sentence containing the speaker name is determined as the speaker-related sentence of the target dialogue sentence;
[0020] If both adjacent sentences of the target dialogue sentence contain the speaker name, and the adjacent sentence in front of the target dialogue sentence has not been determined as the speaker-related sentence corresponding to other dialogue sentences, then the adjacent sentence in front is determined as the speaker-related sentence of the target dialogue sentence;
[0021] If both adjacent sentences of the target dialogue sentence contain the speaker name, and the adjacent sentence in front of the target dialogue sentence has been determined as the speaker-related sentence corresponding to other dialogue sentences, then the adjacent sentence behind the target dialogue sentence is determined as the speaker-related sentence of the target dialogue sentence;
[0022] If there is no adjacent sentence containing the speaker name for the target dialogue sentence, or the adjacent sentence containing the speaker name of the target dialogue sentence has been determined as the speaker-related sentence corresponding to other dialogue sentences, then it is determined that the target dialogue sentence has no speaker-related sentence.
[0023] In a possible implementation, after determining the speaker-related sentence of the target dialogue sentence in the adjacent sentence based on the speaker name recognition result of the adjacent sentence, the method further includes:
[0024] If the speaker-related sentence of the target dialogue sentence is not determined, it is determined that the target dialogue sentence has no corresponding speaker name.
[0025] In a possible implementation, the speaker recognition model is a machine reading comprehension MRC model;
[0026] Inputting the target dialogue sentence and the speaker-related sentence into the trained speaker recognition model to obtain the speaker information of the target dialogue sentence includes:
[0027] Combining the target dialogue sentence and a preset question sentence to form a question field, and combining the target dialogue sentence and the speaker-related sentence to form a question-related text field, where the question sentence is used to query the speaker name of the target dialogue sentence;
[0028] Inputting the question field and the question-related text field into the MRC model to obtain the speaker information of the target dialogue sentence.
[0029] In a second aspect, a method for training a speaker recognition model is provided, and the method includes:
[0030] Obtaining sample dialogue sentences in the sample text;
[0031] Based on the sample speaker list corresponding to the sample text, performing speaker name recognition on the adjacent sentences of the sample dialogue sentences, where the sample speaker list includes at least one speaker name;
[0032] Based on the speaker name recognition result of the adjacent sentence, determining the sample speaker-related sentence corresponding to the sample dialogue sentence in the adjacent sentence of the sample dialogue sentence;
[0033] Using the sample dialogue sentence and the sample speaker-related sentence corresponding to the sample dialogue sentence as input sample data to train and adjust the parameters of the speaker recognition model to be trained, and obtaining the speaker recognition model with adjusted parameters;
[0034] If the training and parameter adjustment meet the preset end condition, determining the speaker recognition model with adjusted parameters as the trained speaker recognition model.
[0035] In a third aspect, a device for recognizing a speaker from text is provided, and the device includes:
[0036] An acquisition module, configured to acquire a target dialogue sentence in a target text;
[0037] a recognition module, configured to perform speaker name recognition on adjacent sentences of the target dialogue sentence based on a speaker list corresponding to the target text, wherein the speaker list includes at least one speaker name;
[0038] A determination module, configured to determine a speaker-related sentence of the target dialogue sentence in the adjacent sentences based on a speaker name recognition result of the adjacent sentences;
[0039] The output module is used for inputting the target dialogue sentence and the speaker-related sentence into a trained speaker recognition model if the speaker-related sentence of the target dialogue sentence is determined, so as to obtain the speaker information of the target dialogue sentence, wherein the speaker information is the speaker name of the target dialogue sentence or indication information for indicating that the target dialogue sentence has no corresponding speaker name.
[0040] In a possible implementation manner, the determining module is further configured to:
[0041] Segment the target text based on the sentence-end punctuation marks to obtain a plurality of sentences;
[0042] For a sentence in the plurality of sentences that contains quotation marks, dividing the content in the quotation marks and the content outside the quotation marks into different sentences;
[0043] Among all the sentences obtained by dividing the target text into sentences, the content in quotation marks is determined as the dialogue sentences in the target text.
[0044] In a possible implementation, the acquisition module is used to:
[0045] Among all the dialogue sentences in the target text, target dialogue sentences are obtained from front to back according to their positions in the target text.
[0046] In a possible implementation, the identification module is used to:
[0047] In the adjacent sentences of the target dialogue sentence, the speaker name in the speaker list corresponding to the target text is searched, and the adjacent sentences containing the speaker name are determined as the adjacent sentences of the target dialogue sentence.
[0048] In a possible implementation manner, the determining module is used to:
[0049] If only one of the adjacent sentences of the target dialogue sentence contains a speaker name, and the adjacent sentence containing the speaker name has not been determined as a speaker-related sentence corresponding to other dialogue sentences other than the target dialogue sentence, then determining the adjacent sentence containing the speaker name as a speaker-related sentence of the target dialogue sentence;
[0050] If both of the two adjacent sentences of the target dialogue sentence contain the speaker name, and the adjacent sentence in front of the target dialogue sentence has not been determined as the speaker-related sentence corresponding to other dialogue sentences, then determine the adjacent sentence in front as the speaker-related sentence of the target dialogue sentence;
[0051] If the adjacent sentences of the target dialogue sentence all contain the speaker name, and the adjacent sentence in front of the target dialogue sentence has been determined as the speaker-related sentence corresponding to other dialogue sentences, then determine the adjacent sentence behind the target dialogue sentence as the speaker-related sentence of the target dialogue sentence;
[0052] If the target dialogue sentence has no adjacent sentence containing the speaker name, or the adjacent sentence containing the speaker name of the target dialogue sentence has been determined as the speaker-related sentence corresponding to other dialogue sentences, then determine that the target dialogue sentence has no speaker-related sentence.
[0053] In a possible implementation manner, the determining module is further configured to:
[0054] If the speaker-related sentence of the target dialogue sentence is not determined, then determine that the target dialogue sentence has no corresponding speaker name.
[0055] In a possible implementation manner, the speaker recognition model is a machine reading comprehension MRC model;
[0056] The output module is configured to: form a question field with the target dialogue sentence and a preset question sentence, and form a question-related text field with the target dialogue sentence and the speaker-related sentence, where the question sentence is used to query the speaker name of the target dialogue sentence;
[0057] Input the question field and the question-related text field into the MRC model to obtain the speaker information of the target dialogue sentence.
[0058] In a fourth aspect, a device for training a speaker recognition model is provided, and the device includes:
[0059] An obtaining module, configured to obtain sample dialogue sentences in a sample text;
[0060] An identifying module, configured to perform speaker name recognition on adjacent sentences of the sample dialogue sentences based on the sample speaker list corresponding to the sample text, where the sample speaker list includes at least one speaker name;
[0061] A determining module, configured to determine the sample speaker-related sentence corresponding to the sample dialogue sentence among the adjacent sentences of the sample dialogue sentence based on the speaker name recognition result of the adjacent sentence;
[0062] A training module, configured to use the sample dialogue sentence and the sample speaker-related sentence corresponding to the sample dialogue sentence as input sample data to train and adjust the parameters of the speaker recognition model to be trained, so as to obtain an adjusted speaker recognition model;
[0063] An end module, configured to determine the adjusted speaker recognition model as the trained speaker recognition model if the training parameter adjustment meets a preset end condition.
[0064] In a fifth aspect, a computer device is provided. The computer device includes a processor and a memory. At least one instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the methods of the first aspect and its possible implementation manners and the second aspect and its possible implementation manners.
[0065] In a sixth aspect, a computer-readable storage medium is provided. At least one instruction is stored in the storage medium, and the instruction is loaded and executed by the processor to implement the methods of the first aspect and its possible implementation manners and the second aspect and its possible implementation manners.
[0066] In a seventh aspect, a computer program product is provided. At least one instruction is included in the computer program product, and the at least one instruction is loaded and executed by the processor to implement the methods of the first aspect and its possible implementation manners and the second aspect and its possible implementation manners.
[0067] The beneficial effects brought by the technical solutions provided in the embodiments of the present application at least include:
[0068] In the embodiments of the present application, a dialogue sentence and a corresponding speaker-related sentence are obtained from a target text and input into a speaker recognition model to obtain corresponding speaker information. By performing speaker recognition in this way, the speaker recognition model performs recognition on a sentence-by-sentence basis, and the corresponding training samples are also sentence-based. Usually, there are many dialogue sentences and speaker-related sentences in an article. Therefore, only a small number of articles are needed to obtain a large number of dialogue sentences and speaker-related sentences. Thus, when the number of sample articles is small, the method of the embodiments of the present application can improve the recognition accuracy of the speaker recognition model. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0070] Figure 1 It is a flowchart of a method for recognizing a speaker from text provided by an embodiment of the present application;
[0071] Figure 2 It is a schematic process diagram for identifying the names of speakers of adjacent sentences of a target dialogue sentence in a target text provided by an embodiment of the present application;
[0072] Figure 3 It is a schematic process diagram for determining a speaker-related sentence of a target dialogue sentence in adjacent sentences provided by an embodiment of the present application;
[0073] Figure 4 It is a schematic process diagram for inputting a target dialogue sentence and a speaker-related sentence into a trained speaker recognition model provided by an embodiment of the present application;
[0074] Figure 5 It is a flowchart of a method for training a speaker recognition model provided by an embodiment of the present application;
[0075] Figure 6 It is a schematic structural diagram of a device for recognizing a speaker of text provided by an embodiment of the present application;
[0076] Figure 7 It is a schematic structural diagram of a device for training a speaker recognition model provided by an embodiment of the present application;
[0077] Figure 8 It is a block diagram of the structure of a server provided by an embodiment of the present application. Detailed implementation manners
[0078] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0079] An embodiment of the present application provides a method for recognizing a speaker of text, and this method can be implemented by a computer device. The computer device can be a server or a terminal, etc. The server can be a single server or a server cluster composed of multiple servers. The terminal can be a desktop computer, a laptop computer, a mobile phone, a tablet computer, etc.
[0080] The computer device can include a processor, a memory, a communication component, etc., and the processor is respectively connected to the memory and the communication component.
[0081] The processor can be a CPU (Central Processing Unit). The processor can be used to read the text content and process data. For example, perform sentence splitting on the target text, determine dialogue sentences, determine speaker-related sentences, determine the speaker information corresponding to the dialogue sentences based on the dialogue sentences and the speaker-related sentences, train the speaker recognition model to be trained, and so on.
[0082] The memory may include a ROM (Read-Only Memory), a RAM (Random Access Memory), a CD-ROM (Compact Disc Read-Only Memory), a magnetic disk, an optical data storage device, etc. The memory can be used for data storage. For example, it stores the pre-stored data required for clause segmentation of the target text, stores the intermediate data generated during the clause segmentation of the target text, stores the obtained dialogue sentences and the corresponding speaker-related sentences, stores the intermediate data generated during the storage of the obtained dialogue sentences and the corresponding speaker-related sentences, stores the pre-stored data required for training the speaker recognition model to be trained, and so on.
[0083] The communication component can be a wired network connector, a WiFi (Wireless Fidelity) module, a Bluetooth module, a cellular network communication module, etc. The communication component can be used to receive and send signals.
[0084] Figure 1 It is a flowchart of a method for identifying a speaker from text provided by an embodiment of the present application. Refer to Figure 1 This embodiment includes:
[0085] 101. Perform clause segmentation on the target text, and among the multiple sentences obtained after clause segmentation, determine all the dialogue sentences in the target text.
[0086] 102. Obtain the target dialogue sentence in the target text.
[0087] 103. Based on the speaker list corresponding to the target text, perform speaker name recognition on the adjacent sentences of the target dialogue sentence.
[0088] 104. Based on the speaker name recognition result of the adjacent sentences, determine the speaker-related sentences of the target dialogue sentence among the adjacent sentences.
[0089] 105. Input the target dialogue sentence and the speaker-related sentences into the trained speaker recognition model to obtain the speaker information of the target dialogue sentence.
[0090] Next, the specific process of speaker recognition for text will be introduced in detail:
[0091] In step 101, the computer device performs clause segmentation on the target text, and among the multiple sentences obtained after clause segmentation, determines the dialogue sentences in the target text. Among them, the target text can be the text of an article containing dialogues, such as a novel text, a narrative text, a play text, etc. Clause segmentation means splitting the target text into multiple sentences based on the punctuation marks in the target text.
[0092] In implementation, the initial text can be preprocessed to obtain the target text. Among them, the initial text can be the original unprocessed article text, which may contain some low-level errors that need to be removed through preprocessing. The preprocessing can be to remove redundant punctuation marks, incorrect punctuation marks, and other symbols that should not appear in the initial text. Redundant punctuation marks can be the continuous use of multiple identical punctuation marks (such as commas, semicolons), incorrect punctuation marks can be using multiple commas together as an ellipsis or using a comma at the end of a paragraph, and symbols that should not appear can be spaces, etc. Technicians can first collect these common low-level errors and the corresponding modification rules for each low-level error, and store these low-level errors and the corresponding modification rules in a computer device in advance. Then, traverse the initial text. When a certain error content is traversed, modify the error content accordingly based on the modification rule corresponding to the error content. Among them, for the problem of having redundant punctuation marks, the redundant punctuation marks can be deleted. For example, if there are two consecutive full stops, one of the full stops is deleted, etc. For the problem of using incorrect punctuation marks, the incorrect punctuation marks can be replaced with correct punctuation marks. For example, replace multiple consecutive commas with an ellipsis, etc. For the problem of having symbols that should not appear, these symbols that should not appear can be directly deleted. For example, delete the spaces in the text, etc. In this way, the target text no longer has the above-mentioned low-level errors. Therefore, when the computer device performs subsequent processing on relevant texts, it will not be affected by low-level errors, and the accuracy of the computer device in performing subsequent processing on texts can be improved.
[0093] After obtaining the target text, the target text can be clause-separated, and then all the dialogue sentences in the target text can be determined. The specific process is as follows:
[0094] First step, based on the positions of the end punctuation marks in the target text, the target text is clause-separated to obtain multiple clauses. Among them, the end punctuation marks can be full stops, exclamation marks, question marks, ellipses and other punctuation marks.
[0095] In implementation, technicians can first collect all types of end punctuation marks and pre-store all types of end punctuation marks in a computer device. Then, traverse the target text in the order from front to back. When an end punctuation mark in the target text is traversed, a number can be assigned to the currently traversed end punctuation mark. If the currently traversed end punctuation mark is the first end punctuation mark in the target text, the currently traversed end punctuation mark and the text before it can be determined as a sentence. If the currently traversed end punctuation mark is not the first end punctuation mark in the target text, the currently traversed end punctuation mark and the text after the currently traversed end punctuation mark and after the previous traversed end punctuation mark can be determined as a sentence. In this way, after the above preliminary sentence splitting process on the target text, multiple sentences can be obtained. The number of multiple sentences can be N, where N is a positive integer.
[0096] Second step, for sentences containing quotation marks among the multiple sentences, the content inside the quotation marks and the content outside the quotation marks are divided into different sentences. Among them, the quotation marks can be double quotation marks, including the opening quotation mark and the closing quotation mark.
[0097] In implementation, after the above preliminary sentence splitting process on the target text, N sentences can be obtained. Among these N sentences, there may be T sentences containing quotation marks (T is a positive integer and T is less than N). Further sentence splitting processing can be performed on these T sentences containing quotation marks. Among these T sentences containing double quotation marks, based on the position of the double quotation marks, the opening quotation mark, the closing quotation mark, and the text inside the opening and closing quotation marks can be divided into separate sentences. If there is still text outside the double quotation marks in the sentence containing double quotation marks, the text before the opening quotation mark and the text after the closing quotation mark are both divided into separate sentences. In this way, after the above further sentence splitting process on the target text, all sentences of the target text can be obtained. The number of sentences can be M, where M is a positive integer and M is greater than N.
[0098] Third step, among all the sentences obtained by splitting the target text, the sentences corresponding to the content inside the double quotation marks are determined as the dialogue sentences in the target text.
[0099] In implementation, the computer device can record the sentence identifiers of the sentences corresponding to the content inside the double quotation marks (the sentence identifiers can be sequence numbers or the position information of the double quotation marks in the text, etc.), and use these sentence identifiers as the sentence identifiers of the dialogue sentences. Optionally, when the computer device determines all the dialogue sentences in the target text, each dialogue sentence can be numbered according to its order in the target text for subsequent calls.
[0100] In step 102, obtain the target dialogue sentences in the target text.
[0101] Among them, the target dialogue sentence can be any one of all the dialogue sentences. The target dialogue sentence is the dialogue sentence for which the speaker name needs to be recognized currently. In implementation, after determining all the dialogue sentences of the target text, the dialogue sentences can be obtained one by one in the order from front to back according to their positions in the target text, so as to perform the subsequent speaker name recognition process.
[0102] In step 103, based on the speaker list corresponding to the target text, the speaker names of the adjacent sentences of the target dialogue sentence in the target text are recognized.
[0103] Among them, the speaker list includes at least one speaker name, and these speaker names are the names of all speakers in the target text. The speaker list can be a list composed of the names of all characters in a novel text, or a list composed of the names of all actors in a script text. Speaker name recognition refers to finding the speaker name in a sentence.
[0104] In implementation, the speaker list can be pre-stored in a computer device. The computer device can first determine the position of the target dialogue sentence. Correspondingly, in the adjacent sentences of the target dialogue sentence, the adjacent sentence located in front of the target dialogue sentence can be determined as the first adjacent sentence, and the adjacent sentence located behind the target dialogue sentence can be determined as the second adjacent sentence. Then, word segmentation processing is performed on the contents of the first adjacent sentence and the second adjacent sentence. Then, the computer device compares each word of the first adjacent sentence with the speaker names in the speaker list respectively to determine whether the corresponding speaker name is included in the first adjacent sentence, and compares each word of the second adjacent sentence with the speaker names in the speaker list respectively to determine whether the corresponding speaker name is included in the second adjacent sentence. As Figure 2 shown, a specific example is given. After obtaining the target dialogue sentence, the speaker names of the adjacent sentences of the target dialogue sentence are recognized.
[0105] Optionally, if the computer device finds the same character name as in the speaker list in the first adjacent sentence or the second adjacent sentence of the target dialogue sentence, the character name in the first adjacent sentence or the second adjacent sentence can be replaced with a special character. Among them, the special character can be a character with a label. For example, PERSON i, where i is a positive integer.
[0106] In implementation, the preset character PERSON can be pre-stored in the vocabulary of the computer device. For example, the preset character PERSON can be pre-stored in the Bert (vocabulary name) vocabulary of the computer device. Before performing step 103, the computer device can first read the speaker list and number each different speaker name in the order in the speaker list. For example, speaker names such as Xiaoming, Xiaohong,..., Xiaogang are numbered 1, 2,..., i in sequence. Here, i is a positive integer. Then, the preset character PERSON is combined with the number corresponding to the speaker name to form the special character PERSONi. According to the one-to-one correspondence between different speaker names and the special character PERSONi, a speaker name - special character correspondence table is constructed. It can be seen that the special character is used to uniquely identify the speaker, so the special character can also be called the speaker unique identifier.
[0107] Table 1
[0108] Name Special Character Xiaoming PERSON 1 Xiaohong PERSON 2 Xiaogang PERSON 3 …… ……
[0109] After assigning special characters to the speaker names in the speaker list, the computer device performs speaker name recognition on the adjacent sentences of the target dialogue sentence. When the computer device performs speaker name recognition on the adjacent sentences of the target dialogue sentence, if a certain speaker name is recognized, the special character corresponding to the speaker name can be searched in the above correspondence table. Then, the speaker name in this adjacent sentence is replaced with the found special character. For example, if the speaker name "Xiaoming" is recognized in the adjacent sentence of the target dialogue sentence, and according to the above correspondence table, the special character corresponding to the speaker name "Xiaoming" is PERSON 1, then the speaker name "Xiaoming" can be replaced with the special character PERSON 1.
[0110] In step 104, based on the recognition result of the speaker name recognition, the speaker-related sentence of the target dialogue sentence is determined in the adjacent sentence.
[0111] Among them, the speaker-related sentence is a sentence used to reflect the speaker name of the dialogue sentence, generally located in the front or rear of the dialogue sentence and adjacent to the dialogue sentence. As Figure 3 shown, a specific example is given to determine one of the adjacent sentences of the target dialogue sentence as the speaker-related sentence of the target dialogue sentence.
[0112] The following gives several possible recognition results and the corresponding determination methods of the speaker-related sentence:
[0113] (1) If only one of the adjacent sentences of the first dialogue sentence contains the speaker name, and the adjacent sentence containing the speaker name is not determined as the speaker-related sentence corresponding to other dialogue sentences outside the first dialogue sentence, then the adjacent sentence containing the speaker name is determined as the speaker-related sentence of the first dialogue sentence.
[0114] Among them, only one of the adjacent sentences of the target dialogue sentence contains the speaker name, including the following two cases: The target dialogue sentence in the target text has only one adjacent sentence, and the adjacent sentence contains the speaker name; The target dialogue sentence has two adjacent sentences, but only one of the adjacent sentences contains the speaker name. The target dialogue sentence in the target text has only one adjacent sentence, and the adjacent sentence contains the speaker name, which may be the case where the target dialogue sentence is the first or last sentence of the target text. The target dialogue sentence has two adjacent sentences, but only one of the adjacent sentences contains the speaker name, which may be the case where the target dialogue sentence is neither the first nor the last sentence of the target text, etc.
[0115] In implementation, when the target dialogue sentence has only one adjacent sentence, if the adjacent sentence contains the speaker name and the adjacent sentence is not determined as the speaker-related sentence corresponding to other dialogue sentences, then the adjacent sentence is determined as the speaker-related sentence of the target dialogue sentence.
[0116] When the target dialogue sentence has two adjacent sentences, if among the first adjacent sentence and the second adjacent sentence, only one adjacent sentence contains the speaker name and the adjacent sentence is not determined as the speaker-related sentence corresponding to other dialogue sentences, then the adjacent sentence is determined as the speaker-related sentence of the target dialogue sentence. Among them, the first adjacent sentence may be the adjacent sentence located in the front of the target dialogue sentence, and the second adjacent sentence may be the adjacent sentence located in the back of the target dialogue sentence.
[0117] (2) If both adjacent sentences of the target dialogue sentence contain the speaker name, and the adjacent sentence in the front of the target dialogue sentence is not determined as the speaker-related sentence corresponding to other dialogue sentences, then the adjacent sentence in the front is determined as the speaker-related sentence of the target dialogue sentence.
[0118] Among them, the adjacent sentence in the front of the target dialogue sentence is hereinafter simply referred to as the first adjacent sentence.
[0119] In implementation, there may be different orders between the target dialogue sentence and the adjacent sentence with the speaker's name. The adjacent sentence with the speaker's name can be before or after the target dialogue sentence. There may be no difference in the meaning between the two, but when determining the speaker-related sentence of the target dialogue sentence, the determination results may be different. Usually, it is a more writing-habitual order that the adjacent sentence with the speaker's name is before the target dialogue sentence. When the adjacent sentence with the speaker's name is before the target dialogue sentence, the adjacent sentence with the speaker's name and the target dialogue sentence can be shown as: [Xiaogang is sitting on the chair. Xiaoming says, "Is today a working day?". Xiaogang nods to indicate yes.]. And when the adjacent sentence with the speaker's name is after the target dialogue sentence, the target dialogue sentence and the adjacent sentence with the speaker's name can be shown as: [Xiaogang is sitting on the chair. "Is today a working day?" Xiaoming says. Xiaogang nods to indicate yes.]. In the above two cases, when ["Is today a working day?"] is the target dialogue sentence, the speaker-related sentence of the target dialogue sentence should be [Xiaoming says:] or [Xiaoming says.]. Since the first case is a more writing-habitual order. Therefore, if both adjacent sentences of the target dialogue sentence contain the speaker's name, and the first adjacent sentence of the target dialogue sentence has not been determined as the speaker-related sentence corresponding to other dialogue sentences, then the first adjacent sentence is determined as the speaker-related sentence of the target dialogue sentence. In this way, the probability of correctly corresponding the target dialogue sentence with the speaker-related sentence can be increased.
[0120] (3) If both adjacent sentences of the target dialogue sentence contain the speaker's name, and the adjacent sentence in front of the target dialogue sentence has been determined as the speaker-related sentence corresponding to other dialogue sentences, then the adjacent sentence behind the target dialogue sentence is determined as the speaker-related sentence of the target dialogue sentence.
[0121] Among them, the adjacent sentence behind the target dialogue sentence is hereinafter referred to as the second adjacent sentence for short.
[0122] In implementation, it is more in line with the writing convention that the relevant sentence containing the speaker's name is located before the dialogue sentence. However, in the target text, there may also be a situation where the relevant sentence containing the speaker's name is located after the dialogue sentence. For example, ["It's a beautiful day today!" said Xiaoming. "Yes, let's go out and play later!" said Xiaohong with a smile.]. In this case, when ["It's a beautiful day today!"] is the target dialogue sentence, there may be only one adjacent sentence containing the speaker's name for the target dialogue sentence. At this time, ["Xiaoming said"] will be determined as the speaker-related sentence. Therefore, when ["Yes, let's go out and play later!"] is the target dialogue sentence, both the first adjacent sentence and the second adjacent sentence of the target dialogue sentence contain the speaker's name. However, since the first adjacent sentence has been determined as the speaker-related sentence for ["It's a beautiful day today!"], at this time, the second adjacent sentence of the target dialogue sentence will be determined as the speaker-related sentence of the target dialogue sentence. In this way, it is possible to avoid the same adjacent sentence being determined as the speaker-related sentence for multiple dialogue sentences, thereby increasing the probability of correctly corresponding the target dialogue sentence with the speaker-related sentence.
[0123] (4) If the target dialogue sentence has no adjacent sentence containing the speaker's name, or the adjacent sentence containing the speaker's name of the target dialogue sentence has been determined as the speaker-related sentence corresponding to other dialogue sentences, then it is determined that the target dialogue sentence has no speaker-related sentence.
[0124] In implementation, there may be a situation where there is a dialogue sentence without a specific speaker. In this case, it is necessary to determine that the target dialogue sentence has no speaker-related sentence. The dialogue sentence without a specific speaker can be: [After the performance, there was enthusiastic applause in the auditorium, "The performance was really good!". "Thank you all!" said Xiaoming, bowing.]. When "The performance was really good!" is the target dialogue sentence, the target dialogue sentence has no adjacent sentence containing the speaker's name. At this time, it is determined that the target dialogue sentence has no speaker-related sentence.
[0125] In implementation, there may also be a situation where the adjacent sentence containing the speaker's name of the target dialogue sentence has been determined as the speaker-related sentence corresponding to other dialogue sentences. In this case, it is necessary to determine that the target dialogue sentence has no speaker-related sentence. The situation where the adjacent sentence containing the speaker's name of the target dialogue sentence has been determined as the speaker-related sentence corresponding to other dialogue sentences can be: [After the performance, there was enthusiastic applause in the auditorium, "Thank you all!" said Xiaoming, bowing. "Good! Good! Good!", there were bursts of cheers from the audience.]. In this case, when "Thank you all!" is the target dialogue sentence, "Xiaoming bowed and said" will be determined as the speaker-related sentence. Therefore, when "Good! Good! Good!" is the target dialogue sentence, the adjacent sentence containing the speaker's name of the target dialogue sentence has been determined as the speaker-related sentence corresponding to other dialogue sentences. Therefore, at this time, it is determined that the target dialogue sentence has no speaker-related sentence.
[0126] In this way, when the adjacent sentence of the target dialogue sentence does not contain the speaker name, or when the adjacent sentence of the target dialogue sentence that contains the speaker name has been determined to be the speaker-related sentence corresponding to other dialogue sentences, it can be determined that the target dialogue sentence has no speaker-related sentence, and then it can be determined that the dialogue sentence has no corresponding speaker name, rather than determining a wrong speaker name for the target dialogue sentence.
[0127] Optionally, during the process of determining the speaker-related sentence of the target dialogue sentence, if the speaker-related sentence of the target dialogue sentence is not determined, it is determined that the target dialogue sentence has no corresponding speaker name. If the speaker-related sentence of the target dialogue sentence is determined, then step 105 is executed.
[0128] In step 105, the target dialogue sentence and the speaker-related sentence are input into the trained speaker recognition model to obtain the speaker information of the target dialogue sentence.
[0129] Among them, the speaker recognition model can be a trained machine learning model, and the specific algorithm adopted by the machine learning model can be set according to actual needs. The speaker recognition model can be a machine reading comprehension MRC (Machine Reading Comprehension) model, etc. The MRC model is a text-based question-and-answer model. The speaker information can be the speaker name of the target dialogue sentence or the indication information used to indicate that the target dialogue sentence has no corresponding speaker name.
[0130] In implementation, the target dialogue sentence can be determined as Q (Quote), the speaker-related sentence can be determined as S (Speaker), and then the target dialogue sentence and the speaker-related sentence are determined as a Q*S pair or an S*Q pair. After determining a Q*S pair or an S*Q pair, the determined Q*S pair or S*Q pair can be input into the trained speaker recognition model.
[0131] For the case of using the MRC model, the target dialogue sentence and the preset question sentence can be combined to form a question field, and the target dialogue sentence and the speaker-related sentence can be combined to form a question-related text field. The question sentence is used to ask about the speaker name of the target dialogue sentence, and the content of the question sentence can be preset, such as "Who said this sentence?". The question field and the question-related text field are the fields required by the MRC model. After determining the question field and the question-related text field, an identifier can be added in front of the question field, and this identifier is used to indicate that a question will be asked of the machine learning model. For example, an identifier is added in front of the question field <cls>Then, a separator can be added between the question field and the question-related text field, as well as at the rear of the question-related text field. This separator is used to distinguish the question field from the question-related text field. For example, a separator is added between the question field and the question-related text field, as well as at the rear of the question-related text field <sep>After adding the above identifiers and delimiters to the question field and the question-related text field, the input data is obtained, as Figure 4 shown. Finally, the input data is input into the trained speaker recognition model to obtain the speaker information of the target dialogue sentence.
[0132] After corresponding prediction processing, the speaker recognition model can output the speaker information of the target dialogue sentence. After completing step 105, the target dialogue sentence can be labeled based on the speaker information.
[0133] Annotation method 1
[0134] Determine the annotation information based on the speaker information, and record the corresponding annotation information for the target dialogue sentence. When there is a speaker for the target dialogue sentence, the content of the annotation information can be the name of the speaker of the target dialogue sentence or the special character corresponding to the speaker name. When there is no speaker for the target dialogue sentence, the content of the annotation information can be an indication information used to indicate that there is no corresponding speaker name for the target dialogue sentence, such as 00000000.
[0135] Annotation method 2
[0136] Pre-establish a correspondence table between the dialogue sentence number and the speaker name. As an example of a correspondence representation shown in Table 2, the correspondence table includes all the speaker names in the speaker list of the target text, and the correspondence table can also include an entry of "no speaker name". After determining the speaker information of the target dialogue sentence, obtain the sequence number of the target dialogue sentence among all the dialogue sentences in the target text, such as 1, 2, 3... If the speaker information is the speaker name, in the above correspondence table, record the sequence number of the target dialogue sentence corresponding to the speaker name. If the speaker information is an indication information used to indicate that there is no corresponding speaker name for the target dialogue sentence, in the above correspondence table, record the sequence number of the target dialogue sentence corresponding to no speaker name.
[0137] Table 2
[0138] Dialogue Sentence Number Speaker Name 1、21…… No Speaker Name 2、4、6、15、17、19…… Xiaohong 3、5、7、8、24、26…… Xiaogang …… ……
[0139] After determining the speaker names of all dialogue sentences in the target text, an audiobook reading audio can be generated based on whether each sentence in the target text is a dialogue sentence, whether each dialogue sentence has a corresponding speaker name, the speaker name corresponding to each dialogue sentence, etc. For sentences in the target text that are not dialogue sentences, the first voice can be used to synthesize the audio to indicate that these sentences in the target text are narration. For dialogue sentences in the target text, when a dialogue sentence does not have a corresponding speaker name, the second voice can be used to synthesize the audio to indicate that these sentences in the target text are dialogue sentences without a specific speaker. When a dialogue sentence has a corresponding speaker name, different voices can be used to synthesize the audio based on different speaker names to distinguish the content of each specific speaker's dialogue sentence in the target text. For example, for dialogue sentences in the target text with the speaker name Xiaohong, the third voice is used to synthesize the audio, and for dialogue sentences in the target text with the speaker name Xiaoming, the fourth voice is used to synthesize the audio, and so on. In this way, during the audiobook reading, a multi-speaker dubbing reading mode can be achieved, thereby improving the user experience.
[0140] An embodiment of the present application provides a method for training a speaker recognition model. Figure 5 It is a flowchart of a method for training a speaker recognition model provided by an embodiment of the present application. Refer to Figure 5 This embodiment includes:
[0141] 501. Obtain sample dialogue sentences in the sample text.
[0142] Among them, the sample text refers to the text used for training the model. The sample dialogue sentence can be any dialogue sentence in the sample text.
[0143] Before step 501, the initial text can be preprocessed and clause-separated to obtain the sample text, and then the dialogue sentences are determined in the sample text. The processes of preprocessing and clause-separating the initial text, and the process of determining the dialogue sentences in the sample text are the same as those of step 101, and will not be repeated here.
[0144] In implementation, after obtaining all the dialogue sentences in the sample text, according to the positions of all the dialogue sentences in the sample text, each sample dialogue sentence can be obtained in order from front to back for the subsequent speaker name recognition processing of the following steps.
[0145] 502. Based on the sample speaker list corresponding to the sample text, perform speaker name recognition on adjacent sentences of the sample dialogue sentences in the sample text.
[0146] Among them, the sample speaker list includes at least one speaker name, which are the names of all speakers in the sample text. The sample speaker list can be a list composed of the names of all characters in the sample novel text, or a list composed of the names of all actors in the sample script text. Speaker name recognition refers to finding the speaker name in a sentence.
[0147] In implementation, the speaker list can be pre-stored in a computer device. The computer device can first determine the position of the target dialogue sentence. Correspondingly, in the adjacent sentences of the target dialogue sentence, the adjacent sentence located in front of the target dialogue sentence can be determined as the first adjacent sentence, and the adjacent sentence located behind the target dialogue sentence can be determined as the second adjacent sentence. Then, word segmentation processing is performed on the content of the first adjacent sentence and the second adjacent sentence. Then, the computer device compares each word of the first adjacent sentence with the speaker names in the speaker list respectively to determine whether the corresponding speaker name is included in the first adjacent sentence, and compares each word of the second adjacent sentence with the speaker names in the speaker list respectively to determine whether the corresponding speaker name is included in the second adjacent sentence.
[0148] 503. Based on the speaker name recognition result of the adjacent sentence, determine the sample speaker-related sentence corresponding to the sample dialogue sentence in the adjacent sentences of the sample dialogue sentence.
[0149] The following gives several possible recognition results and the corresponding determination methods of the speaker-related sentence:
[0150] (1) If only one adjacent sentence of the target dialogue sentence contains a speaker name and this adjacent sentence has not been determined as the speaker-related sentence corresponding to other dialogue sentences, then determine this adjacent sentence as the speaker-related sentence of the target dialogue sentence.
[0151] (2) If both adjacent sentences of the target dialogue sentence contain speaker names and the adjacent sentence in front of the target dialogue sentence has not been determined as the speaker-related sentence corresponding to other dialogue sentences, then determine the adjacent sentence in front as the speaker-related sentence of the target dialogue sentence.
[0152] (3) If both adjacent sentences of the target dialogue sentence contain speaker names and the adjacent sentence in front of the target dialogue sentence has been determined as the speaker-related sentence corresponding to other dialogue sentences, then determine the adjacent sentence behind the target dialogue sentence as the speaker-related sentence of the target dialogue sentence.
[0153] (4) If there is no adjacent sentence containing a speaker name for the target dialogue sentence, or the adjacent sentence containing a speaker name of the target dialogue sentence has been determined as the speaker-related sentence corresponding to other dialogue sentences, then determine that the target dialogue sentence has no speaker-related sentence.
[0154] The possible situations of the above several recognition results are the same as those in step 104, and will not be repeated here.
[0155] 504. Using the sample dialogue sentence and the sample speaker-related sentence corresponding to the sample dialogue sentence as input sample data, train and tune the parameters of the speaker recognition model to be trained, and obtain the speaker recognition model after parameter tuning.
[0156] In implementation, a technician can first determine the sample speaker-related sentence and the sample dialogue sentence as input sample data. Based on the content of the sample speaker-related sentence and the sample dialogue sentence, first label the correct speaker name or indication information (in the case of no corresponding speaker name) for the input sample data as the corresponding reference speaker information. Then, input the input sample data into the speaker recognition model to be trained to obtain the predicted speaker information. Then, the predicted speaker information and the reference speaker information can be input into a preset loss function to calculate the corresponding loss value. Further, based on the loss value, calculate the adjustment value corresponding to each parameter to be adjusted in the speaker recognition model to be trained, and adjust the corresponding parameter to be adjusted based on the adjustment value to obtain the speaker recognition model after parameter tuning.
[0157] 505. If the training parameter tuning meets the preset end condition, determine the speaker recognition model after parameter tuning as the trained speaker recognition model.
[0158] In implementation, a technician can train and tune the parameters of the speaker recognition model to be trained based on the calculated loss value, so that the speaker prediction result output by the speaker recognition model to be trained is more accurate, thereby obtaining an accurate speaker recognition model. Use different sample dialogue sentences and corresponding sample speaker-related sentences to train the speaker recognition model to be trained multiple times until the preset end condition is reached, and then stop training. There can be various preset end conditions. The following gives several situations that meet the preset end condition:
[0159] (1) The number of times of training the speaker recognition model to be trained using different sample dialogue sentences and corresponding sample speaker-related sentences reaches the training times threshold.
[0160] Among them, the training times threshold can be any reasonable value. For example, it can be 5000, or it can be 10000, etc. The embodiments of the present application do not limit this. In implementation, a technician can preset the training times threshold in advance. When the training times reach the training times threshold, training can be stopped, and the speaker recognition model obtained after the last training can be determined as the trained speaker recognition model.
[0161] (2) The loss value in consecutive preset numbers of trainings is less than the preset loss value threshold.
[0162] Among them, both the preset number and the preset loss value threshold can be any reasonable values. For example, the preset number can be 50 or 100, and the embodiments of the present application do not limit this. During implementation, through successive parameter adjustments, in general, the loss value between the predicted speaker information output by the speaker recognition model being trained and the reference speaker information will gradually decrease. If after a certain number of training sessions (such as after 10,000 training sessions), the loss values in consecutive preset numbers of training sessions are all less than the preset loss value threshold, then the current model can be considered a trained speaker recognition model.
[0163] (3) In consecutive preset numbers of training sessions, the model recognition accuracy exceeds the preset accuracy threshold.
[0164] Among them, both the preset number and the preset accuracy threshold can be any reasonable values. For example, the preset number can be 50 or 100, etc., and the preset accuracy threshold can be 90%, etc., and the embodiments of the present application do not limit this.
[0165] During implementation, through successive parameter adjustments, in general, the recognition accuracy of the speaker recognition model being trained will gradually increase. If after a certain number of training sessions (such as after 1,000 training sessions), in consecutive preset numbers of training sessions, the model recognition accuracy exceeds the preset accuracy threshold, then the current model can be considered a trained speaker recognition model.
[0166] All of the above optional technical solutions can be combined arbitrarily to form optional embodiments of the present application, which will not be elaborated one by one here.
[0167] The beneficial effects brought by the technical solutions provided by the embodiments of the present application are as follows: In the embodiments of the present application, the dialogue sentences and the corresponding speaker-related sentences are obtained from the target text and input into the speaker recognition model to obtain the corresponding speaker information. By performing speaker recognition in this way, the speaker recognition model performs recognition on a sentence-by-sentence basis, and the corresponding training samples are also sentence-based. Usually, there are many dialogue sentences and speaker-related sentences in an article. Therefore, only a small number of articles are needed to obtain a large number of dialogue sentences and speaker-related sentences. Thus, when the number of sample articles is small, by using the method of the embodiments of the present application, the recognition accuracy of the speaker recognition model can be improved.
[0168] The embodiments of the present application provide a device for recognizing a speaker from text. The device can be the computer device in the above embodiments, such as Figure 6 as shown, the device includes:
[0169] An acquisition module 610, configured to acquire the target dialogue sentences in the target text;
[0170] A recognition module 620, configured to perform speaker name recognition on adjacent sentences of the target dialogue sentence based on a speaker list corresponding to the target text, wherein the speaker list includes at least one speaker name;
[0171] A determination module 630, configured to determine a speaker-related sentence of the target dialogue sentence in the adjacent sentences based on the speaker name recognition result of the adjacent sentences;
[0172] The output module 640 is used to input the target dialogue sentence and the speaker-related sentence into a trained speaker recognition model if the speaker-related sentence of the target dialogue sentence is determined, so as to obtain the speaker information of the target dialogue sentence, wherein the speaker information is the speaker name of the target dialogue sentence or indication information for indicating that the target dialogue sentence has no corresponding speaker name.
[0173] In a possible implementation, the acquisition module 610 is used to:
[0174] Segment the target text based on the sentence-end punctuation marks to obtain a plurality of sentences;
[0175] For a sentence in the plurality of sentences that contains quotation marks, dividing the content in the quotation marks and the content outside the quotation marks into different sentences;
[0176] Among all the sentences obtained by dividing the target text into sentences, the content in quotation marks is determined as the dialogue sentences in the target text.
[0177] In a possible implementation, the acquisition module 610 is further configured to:
[0178] Among all the dialogue sentences in the target text, target dialogue sentences are obtained from front to back according to their positions in the target text.
[0179] In a possible implementation, the identification module 620 is configured to:
[0180] In the adjacent sentences of the target dialogue sentence, the speaker name in the speaker list corresponding to the target text is searched, and the adjacent sentences containing the speaker name are determined as the adjacent sentences of the target dialogue sentence.
[0181] In a possible implementation, the determining module 630 is configured to:
[0182] If only one of the adjacent sentences of the target dialogue sentence contains a speaker name, and the adjacent sentence containing the speaker name has not been determined as a speaker-related sentence corresponding to other dialogue sentences other than the target dialogue sentence, then determining the adjacent sentence containing the speaker name as a speaker-related sentence of the target dialogue sentence;
[0183] If both of the two adjacent sentences of the target dialogue sentence contain the speaker name, and the adjacent sentence in front of the target dialogue sentence has not been determined as the speaker-related sentence corresponding to other dialogue sentences, then determine the adjacent sentence in front as the speaker-related sentence of the target dialogue sentence;
[0184] If the adjacent sentences of the target dialogue sentence all contain the speaker name, and the adjacent sentence in front of the target dialogue sentence has been determined as the speaker-related sentence corresponding to other dialogue sentences, then determine the adjacent sentence behind the target dialogue sentence as the speaker-related sentence of the target dialogue sentence;
[0185] If the target dialogue sentence has no adjacent sentence containing the speaker name, or the adjacent sentence containing the speaker name of the target dialogue sentence has been determined as the speaker-related sentence corresponding to other dialogue sentences, then determine that the target dialogue sentence has no speaker-related sentence.
[0186] In a possible implementation manner, the determining module 630 is further configured to:
[0187] If the speaker-related sentence of the target dialogue sentence is not determined, then determine that the target dialogue sentence has no corresponding speaker name.
[0188] In a possible implementation manner, the speaker recognition model is a machine reading comprehension MRC model;
[0189] The output module 640 is configured to: form a question field with the target dialogue sentence and a preset question sentence, and form a question-related text field with the target dialogue sentence and the speaker-related sentence, where the question sentence is used to inquire about the speaker name of the target dialogue sentence;
[0190] Input the question field and the question-related text field into the MRC model to obtain the speaker information of the target dialogue sentence.
[0191] The embodiment of the present application provides a device for training a speaker recognition model. The device may be the computer device in the above embodiment, such as Figure 6 As shown, the device includes:
[0192] An obtaining module 710, configured to obtain sample dialogue sentences in the sample text;
[0193] An identifying module 720, configured to perform speaker name recognition on the adjacent sentences of the sample dialogue sentences based on the sample speaker list corresponding to the sample text, where the sample speaker list includes at least one speaker name;
[0194] A determination module 730, configured to determine a sample speaker-related sentence corresponding to the sample dialogue sentence from adjacent sentences of the sample dialogue sentence based on the speaker name recognition result of the adjacent sentences;
[0195] A training module 740, configured to use the sample dialogue sentence and the sample speaker-related sentence corresponding to the sample dialogue sentence as input sample data to train and adjust the parameters of a speaker recognition model to be trained, so as to obtain an adjusted speaker recognition model;
[0196] An end module 750, configured to determine the adjusted speaker recognition model as a trained speaker recognition model if the training and parameter adjustment meet a preset end condition.
[0197] It should be noted that when the apparatus for recognizing a speaker from text provided in the above embodiment performs functions, only the division of the above function modules is used as an example for illustration. In practical applications, the above functions can be allocated to different function modules according to needs, that is, the internal structure of the apparatus is divided into different function modules to complete all or part of the functions described above. In addition, the apparatus for recognizing a speaker from text provided in the above embodiment and the method embodiment for recognizing a speaker from text belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be elaborated here.
[0198] Figure 8 FIG. is a schematic structural diagram of a server provided in an embodiment of the present application. The server 800 may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPUs) 801 and one or more memories 802. Among them, at least one instruction is stored in the memory 802, and the at least one instruction is loaded and executed by the processor 801 to implement the methods provided in the above various method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input / output. The server may also include other components for implementing device functions, which will not be elaborated here.
[0199] In an exemplary embodiment, a computer-readable storage medium is further provided. For example, a memory including instructions, and the above instructions can be executed by a processor in a terminal to complete the method for controlling a drone in the above embodiment. The computer-readable storage medium may be non-transitory. For example, the computer-readable storage medium may be a ROM (read-only memory), a RAM (random access memory), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0200] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and the storage medium mentioned above can be a read-only memory, a disk, an optical disc, or the like.
[0201] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals (including but not limited to signals transmitted between a user terminal and other devices, etc.) involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.
[0202] The above are only alternative embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.< / sep> < / cls>
Claims
1. A method for identifying the speaker of text, characterized in that, The method includes: Obtaining a target dialogue sentence in the target text, where the target dialogue sentence is any dialogue sentence in the target text; Based on the speaker list corresponding to the target text, performing speaker name recognition on adjacent sentences of the target dialogue sentence, where the speaker list includes at least one speaker name; Based on the speaker name recognition result of the adjacent sentences, determining the speaker-related sentence of the target dialogue sentence in the adjacent sentences; If the speaker-related sentence of the target dialogue sentence is determined, inputting the target dialogue sentence and the speaker-related sentence into a trained speaker recognition model to obtain the speaker information of the target dialogue sentence, where the speaker information is the speaker name of the target dialogue sentence or indication information for indicating that there is no corresponding speaker name for the target dialogue sentence; Where the determining the speaker-related sentence of the target dialogue sentence in the adjacent sentences based on the speaker name recognition result of the adjacent sentences includes: If there is only one adjacent sentence of the target dialogue sentence that contains a speaker name, and the adjacent sentence containing the speaker name has not been determined as the speaker-related sentence corresponding to other dialogue sentences outside the target dialogue sentence, then determining the adjacent sentence containing the speaker name as the speaker-related sentence of the target dialogue sentence; If both adjacent sentences of the target dialogue sentence contain speaker names, and the adjacent sentence in front of the target dialogue sentence has not been determined as the speaker-related sentence corresponding to other dialogue sentences, then determining the adjacent sentence in front as the speaker-related sentence of the target dialogue sentence; If both adjacent sentences of the target dialogue sentence contain speaker names, and the adjacent sentence in front of the target dialogue sentence has been determined as the speaker-related sentence corresponding to other dialogue sentences, then determining the adjacent sentence behind the target dialogue sentence as the speaker-related sentence of the target dialogue sentence; If there is no adjacent sentence of the target dialogue sentence that contains a speaker name, or the adjacent sentence containing the speaker name of the target dialogue sentence has been determined as the speaker-related sentence corresponding to other dialogue sentences, then it is determined that the target dialogue sentence has no speaker-related sentence.
2. The method according to claim 1, characterized in that Before obtaining the target dialogue sentence in the target text, the method further includes: Performing sentence segmentation on the target text based on the end punctuation marks to obtain multiple sentences; For the sentences containing quotation marks among the multiple sentences, dividing the content inside the quotation marks and the content outside the quotation marks into different sentences; Among all the sentences obtained by performing sentence segmentation on the target text, determining the content inside the quotation marks as the dialogue sentences in the target text.
3. The method according to claim 1, wherein The obtaining the target dialogue sentence in the target text includes: Obtaining the target dialogue sentence in the order from the front to the back in terms of the position in the target text among all the dialogue sentences in the target text.
4. The method according to claim 1, wherein The performing speaker name recognition on the adjacent sentences of the target dialogue sentence based on the speaker list corresponding to the target text includes: In the adjacent sentences of the target dialogue sentence, searching for the speaker names in the speaker list corresponding to the target text, and determining the adjacent sentences containing the speaker names as the adjacent sentences of the target dialogue sentence.
5. The method according to any one of claims 1-4, characterized in that, After determining the speaker-related sentence of the target dialogue sentence in the adjacent sentences based on the speaker name recognition result of the adjacent sentences, the method further includes: If the speaker-related sentence of the target dialogue sentence is not determined, it is determined that the target dialogue sentence has no corresponding speaker name.
6. The method according to any one of claims 1-4, characterized in that The speaker recognition model is a machine reading comprehension MRC model; The inputting the target dialogue sentence and the speaker-related sentence into a trained speaker recognition model to obtain the speaker information of the target dialogue sentence includes: Combining the target dialogue sentence and a preset question sentence to form a question field, and combining the target dialogue sentence and the speaker-related sentence to form a question-related text field, where the question sentence is used to query the speaker name of the target dialogue sentence; Inputting the question field and the question-related text field into the MRC model to obtain the speaker information of the target dialogue sentence.
7. A method for training a speaker recognition model, characterized in that The method includes: Obtaining sample dialogue sentences in a sample text; Based on the sample speaker list corresponding to the sample text, performing speaker name recognition on the adjacent sentences of the sample dialogue sentence, where the sample speaker list includes at least one speaker name; Based on the speaker name recognition result of the adjacent sentences, determining the sample speaker-related sentence corresponding to the sample dialogue sentence in the adjacent sentences of the sample dialogue sentence; Using the sample dialogue sentence and the sample speaker-related sentence corresponding to the sample dialogue sentence as input sample data to perform training and parameter tuning on a speaker recognition model to be trained, and obtaining a speaker recognition model after parameter tuning; If the training and parameter tuning meet a preset end condition, determining the speaker recognition model after parameter tuning as the trained speaker recognition model; Wherein the determining the sample speaker-related sentence corresponding to the sample dialogue sentence in the adjacent sentences of the sample dialogue sentence based on the speaker name recognition result of the adjacent sentences includes: If only one adjacent sentence of the sample dialogue sentence contains a speaker name, and the adjacent sentence containing the speaker name is not determined as the speaker-related sentence corresponding to other dialogue sentences outside the sample dialogue sentence, then determining the adjacent sentence containing the speaker name as the speaker-related sentence of the sample dialogue sentence; If both adjacent sentences of the sample dialogue sentence contain speaker names, and the adjacent sentence in front of the sample dialogue sentence is not determined as the speaker-related sentence corresponding to other dialogue sentences, then determining the adjacent sentence in front as the speaker-related sentence of the sample dialogue sentence; If both adjacent sentences of the sample dialogue sentence contain speaker names, and the adjacent sentence in front of the sample dialogue sentence has been determined as the speaker-related sentence corresponding to other dialogue sentences, then determining the adjacent sentence behind the sample dialogue sentence as the speaker-related sentence of the sample dialogue sentence; If the sample dialogue sentence has no adjacent sentence containing a speaker name, or the adjacent sentence containing the speaker name of the sample dialogue sentence has been determined as the speaker-related sentence corresponding to other dialogue sentences, then it is determined that the sample dialogue sentence has no speaker-related sentence.
8. An apparatus for identifying a speaker of text, characterized in that, The device includes: An acquisition module, configured to acquire a target dialogue sentence in a target text; An identification module, configured to identify the speaker names of adjacent sentences of the target dialogue sentence based on the speaker list corresponding to the target text, where the speaker list includes at least one speaker name; A determination module, configured to determine the speaker-related sentence of the target dialogue sentence in the adjacent sentences based on the speaker name identification result of the adjacent sentences; An output module, configured to, if the speaker-related sentence of the target dialogue sentence is determined, input the target dialogue sentence and the speaker-related sentence into a trained speaker identification model to obtain the speaker information of the target dialogue sentence, where the speaker information is the speaker name of the target dialogue sentence or an indication information for indicating that there is no corresponding speaker name for the target dialogue sentence; The determination module is configured to: If only one adjacent sentence of the target dialogue sentence contains a speaker name, and the adjacent sentence containing the speaker name has not been determined as the speaker-related sentence corresponding to other dialogue sentences outside the target dialogue sentence, then determine the adjacent sentence containing the speaker name as the speaker-related sentence of the target dialogue sentence; If both adjacent sentences of the target dialogue sentence contain speaker names, and the adjacent sentence in front of the target dialogue sentence has not been determined as the speaker-related sentence corresponding to other dialogue sentences, then determine the adjacent sentence in front as the speaker-related sentence of the target dialogue sentence; If both adjacent sentences of the target dialogue sentence contain speaker names, and the adjacent sentence in front of the target dialogue sentence has been determined as the speaker-related sentence corresponding to other dialogue sentences, then determine the adjacent sentence behind the target dialogue sentence as the speaker-related sentence of the target dialogue sentence; If there is no adjacent sentence containing a speaker name for the target dialogue sentence, or the adjacent sentence containing the speaker name of the target dialogue sentence has been determined as the speaker-related sentence corresponding to other dialogue sentences, then determine that there is no speaker-related sentence for the target dialogue sentence.
9. An apparatus for training a speaker recognition model, characterized in that The apparatus includes: An acquisition module, configured to acquire sample dialogue sentences in the sample text; An identification module, configured to identify the speaker names of adjacent sentences of the sample dialogue sentence based on the sample speaker list corresponding to the sample text, where the sample speaker list includes at least one speaker name; A determination module, configured to determine the sample speaker-related sentence corresponding to the sample dialogue sentence in the adjacent sentences of the sample dialogue sentence based on the speaker name identification result of the adjacent sentences; A training module, configured to use the sample dialogue sentence and the sample speaker-related sentence corresponding to the sample dialogue sentence as input sample data to perform training and parameter adjustment on a speaker identification model to be trained, and obtain a speaker identification model after parameter adjustment; An end module, configured to, if the training and parameter adjustment meet a preset end condition, determine the speaker identification model after parameter adjustment as the trained speaker identification model; The determination module is configured to: If only one of the adjacent sentences of the sample dialogue sentence contains the speaker name, and the adjacent sentence containing the speaker name is not determined as the speaker-related sentence corresponding to other dialogue sentences outside the sample dialogue sentence, then the adjacent sentence containing the speaker name is determined as the speaker-related sentence of the sample dialogue sentence; If both adjacent sentences of the sample dialogue sentence contain the speaker name, and the adjacent sentence in front of the sample dialogue sentence is not determined as the speaker-related sentence corresponding to other dialogue sentences, then the adjacent sentence in front is determined as the speaker-related sentence of the sample dialogue sentence; If both adjacent sentences of the sample dialogue sentence contain the speaker name, and the adjacent sentence in front of the sample dialogue sentence has been determined as the speaker-related sentence corresponding to other dialogue sentences, then the adjacent sentence behind the sample dialogue sentence is determined as the speaker-related sentence of the sample dialogue sentence; If the sample dialogue sentence has no adjacent sentence containing the speaker name, or the adjacent sentence containing the speaker name of the sample dialogue sentence has been determined as the speaker-related sentence corresponding to other dialogue sentences, then it is determined that the sample dialogue sentence has no speaker-related sentence.
10. A computer device, characterized in that, The computer device includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the operations performed by the method according to any one of claims 1 to 7.
11. A computer-readable storage medium, characterized in that, At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the operations performed by the method according to any one of claims 1 to 7.
12. A computer program product, characterized in that, The computer program product includes at least one instruction, and the at least one instruction is loaded and executed by a processor to implement the operations performed by the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Text role labeling method and device, electronic equipment and storage medium
CN112269862A