Method and apparatus for determining speaker role in text
By applying a character prediction model and preset character replacement technology to novels, the speaker's role can be automatically identified, solving the problem of low efficiency in manual character differentiation and achieving efficient character differentiation and automated processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-09
- Publication Date
- 2026-03-27
AI Technical Summary
In long novels, manually distinguishing characters is inefficient, especially for frequently appearing protagonists and infrequently appearing supporting characters.
By acquiring dialogue sentences from the target text, and using a role prediction model and preset character replacement technology, the confidence level of candidate protagonists is automatically identified, and the speaker's role is determined, including the application of tagging and prediction models.
It can efficiently distinguish speaker roles in text without human intervention, improving role differentiation efficiency and enhancing the accuracy and efficiency of automated processing.
Smart Images

Figure CN116011462B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of text recognition, and in particular to a method and device for determining a speaker role in text. BACKGROUND
[0002] Speech synthesis technology is widely used in many fields, and a common application is audio novels. An audio novel is a novel whose dialogues are read by different voices using speech synthesis technology. However, a novel usually has a large number of characters, and some of the supporting characters only appear in a few chapters, or even appear only a few times. Therefore, the same voice can be used to read the dialogues of these supporting characters. However, the main characters appear in most chapters, and different voices can be used to read the dialogues of the main characters to provide a better listening experience for users.
[0003] Currently, technical personnel manually distinguish the characters in a novel into main characters and supporting characters, and manually mark the speaker of each dialogue as a main character or a supporting character.
[0004] However, a novel usually has a large amount of content, and manual character distinction is inefficient. SUMMARY
[0005] Embodiments of the present application provide a method and device for determining a speaker role in text, which can improve the efficiency of character distinction. The technical solution is as follows:
[0006] In a first aspect, a method for determining a speaker role in text is provided, and the method includes:
[0007] obtaining a target dialogue sentence in a target text, wherein the target dialogue sentence is any dialogue sentence in the target text;
[0008] labeling the target dialogue sentence using a first preset character;
[0009] obtaining a plurality of sentences in the context before and after the target dialogue sentence as reference sentences;
[0010] determining the main characters included in the reference sentences as candidate main characters according to a main character list corresponding to the target text, wherein the main character list includes the character names of at least one main character in the target text;
[0011] for each candidate main character, replacing the name of the candidate main character with a second preset character in the reference sentences, inputting the reference sentences after the replacement and the target dialogue sentence after the labeling into a character prediction model as a predicted text corresponding to the candidate main character, and obtaining a confidence degree of the candidate main character as the speaker of the target dialogue sentence.
[0012] According to the confidence that each candidate leading actor is a speaker of the target dialogue sentence, among the candidate leading actors, a candidate leading actor satisfying a preset leading actor condition is determined as a speaker of the target dialogue sentence, and if there is no candidate leading actor satisfying the preset leading actor condition, it is determined that the speaker of the target dialogue sentence is a supporting actor.
[0013] In a possible implementation, the replacing the name of the candidate leading actor with the second preset character in the reference sentence, and the target dialogue sentence after the marking processing, are used as the predicted text corresponding to the candidate leading actor.
[0014] The name of the candidate leading actor is replaced with the second preset character in the reference sentence, and a third preset character is added at a preset position of the reference sentence, and the reference sentence after the replacing and adding processing and the target dialogue sentence after the marking processing are used as the predicted text corresponding to the candidate leading actor.
[0015] In a possible implementation, the confidence that the candidate leading actor is a speaker of the target dialogue sentence is obtained, including:
[0016] The first predicted probability corresponding to the second preset character and the second predicted probability corresponding to the third preset character are obtained, and the first predicted probability is used as the confidence that the candidate leading actor is a speaker of the target dialogue sentence.
[0017] The confidence that each candidate leading actor is a speaker of the target dialogue sentence, among the candidate leading actors, a candidate leading actor satisfying a preset leading actor condition is determined as a speaker of the target dialogue sentence, including:
[0018] For each candidate leading actor, if the confidence that the candidate leading actor is a speaker of the target dialogue sentence is greater than the second predicted probability corresponding to the third preset character in the predicted text corresponding to the candidate leading actor, it is determined that the candidate leading actor satisfies a preset leading actor condition.
[0019] If there are multiple candidate leading actors satisfying the preset leading actor condition, the candidate leading actor corresponding to the maximum confidence is determined as the speaker of the target dialogue sentence.
[0020] In a possible implementation, the confidence that the candidate leading actor is a speaker of the target dialogue sentence is greater than the second predicted probability corresponding to the third preset character in the predicted text corresponding to the candidate leading actor, including:
[0021] If the candidate main character corresponds to a predicted text including N second preset characters, in first prediction probabilities respectively corresponding to the N second preset characters, a target number of first prediction probabilities greater than the second prediction probabilities is determined, and if the target number is greater than N / 2, it is determined that the candidate main character satisfies a preset main character condition, where N is a positive integer greater than 1.
[0022] In a possible implementation, the marking the target dialogue sentence using the first preset character includes:
[0023] The target dialogue sentence is replaced by the first preset character.
[0024] In a possible implementation, the marking the target dialogue sentence using the first preset character includes:
[0025] The quotation marks of the target dialogue sentence are replaced by the first preset character.
[0026] The method further includes:
[0027] The quotation marks of the dialogue sentence in the reference sentence are replaced by the fourth preset character.
[0028] In a possible implementation, the method further includes:
[0029] The occurrence number of the role name of each role in the target text is determined, and the role name with an occurrence number in the top M is added to a main character list corresponding to the target text as a role name of a main character of the target text.
[0030] In a possible implementation, before the obtaining, as reference sentences, a plurality of sentences respectively in the context above and below the target dialogue sentence, the method further includes:
[0031] The target text is segmented according to a sentence ending punctuation;
[0032] For a sentence containing quotation marks obtained by segmentation, the content inside the quotation marks and the content outside the quotation marks are divided into different sentences.
[0033] In a second aspect, a device for determining a speaker role in a text is provided, and the device includes:
[0034] An obtaining module is configured to obtain a target dialogue sentence in a target text, where the target dialogue sentence is any dialogue sentence of the target text.
[0035] A marking module is configured to mark the target dialogue sentence using a first preset character.
[0036] The obtaining module is further configured to obtain, as reference sentences, a plurality of sentences respectively in the context above and below the target dialogue sentence.
[0037] determining a main character included in the reference sentence as a candidate main character according to a main character list corresponding to the target text, wherein the main character list includes a character name of at least one main character in the target text;
[0038] predicting, for each candidate main character, a prediction text corresponding to the candidate main character by replacing a name of the candidate main character with a second preset character in the reference sentence, and inputting the prediction text corresponding to the candidate main character into a character prediction model to obtain a confidence degree of the candidate main character being a speaker of the target dialogue sentence;
[0039] determining, according to the confidence degree of each candidate main character being the speaker of the target dialogue sentence, a candidate main character satisfying a preset main character condition as the speaker of the target dialogue sentence from the candidate main characters, and determining that the speaker of the target dialogue sentence is a supporting character if there is no candidate main character satisfying the preset main character condition.
[0040] In a possible implementation, the prediction module is configured to:
[0041] replace the name of the candidate main character with a second preset character in the reference sentence, and add a third preset character to a preset position of the reference sentence, and predict, as the prediction text corresponding to the candidate main character, the reference sentence after the replacement and addition and the target dialogue sentence after the marking processing.
[0042] In a possible implementation, the prediction module is configured to:
[0043] obtain a first prediction probability corresponding to the second preset character and a second prediction probability corresponding to the third preset character, and take the first prediction probability as the confidence degree of the candidate main character being the speaker of the target dialogue sentence.
[0044] The decision module is configured to:
[0045] for each candidate main character, if the confidence degree of the candidate main character being the speaker of the target dialogue sentence is greater than the second prediction probability corresponding to the third preset character in the prediction text corresponding to the candidate main character, determine that the candidate main character satisfies the preset main character condition.
[0046] if there are multiple candidate main characters satisfying the preset main character condition, determine the candidate main character corresponding to the maximum confidence degree as the speaker of the target dialogue sentence.
[0047] In a possible implementation, the decision module is configured to:
[0048] If the candidate main character corresponds to a predicted text including N second preset characters, in first prediction probabilities corresponding to the N second preset characters respectively, a target number of first prediction probabilities greater than the second prediction probabilities is determined, if the target number is greater than N / 2, it is determined that the candidate main character satisfies a preset main character condition, wherein N is a positive integer greater than 1.
[0049] In a possible implementation, the marking module is configured to:
[0050] replace the target dialogue sentence with a first preset character.
[0051] In a possible implementation, the marking module is configured to:
[0052] replace a quotation mark of the target dialogue sentence with a first preset character;
[0053] The prediction module is further configured to:
[0054] replace a quotation mark of the dialogue sentence in the reference sentence with a fourth preset character.
[0055] In a possible implementation, the determination module is further configured to:
[0056] determine a number of occurrences of a role name of each role in the target text, and add a role name with a number of occurrences in a top M to a main character list corresponding to the target text as a role name of a main character of the target text.
[0057] In a possible implementation, the apparatus further includes a sentence splitting module configured to:
[0058] split the target text according to a sentence ending punctuation;
[0059] for a sentence obtained by splitting and containing a quotation mark, split content inside the quotation mark and content outside the quotation mark into different sentences.
[0060] In a third aspect, an electronic device is provided, the terminal including a processor and a memory, the memory storing at least one instruction, the at least one instruction being loaded and executed by the processor to implement the method for determining a speaker role in a text as described in the first aspect above.
[0061] In a fourth aspect, a computer readable storage medium is provided, the storage medium storing at least one instruction, the at least one instruction being loaded and executed by the processor to implement the method for determining a speaker role in a text as described in the first aspect above.
[0062] In a fifth aspect, a computer program product is provided, which comprises at least one instruction loaded and executed by the processor to implement the method for determining a speaker role in a text according to the first aspect.
[0063] The technical scheme provided by the embodiments of the present application has at least the following beneficial effects:
[0064] In the embodiments of the present application, a target dialogue sentence in a target text is obtained, which is any dialogue sentence in the target text. A plurality of sentences are obtained as reference sentences in the context of the target dialogue sentence. Then, the target dialogue sentence is marked to distinguish it from comparative dialogue sentences in the reference sentences. Next, candidate main characters appearing in the reference sentences are determined according to a main character list of the target text. Then, for each candidate main character, a corresponding predicted text is generated. Specifically, the name of the candidate main character is replaced by a preset character in the reference sentences, and the reference sentences after the replacement and the target dialogue sentence after the marking are taken as the predicted text corresponding to the candidate main character. The predicted text corresponding to the candidate main character is input into a role prediction model to obtain the confidence that the candidate main character is the speaker of the target dialogue sentence. Further, the candidate main character that meets a preset main character condition can be determined as the speaker of the target dialogue sentence from the candidate main characters according to the confidence that each candidate main character is the speaker of the target dialogue sentence. If there is no candidate main character that meets the preset main character condition, it can be determined that the speaker of the target dialogue sentence is a supporting character. In this way, the speaker role of a dialogue sentence in a text can be determined without human intervention, which is more efficient than manual calibration. BRIEF DESCRIPTION OF DRAWINGS
[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0066] Figure 1 is a flowchart of a method for determining a speaker role in a text provided by the embodiments of the present application;
[0067] Figure 2 is a sentence division schematic diagram provided by the embodiments of the present application;
[0068] Figure 3 is a reference sentence schematic diagram provided by the embodiments of the present application;
[0069] Figure 4 is a target dialogue sentence marking schematic diagram provided by the embodiments of the present application;
[0070] Figure 5 is a target dialogue marking schematic diagram provided by an embodiment of the present application;
[0071] Figure 6 is a test text schematic diagram provided by an embodiment of the present application;
[0072] Figure 7 is a test text schematic diagram provided by an embodiment of the present application;
[0073] Figure 8 is a structure schematic diagram of a device for determining a speaker role in a text provided by an embodiment of the present application;
[0074] Figure 9 is a structure schematic diagram of an electronic device provided by an embodiment of the present application;
[0075] Figure 10 is a structure schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0076] In order to make the purpose, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings.
[0077] The embodiments of the present application provide a method for determining a speaker role in a text, which can be applied in the field of audio books and the like. The method can be implemented by an electronic device. The electronic device can be a terminal or a server. The terminal can be a mobile phone, a desktop computer, a tablet computer, a notebook computer and the like. The server can be a single server or a server cluster.
[0078] In the case of being implemented by a terminal, the terminal can be installed with an audio book application program, and a user can select a target audio book to be listened to in the audio book application program. Then, the terminal can obtain a target text of the target audio book, and further determine a speaker role of a dialogue sentence in the text by using the method provided by the embodiments of the present application. Thus, the dialogue sentence of a main character can be read with different tones, and the dialogue sentence of a supporting character can be read with the same tone. The audio book application program can be an application program specially used for providing audio books for users, or an application program with an audio book function.
[0079] In the case of being implemented by the server, the server can pre-acquire the target text to be made into an audiobook, and then can determine the speaker roles of each dialogue sentence in the target text by using the method provided in the embodiments of the present application. Thus, the audiobook of the target text can be generated by reading the dialogue sentences of the main characters with different tones and reading the dialogue sentences of the supporting characters with the same tone. In this way, when the user selects the target audiobook to be listened to through the audiobook application program in the terminal, the server can issue the target audiobook that has been pre-generated to the terminal.
[0080] In the method provided in the embodiments of the present application, a target dialogue sentence in the target text is acquired, which is any dialogue sentence in the target text. In the upper and lower texts of the target dialogue sentence, a plurality of sentences are acquired as reference sentences. Then, the target dialogue sentence is marked so as to be distinguished from the contrast dialogue sentence in the reference sentences. Next, according to the list of main characters of the target text, a candidate main character appearing in the reference sentences is determined. Then, for each candidate main character, a corresponding prediction text is generated. Specifically, the name of the candidate main character is replaced by a preset character in the reference sentences, and the reference sentences after the replacement processing and the target dialogue sentence after the marking processing are taken as the prediction text corresponding to the candidate main character. The prediction text corresponding to the candidate main character is input into the role prediction model, and the confidence that the candidate main character is the speaker of the target dialogue sentence can be obtained. Then, according to the confidence that each candidate main character is the speaker of the target dialogue sentence, the candidate main character that satisfies the preset main character condition is determined as the speaker of the target dialogue sentence from the candidate main characters. If there is no candidate main character that satisfies the preset main character condition, it can be determined that the speaker of the target dialogue sentence is a supporting character. In this way, without human participation, the speaker role of the dialogue sentence in the text can be determined, which is more efficient than manual calibration.
[0081] The method for determining the speaker role in the text provided in the embodiments of the present application will be described below with reference to the accompanying drawings, as shown in FIG. 1, the processing flow of the method can include the following steps: Figure 1
[0082] Step 101, a target dialogue sentence in a target text is acquired, wherein the target dialogue sentence is any dialogue sentence in the target text.
[0083] In implementation, in the case that the method is implemented by the terminal, when the user wants to listen to the target audiobook, the user can open the audiobook application program in the terminal and select the target audiobook in the audiobook application program. Then, the terminal can send a target audiobook download request to the background server of the audiobook application program, and the background server returns the target text of the target audiobook to the terminal after receiving the target audiobook download request. In this way, the terminal can acquire the target text.
[0084] When this method is implemented by a server, the server can obtain the target text that needs to be made into an audiobook via the Internet or local storage devices.
[0085] After obtaining the target text, the terminal and server determine the speaker's role in the same way, which will not be described separately below. For ease of description, the method will be described below using an electronic device as the execution subject; the electronic device can be a terminal or a server.
[0086] After acquiring the target text, the electronic device can first perform data cleaning, removing redundant spaces and punctuation marks. Then, the cleaned text is segmented into sentences. Sentence segmentation can be done as follows: sentences are segmented at the end of punctuation marks; if the punctuation mark is within quotation marks, it is not segmented. For sentences containing quotation marks, the content within the quotation marks and the content outside the quotation marks are separated into different sentences. Ending punctuation marks include periods, ellipses, exclamation marks, and question marks. An example is provided below to illustrate the sentence segmentation process.
[0087] See Figure 2 The sentence is split at the end of the quotation marks. Even if there is end-of-sentence punctuation within the quotation marks, no split is made. For example, Figure 2 In the sentence "Did they get discovered? Could my luck be that bad?", although there is a question mark inside the quotation marks, it is not segmented because it is within the quotation marks. Therefore, "Did they get discovered? Could my luck be that bad?" is treated as a single sentence. Furthermore, for sentences containing quotation marks, the content inside the quotation marks and the content outside the quotation marks are treated as separate sentences. For example, Figure 2 In the passage, Li Si whispers, "Could we have been discovered? Could my luck be that bad?" The phrase "Li Si whispered" outside the quotation marks is treated as a separate sentence. Through clause segmentation, Figure 2 The text can be divided into 14 sentences, namely:
[0088] Sentence 1: Zhang San slowly walked out of the square. His somewhat lonely figure seemed much more relaxed than when he arrived.
[0089] Sentence 2: At this moment, a loud noise suddenly rang out in the square.
[0090] Sentence 3: Upon hearing the sound from the square, Zhang San's steps, which were about to descend the steps, suddenly froze.
[0091] Sentence 4: Zhang San was facing away from the square, tilted his head back and took a long breath, his fist clenched slightly in his sleeve.
[0092] Sentence 5: On the distant mountaintop, Li Si's gaze swept towards the square at that moment.
[0093] Sentence six, the faces of Wang Wu and others on the square are very strange.
[0094] Sentence seven, Li Si whispered
[0095] Sentence eight, "Did you get caught? Can't be so bad luck?".
[0096] Sentence nine, then, the breath in Li Si's body also quietly circulated.
[0097] Sentence ten, on the square, people's eyes once again looked at the lonely figure about to walk down the stairs.
[0098] Sentence eleven, Lao Six looked up at Zhang San's back and said
[0099] Sentence twelve, "Today, I lost."
[0100] Sentence thirteen, "Lao Six, you stand aside for a while."
[0101] Sentence fourteen, the eldest brother waved his hand and said.
[0102] After the target text is divided into sentences, each dialogue sentence in the target text is obtained in turn as a target dialogue sentence. For example, in the above fourteen sentences, sentence eight, sentence twelve, and sentence thirteen are dialogue sentences, and these three sentences are obtained in turn as target dialogue sentences.
[0103] Step 102, in the context above and below the target dialogue sentence, a plurality of sentences are obtained as reference sentences.
[0104] In implementation, the first preset number of sentences closest to the target dialogue sentence are selected in the context above the target dialogue sentence, and the second preset number of sentences closest to the target dialogue sentence are selected in the context below the target dialogue sentence. The selected first preset number of sentences and second preset number of sentences are used as reference sentences corresponding to the target dialogue sentence. The first preset number and the second preset number can be the same or different, for example, the first preset number and the second preset number can both be 5.
[0105] The acquisition of reference sentences is described below through an example.
[0106] Referring to Figure 3 , in combination with Figure 2, assuming the target quoted sentence is sentence eight "Did you get caught? Can't be that unlucky, can you?". The five sentences closest to sentence eight in the context of sentence eight are sentence three, sentence four, sentence five, sentence six and sentence seven. The five sentences closest to sentence eight in the context of sentence eight are sentence nine, sentence ten, sentence eleven, sentence twelve and sentence thirteen. Sentence three, sentence four, sentence five, sentence six, sentence seven, sentence nine, sentence ten, sentence eleven, sentence twelve and sentence thirteen are reference sentences of sentence eight.
[0107] Step 103, marking the target quoted sentence using a first preset character.
[0108] In implementation, the target quoted sentence can be marked to distinguish the target quoted sentence from the quoted sentences in the corresponding reference sentences. There are various marking methods, some of which are listed below for illustration.
[0109] Marking method one, replacing the content in the double quotes of the target quoted sentence with a first preset character.
[0110] Referring to Figure 4 , assuming the target quoted sentence is "Did you get caught? Can't be that unlucky, can you?". The first preset character is [TARGET_QUOTE]. Accordingly, the processing of marking the target quoted sentence is to replace the content in the double quotes of the target quoted sentence with [TARGET_QUOTE].
[0111] Marking method two, replacing the double quotes of the target quoted sentence with a first preset character, and replacing the double quotes of the quoted sentences in the corresponding reference sentences with a second preset character.
[0112] Referring to Figure 5 , assuming the target quoted sentence is "Did you get caught? Can't be that unlucky, can you?". The first preset character is [T] and [ / T], and the second preset character is [Q] and [ / Q]. Accordingly, the processing of marking the target quoted sentence is to replace the left half of the double quotes of the target quoted sentence with [T], replace the right half of the double quotes of the target quoted sentence with [ / T], replace the left half of the double quotes of the quoted sentences in the reference sentences with [Q], and replace the right half of the double quotes of the quoted sentences in the reference sentences with [ / Q].
[0113] Step 104, determining the characters included in the reference sentences as candidate characters according to the list of main characters corresponding to the target text, wherein the list of main characters includes the character names of at least one main character in the target text.
[0114] In implementation, the role names of all roles in the target text are determined through named entity recognition, and the roles with the top M occurrence times in the target text are taken as the main roles of the target text according to the occurrence times of the roles in the target text. The role names of the main roles of the target text are combined to form a main role list of the target text. M is a preset value, which can be configured by a technician according to the actual situation of the target text, for example, M = 30.
[0115] It should be noted here that the determination of the main role list of the target text can be performed after the target text is obtained in step 101, or before step 104.
[0116] The role names of the roles included in the reference sentence corresponding to the target dialogue sentence are obtained, and the main role list is queried to perform role matching. The main role matched in the reference sentence is taken as the candidate main role of the target dialogue sentence.
[0117] In step 105, for each candidate main role, the name of the candidate main role is replaced with a third preset character in the reference sentence, and the replaced reference sentence and the marked target dialogue sentence are taken as the prediction text corresponding to the candidate main role. The prediction text corresponding to the candidate main role is input into the role prediction model to obtain the confidence of the candidate main role as the speaker of the target dialogue sentence.
[0118] In implementation, for each candidate main role of the target dialogue sentence, the prediction text corresponding to the candidate main role is generated. The generation method of the prediction text can be as follows:
[0119] For each candidate main role, the role name of the candidate main role is replaced with a third preset character in the reference sentence, and the replaced reference sentence and the marked target dialogue sentence in step 103 are taken as the prediction text corresponding to the candidate main role. The following will illustrate the processing of the prediction text through an example.
[0120] Referring to Figure 6 , the marked target dialogue sentence and the corresponding reference sentence are shown. It is assumed that the candidate main role is “Li Si”, and an example of generating the prediction text corresponding to Li Si is described. The “Li Si” in the reference sentence is replaced with a third preset character, for example, the third preset character is: [MASK]. In this way, the prediction text corresponding to the candidate main role “Li Si” is obtained.
[0121] For each candidate main character, after obtaining the predicted text of the candidate main character, it can be judged whether the predicted text of the candidate main character is greater than the maximum input character length of the language representation model. If the predicted text of the candidate main character is greater than the maximum input character length of the language representation model, the predicted text is divided into multiple predicted texts smaller than or equal to the maximum input character length in a sliding window manner. The window size and step of the sliding window can be set by the technician according to the actual situation.
[0122] For each candidate main character, the predicted text corresponding to the candidate main character is input into the pre-trained language representation model to obtain the feature vector corresponding to each character in the predicted text. The language representation model can be a BERT (Bidirectional Encoder Representation from Transformers, Bidirectional Encoder Representation from Transformers) model. Here, if the candidate main character corresponds to multiple predicted texts, the multiple predicted texts are input into the language representation model respectively.
[0123] Then, the first feature vector corresponding to the third preset character is input into the full connection layer, and the full connection layer is further used for feature extraction to obtain the second feature vector corresponding to the third preset character. The second feature vector is input into the activation function to obtain the first prediction probability corresponding to the third preset character. The first prediction probability is the confidence degree of the candidate main character as the speaker of the target dialogue sentence.
[0124] The language representation model, the full connection layer, and the activation function form a role prediction model.
[0125] Here, it should be noted that for each candidate main character, if the predicted text corresponding to the candidate main character includes multiple third preset characters (may include multiple third preset characters in one predicted text, or may include multiple third preset characters in multiple predicted texts), the role prediction model can obtain the first prediction probability corresponding to each third preset character. The values of the multiple first prediction probabilities can be different.
[0126] In step 106, according to the confidence degree of each candidate main character as the speaker of the target dialogue sentence, a candidate main character satisfying the preset main character condition is determined from the candidate main characters as the speaker of the target dialogue sentence. If there is no candidate main character satisfying the preset main character condition, the speaker of the target dialogue sentence is indeed a supporting actor.
[0127] In the implementation, for each candidate main character, if only one first prediction probability is obtained by the role prediction model in step 105, it is judged whether the first prediction probability is greater than the preset confidence threshold. If the first prediction probability is greater than the preset confidence threshold, it is determined that the candidate main character satisfies the preset main character condition.
[0128] For each candidate protagonist, if N first predicted probabilities are obtained in step 105 through the role prediction model, where N is a positive integer greater than 1, then the number of first predicted probabilities greater than a preset confidence threshold among the N first predicted probabilities is determined, and this number is denoted as Q. If Q > N / 2, then the candidate protagonist is determined to meet the preset protagonist conditions.
[0129] If only one candidate protagonist meets the preset protagonist conditions, then that candidate protagonist will be the speaker of the target dialogue sentence.
[0130] If multiple candidate protagonists meet the preset protagonist criteria, the confidence levels of each candidate protagonist who meets the preset criteria are compared, and the candidate protagonist with the highest confidence level is selected as the speaker of the target dialogue sentence. Optionally, the candidate protagonist with the highest average confidence level can also be selected as the speaker of the target dialogue sentence.
[0131] In one possible implementation, step 105 above can also be processed as follows:
[0132] The candidate protagonist's name is replaced with a third preset character in the reference sentence, and a fourth preset character is added at a preset position in the reference sentence. The reference sentence with the replacement and addition, along with the target dialogue sentence with the added tags, are used as the predicted text corresponding to the candidate protagonist. The preset position can be the beginning of the first sentence in the reference sentence.
[0133] See Figure 7 This shows the target dialogue sentence and its corresponding reference sentence after the tagging process. Assuming the candidate protagonist is "Li Si," the example of generating the predicted text corresponding to Li Si is used. "Li Si" in the reference sentence is replaced with [MASK]. A fourth preset character is added to the beginning of the first sentence in the reference sentence; for example, the fourth preset character is... <cls>Thus, the prediction text corresponding to the candidate protagonist "Li Si" is obtained.
[0134] For each candidate protagonist, after obtaining the prediction text of the candidate protagonist, it can be judged whether the prediction text of the candidate protagonist is greater than the maximum input character length of the language representation model. If the prediction text of the candidate protagonist is greater than the maximum input character length of the language representation model, the prediction text is divided in a sliding window manner to obtain multiple divided texts, and the fourth preset character is added before the first character of the divided text except the divided text including the fourth preset character, so that multiple prediction texts less than or equal to the maximum input character length (each divided text corresponds to a prediction text) are obtained. The window size and step of the sliding window can be set by the technician according to the actual situation.
[0135] For each candidate protagonist, the prediction text corresponding to the candidate protagonist is input into the pre-trained language representation model to obtain the feature vector corresponding to each character in the prediction text. Here, if the candidate protagonist corresponds to multiple prediction texts, the multiple prediction texts are input into the language representation model respectively.
[0136] Then, the first feature vector corresponding to the third preset character and the third feature vector corresponding to the fourth preset character are input into the full connection layer, and the full connection layer is further used for feature extraction to obtain the second feature vector corresponding to the third preset character and the fourth feature vector corresponding to the fourth preset character. The second feature vector and the fourth feature vector are input into the activation function to obtain the first prediction probability corresponding to the third preset character and the second prediction probability corresponding to the fourth preset character. The first prediction probability is the confidence degree of the candidate protagonist being the speaker of the target dialogue sentence. The language representation model, the full connection layer and the activation function constitute the role prediction model.
[0137] Here, it should be noted that for each candidate protagonist, if the prediction text corresponding to the candidate protagonist includes multiple third preset characters (it may be that a prediction text includes multiple third preset characters, or multiple prediction texts include multiple third preset characters), the role prediction model can obtain multiple first prediction probabilities corresponding to the multiple third preset characters respectively. If the candidate protagonist corresponds to multiple prediction texts, a second prediction probability corresponding to a fourth preset character can be obtained for each prediction text. The values of the multiple first prediction probabilities may be different, and the values of the multiple second prediction probabilities may also be different.
[0138] Correspondingly, in this possible implementation, the processing of step 106 can be as follows:
[0139] For each candidate main character, if there is only one predicted text corresponding to the candidate main character, and only one first prediction probability and one second prediction probability are obtained through the character prediction model, it is determined whether the first prediction probability is greater than the second prediction probability. If the first prediction probability is greater than the second prediction probability, it is determined that the candidate main character meets the preset main character condition.
[0140] For each candidate main character, if there is only one predicted text corresponding to the candidate main character, and A first prediction probabilities and one second prediction probability are obtained through the character prediction model, where A is a positive integer greater than 1. It is determined that the number of first prediction probabilities greater than the second prediction probability in the A first prediction probabilities, and the number is denoted as B here. If B > A / 2, it is determined that the candidate main character meets the preset main character condition.
[0141] For each candidate main character, if there are C predicted texts corresponding to the candidate main character, where C is a positive integer greater than 1, for each of the C predicted texts, it is determined that the number of first prediction probabilities greater than the second prediction probability obtained through the character prediction model. The number determined for each of the C predicted texts is denoted as D1, D2, …, DC respectively. C The total number of first prediction probabilities obtained from the C predicted texts is denoted as D. If (D1+D2+…+DC) > D / 2, it is determined that the candidate main character meets the preset main character condition. C
[0142] Further, if there is only one candidate main character that meets the preset main character condition, the candidate main character is taken as the speaker of the target dialogue sentence.
[0143] If there are multiple candidate main characters that meet the preset main character condition, the confidence corresponding to each candidate main character that meets the preset main character condition is compared, and the candidate main character corresponding to the maximum confidence is taken as the speaker of the target dialogue sentence. Alternatively, the candidate main character corresponding to the maximum average of the corresponding confidence can also be taken as the speaker of the target dialogue sentence.
[0144] In the method provided in the embodiments of the present application, a target dialogue sentence in a target text is acquired, the target dialogue sentence being any dialogue sentence in the target text. A plurality of sentences are acquired as reference sentences in the context above and below the target dialogue sentence. Then, the target dialogue sentence is marked so as to be distinguished from a comparative dialogue sentence in the reference sentences. Next, a candidate main character appearing in the reference sentences is determined according to a main character list of the target text. Then, for each candidate main character, a corresponding predicted text is generated, specifically, the name of the candidate main character is replaced by a preset character in the reference sentences, and the reference sentences after the replacement and the target dialogue sentence after the marking are taken as the predicted text corresponding to the candidate main character. The predicted text corresponding to the candidate main character is input into a character prediction model, and a confidence degree that the candidate main character is a speaker of the target dialogue sentence can be obtained. Furthermore, the candidate main character that satisfies a preset main character condition can be determined as the speaker of the target dialogue sentence from the candidate main characters according to the confidence degree that each candidate main character is the speaker of the target dialogue sentence. If there is no candidate main character that satisfies the preset main character condition, it can be determined that the speaker of the target dialogue sentence is a supporting character. In this way, the speaker character of a dialogue sentence in a text can be determined without manual participation, and the efficiency is higher than manual calibration.
[0145] Based on the same technical concept, the embodiments of the present application also provide a device for determining a speaker character in a text, which is described in detail in the following Figure 8 The device comprises an acquisition module 710, a marking module 720, a determination module 730, a prediction module 740 and a decision module 750, wherein:
[0146] The acquisition module 710 is configured to acquire a target dialogue sentence in a target text, wherein the target dialogue sentence is any dialogue sentence in the target text.
[0147] The marking module 720 is configured to mark the target dialogue sentence using a first preset character.
[0148] The acquisition module 710 is further configured to acquire a plurality of sentences as reference sentences in the context above and below the target dialogue sentence.
[0149] The determination module 730 is configured to determine a main character included in the reference sentences as a candidate main character according to a main character list corresponding to the target text, wherein the main character list includes a character name of at least one main character in the target text.
[0150] The prediction module 740 is configured to, for each candidate leading role, replace the name of the candidate leading role with a second preset character in the reference sentence, input the reference sentence after the replacement and the target dialogue sentence after the marking into a role prediction model as a predicted text corresponding to the candidate leading role, and obtain a confidence degree of the candidate leading role as a speaker of the target dialogue sentence.
[0151] The decision module 750 is configured to determine, from the candidate leading roles, a candidate leading role satisfying a preset leading role condition as a speaker of the target dialogue sentence according to the confidence degree of each candidate leading role as the speaker of the target dialogue sentence, and determine that the speaker of the target dialogue sentence is a supporting role if there is no candidate leading role satisfying the preset leading role condition.
[0152] In a possible implementation, the prediction module 740 is configured to:
[0153] The name of the candidate leading role is replaced with a second preset character in the reference sentence, and a third preset character is added at a preset position of the reference sentence, and the reference sentence after the replacement and the addition and the target dialogue sentence after the marking are input into a role prediction model as a predicted text corresponding to the candidate leading role.
[0154] In a possible implementation, the prediction module 740 is configured to:
[0155] The first prediction probability corresponding to the second preset character and the second prediction probability corresponding to the third preset character are obtained, and the first prediction probability is taken as the confidence degree of the candidate leading role as the speaker of the target dialogue sentence.
[0156] The decision module 750 is configured to:
[0157] For each candidate leading role, if the confidence degree of the candidate leading role as the speaker of the target dialogue sentence is greater than the second prediction probability corresponding to the third preset character in the predicted text corresponding to the candidate leading role, it is determined that the candidate leading role satisfies a preset leading role condition.
[0158] If there are multiple candidate leading roles satisfying the preset leading role condition, the candidate leading role corresponding to the maximum confidence degree is taken as the speaker of the target dialogue sentence.
[0159] In a possible implementation, the decision module 750 is configured to:
[0160] If the candidate main character corresponds to a predicted text including N second preset characters, in first prediction probabilities corresponding to the N second preset characters respectively, a target number of first prediction probabilities greater than the second prediction probabilities is determined, if the target number is greater than N / 2, it is determined that the candidate main character satisfies a preset main character condition, wherein N is a positive integer greater than 1.
[0161] In a possible implementation, the marking module 720 is configured to:
[0162] replace the target dialogue sentence with a first preset character.
[0163] In a possible implementation, the marking module 720 is configured to:
[0164] replace a quotation mark of the target dialogue sentence with a first preset character;
[0165] The prediction module 740 is further configured to:
[0166] replace a quotation mark of a dialogue sentence in the reference sentence with a fourth preset character.
[0167] In a possible implementation, the determination module 730 is further configured to:
[0168] determine a number of occurrences of a character name of each character in the target text, and add a character name with a number of occurrences in the top M to a main character list corresponding to the target text as a character name of a main character of the target text.
[0169] In a possible implementation, the apparatus further includes a sentence splitting module configured to:
[0170] split the target text according to a trailing punctuation mark;
[0171] for a split sentence including a quotation mark, split content inside the quotation mark and content outside the quotation mark into different sentences.
[0172] As to the apparatus in the above embodiments, specific manners in which various modules perform operations have been described in detail in the embodiments of the method, and will not be described in detail here.
[0173] In the method provided in the embodiments of the present application, a target dialogue sentence in a target text is acquired, the target dialogue sentence being any dialogue sentence in the target text. A plurality of sentences are acquired as reference sentences in the context of the target dialogue sentence. Then, the target dialogue sentence is marked to distinguish the target dialogue sentence from a comparative dialogue sentence in the reference sentences. Next, a candidate main character appearing in the reference sentences is determined according to a main character list of the target text. Then, for each candidate main character, a corresponding predicted text is generated, specifically, the name of the candidate main character is replaced by a preset character in the reference sentences, and the reference sentences after the replacement and the target dialogue sentence after the marking are taken as the predicted text corresponding to the candidate main character. The predicted text corresponding to the candidate main character is input into a character prediction model, and a confidence degree that the candidate main character is a speaker of the target dialogue sentence can be obtained. Furthermore, the candidate main character that satisfies a preset main character condition can be determined as the speaker of the target dialogue sentence from the candidate main characters according to the confidence degree that each candidate main character is the speaker of the target dialogue sentence. If there is no candidate main character that satisfies the preset main character condition, it can be determined that the speaker of the target dialogue sentence is a supporting character. In this way, the speaker character of a dialogue sentence in a text can be determined without manual participation, and the efficiency is higher than manual calibration.
[0174] It should be noted that the device for determining a speaker character in a text provided in the above embodiments is only used as an example to illustrate the division of the above functional modules, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the electronic device is divided into different functional modules to complete all or part of the above described functions. In addition, the device for determining a speaker character in a text and the method for determining a speaker character in a text provided in the above embodiments belong to the same concept, and the specific implementation process is described in the method embodiments, which will not be described here.
[0175] Figure 9 A structural block diagram of an electronic device 800 provided in an example embodiment of the present application is shown. The electronic device 800 can be a portable mobile terminal, such as a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a notebook computer or a desktop computer. The electronic device 800 can also be referred to as a user device, a portable terminal, a laptop terminal, a desktop terminal, and other names.
[0176] Generally, the electronic device 800 includes a processor 801 and a memory 802.
[0177] The processor 801 can include one or more processing cores, such as a 4-core processor, an 8-core processor, and the like. The processor 801 can be implemented in at least one of a hardware form of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), a PLA (Programmable Logic Array). The processor 801 can also include a main processor and a coprocessor, the main processor being a processor for processing data in an awake state, also known as a CPU (Central Processing Unit), and the coprocessor being a low-power processor for processing data in a standby state. In some embodiments, the processor 801 can be integrated with a GPU (Graphics Processing Unit) for rendering and drawing content required to be displayed by the display screen. In some embodiments, the processor 801 can further include an AI (Artificial Intelligence) processor for processing computing operations related to machine learning.
[0178] The memory 802 can include one or more computer-readable storage media that can be non-transitory. The memory 802 can also include a high-speed random access memory, and a nonvolatile memory such as one or more disk storage devices, flash storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 802 is used to store at least one instruction for being executed by the processor 801 to implement the method for determining the role of a speaker in text provided by the method embodiments in the present application.
[0179] In some embodiments, the electronic device 800 can also optionally include a peripheral device interface 803 and at least one peripheral device. The processor 801, the memory 802, and the peripheral device interface 803 can be connected through a bus or a signal line. Each peripheral device can be connected to the peripheral device interface 803 through a bus, a signal line, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 804, a display screen 805, a camera assembly 806, an audio circuit 807, a positioning assembly 808, and a power supply 809.
[0180] The peripheral interface 803 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 801 and the memory 802. In some embodiments, the processor 801, the memory 802 and the peripheral interface 803 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 801, the memory 802 and the peripheral interface 803 can be implemented on a separate chip or circuit board, and the present embodiments are not limited in this regard.
[0181] The radio frequency circuit 804 is used to receive and send RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 804 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 804 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the radio frequency circuit 804 includes an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and the like. The radio frequency circuit 804 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to the World Wide Web, a metropolitan area network, an intranet, various generations of mobile communication networks (2G, 3G, 4G and 5G), a wireless local area network and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 804 can also include NFC (Near Field Communication) related circuitry, and the present application is not limited in this regard.
[0182] The display screen 805 is configured to display a UI (User Interface). The UI can include graphics, text, icons, video, and any combination thereof. When the display screen 805 is a touch display screen, the display screen 805 is further configured to capture touch signals on or above the surface of the display screen 805. The touch signals can be input to the processor 801 as control signals for processing. In this case, the display screen 805 can also be configured to provide virtual buttons and / or virtual keyboard, also known as soft buttons and / or soft keyboard. In some embodiments, the display screen 805 can be one, disposed on the front panel of the electronic device 800; in other embodiments, the display screen 805 can be at least two, respectively disposed on different surfaces of the electronic device 800 or in a folding design; in other embodiments, the display screen 805 can be a flexible display screen, disposed on a curved surface or a folding surface of the electronic device 800. Even, the display screen 805 can also be disposed in an irregular shape, i.e., a special-shaped screen. The display screen 805 can be made of LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), etc.
[0183] The camera assembly 806 is configured to capture images or videos. Optionally, the camera assembly 806 includes a front camera and a rear camera. Typically, the front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, the rear camera is at least two, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, to realize the background blur function by fusing the main camera and the depth-of-field camera, the panoramic shooting and VR (Virtual Reality) shooting function by fusing the main camera and the wide-angle camera, or other fusion shooting functions. In some embodiments, the camera assembly 806 can further include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. The dual-color temperature flash refers to the combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.
[0184] The audio circuit 807 can include a microphone and a speaker. The microphone is used to collect sound waves of a user and an environment, and convert the sound waves into an electrical signal input to the processor 801 for processing, or input to the radio frequency circuit 804 to realize voice communication. For the purpose of stereo sound collection or noise reduction, the microphone can be multiple, respectively arranged at different parts of the electronic device 800. The microphone can also be an array microphone or an omnidirectional collection type microphone. The speaker is used to convert an electrical signal from the processor 801 or the radio frequency circuit 804 into sound waves. The speaker can be a conventional diaphragm speaker, or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, not only can it convert an electrical signal into a sound wave audible to humans, but also can convert an electrical signal into an inaudible sound wave to humans for ranging purposes, etc. In some embodiments, the audio circuit 807 can also include a headphone jack.
[0185] The positioning component 808 is used to position the current geographic location of the electronic device 800 to realize navigation or LBS (Location Based Service). The positioning component 808 can be a positioning component based on the U.S. GPS (Global Positioning System), the Chinese Beidou system or the Russian Galileo system.
[0186] The power supply 809 is used to supply power to various components in the electronic device 800. The power supply 809 can be alternating current, direct current, disposable batteries or rechargeable batteries. When the power supply 809 includes rechargeable batteries, the rechargeable batteries can be wired charging batteries or wireless charging batteries. The wired charging battery is a battery charged through a wired line, and the wireless charging battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0187] In some embodiments, the electronic device 800 further includes one or more sensors 810. The one or more sensors 810 include, but are not limited to, an acceleration sensor 811, a gyroscope sensor 812, a pressure sensor 813, a fingerprint sensor 814, an optical sensor 815 and a proximity sensor 816.
[0188] The acceleration sensor 811 can detect the acceleration magnitude in three coordinate axes of the coordinate system established by the electronic device 800. For example, the acceleration sensor 811 can be used to detect the components of gravitational acceleration in three coordinate axes. The processor 801 can control the display screen 805 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 811. The acceleration sensor 811 can also be used for gaming or user motion data collection.
[0189] The gyroscope sensor 812 can detect the body direction and rotation angle of the electronic device 800, and can collect 3D motions of the user with respect to the electronic device 800 in cooperation with the acceleration sensor 811. The processor 801 can implement the following functions according to the data collected by the gyroscope sensor 812: motion sensing (e.g., changing a UI according to a tilt operation of the user), image stabilization during shooting, game control, and inertial navigation.
[0190] The pressure sensor 813 can be disposed on the side frame of the electronic device 800 and / or the lower layer of the display 805. When the pressure sensor 813 is disposed on the side frame of the electronic device 800, the grip signal of the user with respect to the electronic device 800 can be detected, and the left / right hand recognition or shortcut operation can be performed by the processor 801 according to the grip signal collected by the pressure sensor 813. When the pressure sensor 813 is disposed on the lower layer of the display 805, the operable control on the UI interface can be controlled by the processor 801 according to the pressure operation of the user with respect to the display 805. The operable control includes at least one of a button control, a scroll bar control, an icon control, and a menu control.
[0191] The fingerprint sensor 814 is used to collect the fingerprint of the user, and the identity of the user can be recognized by the processor 801 according to the fingerprint collected by the fingerprint sensor 814, or by the fingerprint sensor 814 according to the collected fingerprint. When the identity of the user is recognized as a trusted identity, the processor 801 authorizes the user to perform a related sensitive operation, which includes unlocking the screen, viewing encrypted information, downloading software, payment, and changing settings, etc. The fingerprint sensor 814 can be disposed on the front, back, or side of the electronic device 800. When the physical button or the manufacturer's logo is disposed on the electronic device 800, the fingerprint sensor 814 can be integrated with the physical button or the manufacturer's logo.
[0192] The optical sensor 815 is used to collect the ambient light intensity. In one embodiment, the processor 801 can control the display brightness of the display 805 according to the ambient light intensity collected by the optical sensor 815. Specifically, when the ambient light intensity is high, the display brightness of the display 805 is increased, and when the ambient light intensity is low, the display brightness of the display 805 is decreased. In another embodiment, the processor 801 can also dynamically adjust the shooting parameters of the camera assembly 806 according to the ambient light intensity collected by the optical sensor 815.
[0193] The proximity sensor 816, also referred to as a distance sensor, is usually arranged on the front panel of the electronic device 800. The proximity sensor 816 is used to collect the distance between the user and the front of the electronic device 800. In an embodiment, when the proximity sensor 816 detects that the distance between the user and the front of the electronic device 800 gradually decreases, the display screen 805 is switched from the bright screen state to the screen-off state under the control of the processor 801; when the proximity sensor 816 detects that the distance between the user and the front of the electronic device 800 gradually increases, the display screen 805 is switched from the screen-off state to the bright screen state under the control of the processor 801.
[0194] Those skilled in the art can understand that the structure shown in the foregoing embodiments does not constitute a limitation on the electronic device 800, and the electronic device 800 can include more or fewer components than those shown in the figure, or combine certain components, or adopt a different component arrangement. Figure 8
[0195] In an exemplary embodiment, a computer readable storage medium is also provided, and the computer readable storage medium stores at least one instruction. The at least one instruction is loaded and executed by a processor to implement the method for determining the role of a speaker in text in the above embodiments. For example, the computer readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, and the like.
[0196] In an exemplary embodiment, a computer program product is also provided, and the computer program product includes at least one instruction. The at least one instruction is loaded and executed by a processor to implement the method for determining the role of a speaker in text.
[0197] Figure 10 FIG. 1 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. The electronic device 900 can be a computer device, which can have great differences due to different configurations or performances, and can include one or more processors (CPU) 901 and one or more memories 902. The memory 902 stores at least one instruction, and the at least one instruction is loaded and executed by the processor 901 to implement the method for determining the role of a speaker in text.
[0198] Those skilled in the art can understand that all or part of the steps of the above-mentioned embodiments can be completed by hardware, or by program instructing relevant hardware to complete, and the program can be stored in a computer readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.
[0199] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0200] It should be noted that the information (including but not limited to user equipment information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals (including but not limited to signals transmitted between user terminals and other devices, etc.) involved in the present application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions. For example, the audiobooks, target texts, etc. involved in the present application are obtained under full authorization.< / cls>
Claims
1. A method for determining the speaker role in a text, characterized in that, The method includes: Obtain the target dialogue sentence in the target text, wherein the target dialogue sentence is any dialogue sentence in the target text; The target dialogue sentence is marked using a first preset character; Multiple sentences are obtained from the context of the target dialogue sentence, both before and after it, as reference sentences. Based on the list of protagonists corresponding to the target text, the protagonists included in the reference sentence are determined as candidate protagonists, wherein the list of protagonists includes the character names of at least one protagonist in the target text; For each candidate protagonist, the name of the candidate protagonist is replaced with a second preset character in the reference sentence. The reference sentence after replacement and the target dialogue sentence after marking are used as the predicted text corresponding to the candidate protagonist. The predicted text corresponding to the candidate protagonist is input into the role prediction model to obtain the confidence that the candidate protagonist is the speaker of the target dialogue sentence. Based on the confidence level of each candidate protagonist as the speaker of the target dialogue, among the candidate protagonists, those who meet the preset protagonist conditions are determined as the speaker of the target dialogue. If there are no candidate protagonists who meet the preset protagonist conditions, then the speaker of the target dialogue is determined to be a supporting character.
2. The method according to claim 1, characterized in that, The step of replacing the candidate protagonist's name with a second preset character in the reference sentence, and using the replaced reference sentence and the marked target dialogue sentence as the predicted text corresponding to the candidate protagonist, includes: The name of the candidate protagonist is replaced with a second preset character in the reference sentence, and a third preset character is added at a preset position in the reference sentence. The reference sentence after replacement and addition, as well as the target dialogue sentence after marking, are used as the predicted text corresponding to the candidate protagonist.
3. The method according to claim 2, characterized in that, The confidence level obtained for the candidate protagonist to be the speaker of the target dialogue includes: The first prediction probability corresponding to the second preset character and the second prediction probability corresponding to the third preset character are obtained, and the first prediction probability is used as the confidence level that the candidate protagonist is the speaker of the target dialogue sentence. The step of determining, based on the confidence level of each candidate protagonist as the speaker of the target dialogue, a candidate protagonist who meets the preset protagonist conditions and is thus selected as the speaker of the target dialogue includes: For each candidate protagonist, if the confidence that the candidate protagonist is the speaker of the target dialogue sentence is greater than the second prediction probability corresponding to the third preset character in the prediction text corresponding to the candidate protagonist, then the candidate protagonist is determined to meet the preset protagonist condition. If multiple candidate protagonists meet the preset protagonist conditions, the candidate protagonist with the highest confidence level will be selected as the speaker of the target dialogue sentence.
4. The method according to claim 3, characterized in that, If the confidence level of the candidate protagonist being the speaker of the target dialogue is greater than the second prediction probability corresponding to the third preset character in the predicted text of the candidate protagonist, then determining that the candidate protagonist meets the preset protagonist condition includes: If the predicted text corresponding to the candidate protagonist includes N second preset characters, then among the first prediction probabilities corresponding to the N second preset characters, determine the number of targets with a first prediction probability greater than the second prediction probability. If the number of targets is greater than N / 2, then determine that the candidate protagonist meets the preset protagonist condition, where N is a positive integer greater than 1.
5. The method according to claim 1, characterized in that, The step of marking the target dialogue sentence using a first preset character includes: Replace the target dialogue with a first preset character.
6. The method according to claim 1, characterized in that, The step of marking the target dialogue sentence using a first preset character includes: Replace the quotation marks in the target dialogue with a first preset character; The method further includes: Replace the quotation marks in the dialogue in the reference sentence with a fourth preset character, wherein the first preset character and the fourth preset character are different.
7. The method according to any one of claims 1-6, characterized in that, The method further includes: Determine the frequency of each character's name in the target text, and select the character names with the highest frequency of occurrence (M) as the main character names of the target text, adding them to the main character list corresponding to the target text.
8. The method according to any one of claims 1-6, characterized in that, Before obtaining multiple sentences from the context of the target dialogue sentence as reference sentences, the method further includes: The target text is segmented into sentences based on the punctuation at the end of each sentence; For sentences containing quotation marks obtained from clause segmentation, the content inside the quotation marks and the content outside the quotation marks are separated into different sentences.
9. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing at least one instruction, which is loaded and executed by the processor to implement the method for determining the speaker role in text as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, which is loaded and executed by a processor to implement the method for determining the speaker role in text as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Text data management method and system
CN110728132A
Audio generation method of dialogue novel, electronic equipment and storage medium
CN114023358A