Method and device for analyzing data of characters in literary works, and electronic device
By acquiring character data from literary works and using a pre-defined personality polarity analysis model to calculate polarity probability values, the problem of time-consuming and inaccurate character data analysis in literary works has been solved, achieving fast and accurate character personality analysis.
Patent Information
- Application Number
- CN202111162882.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-30
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2041-09-30
AI Technical Summary
In existing technologies, data analysis of characters in literary works is too time-consuming and the accuracy of the analysis results is poor.
By acquiring corpus related to the target characters in the target literary works, the polarity probability value of words is calculated using a preset personality polarity analysis model, the polarity value of personality under each personality dimension is determined, and the final polarity value is output.
It enables rapid and accurate character personality analysis, overcomes data volume limitations, and improves the accuracy and efficiency of analysis results.
Smart Images

Figure CN113886583B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and in particular to a data analysis method and device for a character in a literary work, and an electronic device. BACKGROUND
[0002] A script is one of the necessary tools for stage performance or filming, and is a reference language for dialogue between characters in a play. The script is mainly composed of lines and stage directions. That is, the script not only records the dialogue of the character, but also records the description of the behavior of the character, thereby interpreting the detailed situation of the character.
[0003] Generally, the script records a large amount of information, and how to quickly and accurately analyze the data of the script character has always been a big problem. At present, the data analysis of the script character can only be realized by manual reading and understanding, which consumes a lot of time and energy. Moreover, due to the limited memory, all relevant information cannot be remembered during the analysis process, thereby resulting in poor accuracy of the analysis result. SUMMARY
[0004] In view of the above problems, the embodiments of the present application provide a data analysis method and device for a character in a literary work, and an electronic device, to solve the problem in the prior art that the data analysis process of the character in the literary work is too time-consuming and the accuracy of the analysis result is poor.
[0005] In a first aspect of the present application, a data analysis method for a character in a literary work is provided, and the method comprises:
[0006] obtaining a first target corpus in a target literary work, wherein the first target corpus is a corpus associated with a target character in the target literary work;
[0007] determining a plurality of polarity probability values of a word in the first target corpus according to the first target corpus and a preset personality polarity analysis model, wherein the plurality of polarity probability values are probability values of the word representing different personality polarities under each personality dimension;
[0008] determining a polarity value for indicating a polarity degree of the personality of the target character under each personality dimension according to the plurality of polarity probability values of the word in the first target corpus;
[0009] outputting the target character and the polarity value of the target character under each personality dimension.
[0010] Optionally, each personality dimension includes a positive personality polarity and a negative personality polarity in opposite directions.
[0011] determining, according to the plurality of polarity probability values of the words in the first target corpus, a polarity value of the character of the target role in each character dimension for indicating a polarity degree, comprises:
[0012] respectively for each character dimension, determining a case that the character of the target role is biased to the positive character polarity in the character dimension based on the polarity probability value of each word in the first target corpus representing the positive character polarity in the character dimension;
[0013] respectively for each character dimension, determining a case that the character of the target role is biased to the negative character polarity in the character dimension based on the polarity probability value of each word in the first target corpus representing the negative character polarity in the character dimension;
[0014] respectively for each character dimension, determining the polarity value of the character of the target role in the character dimension based on the case that the character of the target role is biased to the positive character polarity in the character dimension and the case that the character of the target role is biased to the negative character polarity in the character dimension.
[0015] Optionally, the determining the case that the character of the target role is biased to the positive character polarity in the character dimension based on the polarity probability value of each word in the first target corpus representing the positive character polarity in the character dimension, comprises:
[0016] calculating a Bayesian mean of the polarity probability value of each word in the first target corpus representing the positive character polarity in the character dimension based on a Bayesian average algorithm;
[0017] determining the Bayesian mean as a bias value of the character of the target role to the positive character polarity in the character dimension;
[0018] Correspondingly, the determining the case that the character of the target role is biased to the negative character polarity in the character dimension based on the polarity probability value of each word in the first target corpus representing the negative character polarity in the character dimension, comprises:
[0019] calculating a Bayesian mean of the polarity probability value of each word in the first target corpus representing the negative character polarity in the character dimension based on a Bayesian average algorithm;
[0020] determining the Bayesian mean as a bias value of the character of the target role to the negative character polarity in the character dimension.
[0021] Optionally, the determining the polarity value of the character of the target role in the character dimension based on the case that the character of the target role is biased towards the positive character polarity and the case that the character of the target role is biased towards the negative character polarity in the character dimension comprises:
[0022] determining an initial polarity value based on the case that the character of the target role is biased towards the positive character polarity and the case that the character of the target role is biased towards the negative character polarity in the character dimension;
[0023] determining an average polarity value based on the case that the characters of the plurality of roles satisfying the preset condition are biased towards the positive character polarity and the case that the characters of the plurality of roles satisfying the preset condition are biased towards the negative character polarity in the character dimension;
[0024] determining the polarity value of the character of the target role in the character dimension based on the initial polarity value, the average polarity value, and a preset character polarity distribution of roles in different character dimensions in literary works.
[0025] Optionally, the plurality of roles satisfying the preset condition comprises all roles in the target literary work whose number of description sentences is less than a preset threshold.
[0026] Optionally, the determining the plurality of polarity probability values of the words in the first target corpus based on the first target corpus and a preset character polarity analysis model comprises:
[0027] inputting each word in the first target corpus into the preset character polarity analysis model respectively to obtain an initial probability value of each word representing different character polarities in each character dimension;
[0028] determining a weight of the word based on the influence of different sentence types on the description of the character and the sentence type of the sentence to which the word belongs;
[0029] multiplying the initial probability value of each word by the weight of the word to obtain a probability value of the word representing different character polarities in each character dimension.
[0030] Optionally, after the determining the polarity value of the character of the target role in each character dimension for indicating the degree of polarity based on the plurality of polarity probability values of the words in the first target corpus, the method further comprises:
[0031] determining a polarity level of the character of the target role in each character dimension based on a preset numerical interval corresponding to each polarity level in each character dimension;
[0032] outputting the polarity level of the character of the target role in each character dimension.
[0033] In a second aspect of the present application, a data analysis device for a character in a literary work is provided, and the device comprises:
[0034] an acquisition module configured to acquire a first target corpus in a target literary work, wherein the first target corpus is a corpus associated with a target character in the target literary work;
[0035] a first determination module configured to determine a plurality of polarity probability values of a word in the first target corpus according to the first target corpus and a preset personality polarity analysis model, wherein the plurality of polarity probability values are probability values of the word representing different personality polarities in each personality dimension;
[0036] a second determination module configured to determine a personality value of the target character in each personality dimension for indicating a polarity degree according to the plurality of polarity probability values of the word in the first target corpus;
[0037] an output module configured to output the target character and the personality value of the target character in each personality dimension.
[0038] In a third aspect of the present application, an electronic device is provided, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory complete mutual communication through the communication bus;
[0039] the memory is configured to store a computer program;
[0040] the processor is configured to execute the program stored on the memory, so as to realize the steps of the data analysis method for a character in a literary work.
[0041] In a fourth aspect of the present application, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program, wherein the computer program is executed by a processor to realize the steps of the data analysis method for a character in a literary work according to any one of the first aspect.
[0042] In a fifth aspect of the present application, a computer program product containing instructions is provided, and when the computer program product is run on a computer, the computer is caused to execute the data analysis method for a character in a literary work.
[0043] Compared with the prior art, the present application has the following advantages:
[0044] The data analysis method of a character in a literary work provided by the present application can obtain a first target corpus in a target literary work, wherein the first target corpus is a corpus associated with a target character in the target literary work. Compared with the manual data analysis method, the limitation of the data volume of the first target corpus can be broken through, the accuracy of the data analysis result can be improved by enriching the first target corpus. According to the first target corpus and a preset personality polarity analysis model, a plurality of polarity probability values of a word in the first target corpus are determined, wherein the plurality of polarity probability values are probability values of the word representing different personality polarities in each personality dimension. The probability values of each word representing different personality polarities in each personality dimension are determined in units of words. According to the plurality of polarity probability values of the word in the first target corpus, a polarity value indicating the degree of polarity of the personality of the target character in each personality dimension is determined. The polarity value calculated from the probability value of the word representing the personality polarity can realize the quantification of the degree of polarity. Through the polarity value, the personality characteristics of the target character are more finely shown. Finally, by outputting the target character and the polarity value of the target character in each personality dimension, the data analysis result of the target character can be provided to the demander, so that the demander can quickly and accurately understand the target character according to the data analysis result, and then better carry out subsequent work. The personality characteristics of the target character are more finely shown by the calculated polarity value, the analysis result of the target character is more accurate, the whole process is completed in an automatic form, the time consumption is short, and the demander can quickly obtain the analysis result of the target character. BRIEF DESCRIPTION OF DRAWINGS
[0045] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0046] Figure 1 The step flow chart of the data analysis method of a character in a literary work provided by the present application is shown in the following.
[0047] Figure 2 The polar line graph of the relationship between the polarity value and the number of characters based on the movie script drawn by the present application is shown in the following.
[0048] Figure 3 The polar line graph of the relationship between the polarity value and the number of characters based on the television script drawn by the present application is shown in the following.
[0049] Figure 4 The column chart of the relationship between the polarity level and the number of characters based on the movie script drawn by the present application is shown in the following.
[0050] Figure 5 A column chart of a relationship between a polarity level and a number of characters drawn based on a script of a TV series is provided for an embodiment of the present application.
[0051] Figure 6 An actual application schematic diagram of the data analysis method of characters in a literary work is provided for an embodiment of the present application.
[0052] Figure 7 A structural block diagram of the data analysis device of characters in a literary work is provided for an embodiment of the present application.
[0053] Figure 8 A structural block diagram of an electronic device is provided for an embodiment of the present application. DETAILED DESCRIPTION
[0054] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0055] It should be understood that the terms “one embodiment” or “an embodiment” mentioned throughout the specification mean that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, “in one embodiment” or “in an embodiment” appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in any suitable manner in one or more embodiments.
[0056] In various embodiments of the present application, it should be understood that the size of the serial number of each process does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0057] Referring to Figure 1 The data analysis method of characters in a literary work provided by the embodiments of the present application comprises:
[0058] Step 101: obtaining a first target corpus in a target literary work.
[0059] It should be noted that the target literary work is any literary work, which is the product of literary creation, such as a novel, a script, etc. Here, the target literary work can be an electronic version file, and when data analysis of the characters in the literary work is performed by an electronic device, the target literary work can be uploaded to the electronic device in advance, and then the data analysis process is performed on the electronic device. The first target corpus is the corpus associated with the target character in the target literary work. Here, the corpus associated with the target character can be understood as the dialogue of the target character and / or the sentence describing the behavior of the target character. The behavior of the target character includes the physical behavior and psychological activity of the target character. For example, the target character is the male lead in the target literary work, and the first target corpus includes all the dialogue of the male lead and / or all the sentences describing the behavior of the male lead. The target character can be any number of virtual character roles in the target literary work. When the target character includes multiple characters in the target literary work, the first target corpus includes the corpus associated with each of the multiple characters. Here, a correspondence is established between the character and the corpus associated with the character, and for each character, its corresponding corpus is used for data analysis.
[0060] Preferably, the target character can be determined according to user input. For example, the user inputs a character keyword according to the demand, matches all the characters in the target literary work according to the character keyword, and determines the matched characters as the target character, but not limited to this. All characters in the target literary work can also be automatically determined as the target character, and then data analysis is performed on all characters. When the first target corpus is obtained, the corpus associated with the character name can be screened in the target literary work according to the character name of the target character, but not limited to this.
[0061] Step 102: determining a plurality of polarity probability values of the words in the first target corpus according to the first target corpus and a preset personality polarity analysis model.
[0062] It should be noted that each personality has a plurality of different personality dimensions, and a comprehensive evaluation of the personality is achieved through different personality dimensions. For example, in the Big Five Personality Model theory, the personality of a person is divided into five different personality dimensions: neuroticism, extroversion, openness, agreeableness, and conscientiousness. It can be understood that each personality dimension corresponds to two different personality polarities, which respectively represent two different extreme performances in the same personality dimension. Here, the plurality of polarity probability values are the probability values of the words representing different personality polarities in each personality dimension. Continuing to take the Big Five Personality Model theory as an example, the two different personality polarities of neuroticism are rationality and neuroticism, the two different personality polarities of extroversion are introversion and extroversion, the two different personality polarities of openness are conservatism and openness, the two different personality polarities of agreeableness are coldness and agreeableness, and the two different personality polarities of conscientiousness are laziness and conscientiousness.
[0063] The preset personality polarity analysis model is used to determine the probability value of the input word representing each personality polarity in a plurality of different personality dimensions. Preferably, each personality dimension corresponds to a different personality polarity analysis model, and each personality polarity analysis model is used to determine the probability value of the input word representing each personality polarity in its corresponding personality dimension. The number of personality dimensions involved here can be determined as needed, for example, it can be five personality dimensions in the Big Five Personality Model theory, and five different personality polarity analysis models are set in advance to process the words in the first target corpus. Specifically, a first personality polarity analysis model is created for agreeableness to determine the probability value of the input word representing coldness and the probability value of the input word representing agreeableness. Similarly, a second personality polarity analysis model is created for neuroticism to determine the probability value of the input word representing rationality and the probability value of the input word representing neuroticism. A third personality polarity analysis model is created for openness to determine the probability value of the input word representing conservatism and the probability value of the input word representing openness. A fourth personality polarity analysis model is created for conscientiousness to determine the probability value of the input word representing laziness and the probability value of the input word representing conscientiousness. A fifth personality polarity analysis model is created for extraversion to determine the probability value of the input word representing introversion and the probability value of the input word representing extraversion. The preset personality polarity analysis model refers to the first personality polarity model, the second personality polarity model, the third personality polarity model, the fourth personality polarity model, and the fifth personality polarity model. Here, the first target corpus needs to be processed into words, and then the words are input into the preset personality polarity analysis model. Here, the sentences in the first target corpus can be processed into a large number of words in the form of word segmentation.
[0064] The preset personality polarity analysis model can use a binary classifier that outputs two probability values, or a ternary classifier that outputs three probability values, considering that some words are not related to personality. The three probability values output by the ternary classifier are: the first probability value representing the positive personality polarity, the second probability value representing the negative personality polarity, and the third probability value representing no personality.
[0065] To ensure the accuracy of the output results of the preset personality polarity analysis model, it needs to be trained. Here, taking the preset personality polarity analysis model as a ternary classifier as an example, the training process is described.
[0066] Step 1: Manually list representative words for 10 personality polarities (cold, affinity, rational, neurotic, conservative, development, lazy, responsible, introverted, and extroverted) in the 5 major personalities, with about 200 representative words for each personality polarity, such as: impulsive, angry, get out of the way, hostile, etc. for the positive personality polarity of neurotic. Manually list irrelevant words that are not related to personality polarity, with about 460 irrelevant words, such as: takeout, basketball, table tennis, computer, keyboard, etc. Step 2: Vectorize all personality polarity representative words and irrelevant words into 300-dimensional vector representations using w2v (word to vector). Step 3: Construct a multi-layer perceptron for supervised training of the 5 major personalities. For convenience, set labels for each personality. label = 0 represents a positive personality polarity, label = 1 represents a negative personality polarity, and label = 2 represents an irrelevant word. Step 4: Obtain 5 three-classifiers, input the w2v vector of the word, and output the result as: [positive polarity probability, negative polarity probability, irrelevant probability]. After training, the accuracy of the 5 three-classifiers is greater than the preset probability threshold. Here, the larger the preset probability threshold, the higher the accuracy of the trained three-classifier, but the longer the training period. To balance, the preset probability threshold can be set to 80%, but is not limited to this. After testing, the accuracy of the 5 trained three-classifiers is shown in Table 1:
[0067]
[0068]
[0069] Table 1
[0070] Step 103: Determine the polarity value of the target character's personality in each personality dimension for indicating the degree of polarity based on the plurality of polarity probability values of the words in the first target corpus.
[0071] It should be noted that different personality dimensions do not affect each other, and the polarity value of the target character's personality in each personality dimension can be determined separately, and then the polarity value of the target character's personality in each personality dimension is obtained. For example, for affinity, the polarity value of the target character in affinity can be determined based on the probability value of the word representing cold and the probability value of the word representing affinity.
[0072] Step 104: Output the target character and the polarity value of the target character in each personality dimension.
[0073] It should be noted that the target role and the polarity value of each character dimension are the data analysis results of the target role, and when outputting the data analysis results, the data analysis results can be directly displayed, and of course the data analysis results can be sent to the preset mailbox in the form of automatically triggered email. The data analysis results can help different demand parties to better complete their own work. For example, when the demand party is a director, according to the polarity value of the target role in each character dimension, the character of the target role can be quickly and accurately understood, and then when selecting a role, the actor who meets the target role can be selected according to the understanding of different actors.
[0074] In the embodiment of the application, the first target corpus in the target literary work can be obtained, wherein the first target corpus is the corpus associated with the target role in the target literary work. Compared with the manual data analysis method, the limitation of the first target corpus in the data quantity can be broken through, the accuracy of the data analysis result can be improved by enriching the first target corpus. According to the first target corpus and the preset character polarity analysis model, a plurality of polarity probability values of the words in the first target corpus are determined, wherein the plurality of polarity probability values are the probability values of the words representing different character polarities in each character dimension. The probability values of each word representing different character polarities in each character dimension are determined in units of words. According to the plurality of polarity probability values of the words in the first target corpus, the polarity value of the character of the target role in each character dimension for indicating the polarity degree is determined. The polarity value calculated from the probability value of the word representing the character polarity can realize the quantification of the polarity degree. Through the polarity value, the character features of the target role are more accurately shown. Finally, by outputting the target role and the polarity value of each character dimension, the data analysis result of the target role can be provided to the demand party, so that the demand party can quickly and accurately understand the target role according to the data analysis result, and then better carry out the subsequent work. The character features of the target role are more accurately shown by the polarity value calculated in the embodiment of the application, the analysis result of the target role is more accurate, and the whole process is completed in an automatic form, which consumes less time and can make the demand party quickly obtain the analysis result of the target role.
[0075] Optionally, each character dimension includes a positive character polarity and a negative character polarity in opposite directions; and the step 103 of determining, according to the plurality of polarity probability values of the words in the first target corpus, the polarity value of the character of the target role in each character dimension for indicating the polarity degree can include:
[0076] For each character dimension, the case that the character of the target role is biased to the positive character polarity in the character dimension is determined based on the polarity probability value of each word representing the positive character polarity in the character dimension in the first target corpus.
[0077] It should be noted that since the number of personality dimensions is multiple, each personality dimension needs to be processed separately. Taking two personality dimensions as an example, based on the polarity probability value of each word in the first target corpus representing the positive personality polarity of the first personality dimension, the case that the personality of the target role is biased to the positive personality polarity in the first personality dimension is determined, and based on the polarity probability value of each word in the first target corpus representing the positive personality polarity of the second personality dimension, the case that the personality of the target role is biased to the positive personality polarity in the second personality dimension is determined.
[0078] For each personality dimension, it can be understood that the polarity probability value of the word representing the positive personality polarity can be used as a quantitative value of the degree of the word biasing to the positive personality polarity. The greater the polarity probability value, the more biased to the positive personality polarity, and vice versa. When a certain role is described using a large number of words biased to the positive personality polarity, the personality of the role will also be biased to the positive personality polarity. Therefore, by using the polarity probability value of each word associated with the target role representing the positive personality polarity, the case that the personality of the target role is biased to the positive personality polarity can be determined. The case of biasing to the positive personality polarity can be the degree of biasing to the positive personality polarity, which can be quantified by a specific numerical value, but is not limited thereto.
[0079] For each personality dimension, based on the polarity probability value of each word in the first target corpus representing the negative personality polarity of the personality dimension, the case that the personality of the target role is biased to the negative personality polarity in the personality dimension is determined.
[0080] It should be noted that since the number of personality dimensions is multiple, each personality dimension needs to be processed separately. Taking two personality dimensions as an example, based on the polarity probability value of each word in the first target corpus representing the negative personality polarity of the first personality dimension, the case that the personality of the target role is biased to the negative personality polarity in the first personality dimension is determined, and based on the polarity probability value of each word in the first target corpus representing the negative personality polarity of the second personality dimension, the case that the personality of the target role is biased to the negative personality polarity in the second personality dimension is determined.
[0081] For each personality dimension, it can be understood that the polarity probability value of the word representing the negative personality polarity can be used as a quantitative value of the degree of the word biasing to the negative personality polarity. The greater the polarity probability value, the more biased to the negative personality polarity, and vice versa. When a certain role is described using a large number of words biased to the negative personality polarity, the personality of the role will also be biased to the negative personality polarity. Therefore, by using the polarity probability value of each word associated with the target role representing the negative personality polarity, the case that the personality of the target role is biased to the negative personality polarity can be determined. The case of biasing to the negative personality polarity can be the degree of biasing to the negative personality polarity, which can be quantified by a specific numerical value, but is not limited thereto.
[0082] For each personality dimension, the polarity value of the personality of the target role in the personality dimension is determined based on the case that the personality of the target role is biased to the positive personality polarity and the case that the personality of the target role is biased to the negative personality polarity in the personality dimension.
[0083] It should be noted that since the number of personality dimensions is multiple, each personality dimension needs to be processed respectively. Similarly, taking two personality dimensions as an example, the polarity value of the personality of the target role in the first personality dimension is determined based on the case that the personality of the target role is biased to the positive personality polarity and the case that the personality of the target role is biased to the negative personality polarity in the first personality dimension, and the polarity value of the personality of the target role in the second personality dimension is determined based on the case that the personality of the target role is biased to the positive personality polarity and the case that the personality of the target role is biased to the negative personality polarity in the second personality dimension.
[0084] For each personality dimension, the case that the personality of the target role is biased to the positive personality polarity determines the degree of bias to the positive personality polarity, and the case that the personality of the target role is biased to the negative personality polarity determines the degree of bias to the negative personality polarity. It can be understood that the personality of a character usually biases to the positive personality polarity and the negative personality polarity to different degrees, and the greater the degree of bias to which personality polarity, the closer the personality is to which personality polarity. Specifically, the polarity value is in a pre-set numerical interval, wherein the greater the polarity value is to the maximum value of the numerical interval, the more the personality of the target role is biased to the positive personality polarity, and the greater the polarity value is to the minimum value of the numerical interval, the more the personality of the target role is biased to the negative personality polarity.
[0085] In the embodiment of the present application, for each personality dimension, the case that the personality of the target role is biased to the positive personality polarity and the case that the personality of the target role is biased to the negative personality polarity are comprehensively considered, which can accurately determine which personality polarity the personality of the target role is closest to and quantify the closeness as a polarity value.
[0086] Optionally, the case that the personality of the target role is biased to the positive personality polarity in the personality dimension is determined based on the polarity probability value of each word representing the positive personality polarity in the personality dimension in the first target corpus, comprising:
[0087] The Bayesian mean value of the polarity probability value of each word representing the positive personality polarity in the personality dimension in the first target corpus is calculated based on a Bayesian average algorithm.
[0088] It should be noted that, for each personality dimension, the average size of the polarity probability value of each word representing the positive personality polarity can indicate that most words are biased towards the positive personality polarity, and in this case, the case that most words are biased towards the positive personality polarity can be taken as the case that the personality of the target role is biased towards the positive personality polarity. Therefore, the average of the polarity probability value of each word representing the positive personality polarity can be taken as a quantitative value for measuring the case that the personality of the target role is biased towards the positive personality polarity, but is not limited thereto.
[0089] Since the proportions of different roles in the target literary work are not completely the same. For example, the proportion of role A is large, and the target literary work includes a large number of sentences associated with role A. The proportion of role B is small, and the target literary work includes a few sentences associated with role B. Therefore, even if the average of the polarity probability value of each word representing the positive personality polarity in the corpus related to role B is greater than the average of the polarity probability value of each word representing the positive personality polarity in the corpus related to role A, it cannot be concluded that the personality of role B is more biased towards the positive personality polarity than the personality of role A. Furthermore, when the polarity value determined based on the average of the polarity probability value is compared between different roles, it has little reference value. In order to make the polarity values of different roles finally calculated have comparability, it is necessary to adjust the average of the polarity probability value, so that the adjusted value can reduce or even eliminate the influence caused by the small amount of corpus associated with the role. For example, after obtaining the average of the polarity probability value, a target coefficient or an adjustment amount can be multiplied, wherein the target coefficient and the adjustment amount are both associated with the number of words in the corpus associated with the role, but are not limited thereto. The Bayesian average algorithm can also be used directly. Specifically, formula one:
[0090]
[0091] wherein, represents the Bayesian mean, represents the sum of the polarity probability values of each word representing the positive personality polarity in the first target corpus, n src represents the number of words in the first target corpus, n represents the first preset parameter, and C represents the second preset parameter. Wherein, the first preset parameter and the second preset parameter are both greater than zero, and preferably the specific values of the first preset parameter and the second preset parameter can be determined according to the distribution of the polarity value-role number.
[0092] The Bayesian mean is determined as the deflection value of the personality of the target role biased towards the positive personality polarity in the personality dimension.
[0093] In this step, the deflection value of the personality of the target role biased towards the positive personality polarity in the personality dimension is the quantitative value of the case that the personality of the target role is biased towards the positive personality polarity in the personality dimension.
[0094] Correspondingly, based on the polarity probability value of each word representing the negative personality polarity of the personality dimension in the first target corpus, the case that the personality of the target role is biased to the negative personality polarity of the personality dimension is determined, including:
[0095] Based on the Bayesian average algorithm, the Bayesian mean of the polarity probability value of each word representing the negative personality polarity of the personality dimension in the first target corpus is calculated.
[0096] It should be noted that the process of calculating the Bayesian mean of the polarity probability value of the negative personality polarity is similar to the process of calculating the Bayesian mean of the polarity probability value of the positive personality polarity, except that the polarity probability value of the negative personality polarity is used when calculating the Bayesian mean of the polarity probability value of the negative personality polarity, which will not be described here.
[0097] The Bayesian mean is determined as the bias value of the personality of the target role to the negative personality polarity of the personality dimension.
[0098] In this step, the bias value of the personality of the target role to the negative personality polarity of the personality dimension is the quantitative value of the case that the personality of the target role is biased to the negative personality polarity of the personality dimension.
[0099] Correspondingly, based on the case that the personality of the target role is biased to the positive personality polarity of the personality dimension and the case that the personality of the target role is biased to the negative personality polarity of the personality dimension, the polarity numerical value of the personality of the target role in the personality dimension is determined, including: the ratio of the bias value of the personality of the target role to the positive personality polarity of the personality dimension to the bias value of the personality of the target role to the negative personality polarity of the personality dimension is taken as the polarity numerical value of the personality of the target role in the personality dimension.
[0100] In the embodiment of the application, the Bayesian mean of the polarity probability value of each word representing the personality polarity is calculated based on the Bayesian average algorithm, and the Bayesian mean is taken as the bias value of the target role to the personality polarity, which improves the comparability between different roles and makes the polarity numerical value between different roles more meaningful.
[0101] Optionally, based on the case that the personality of the target role is biased to the positive personality polarity of the personality dimension and the case that the personality of the target role is biased to the negative personality polarity of the personality dimension, the polarity numerical value of the personality of the target role in the personality dimension is determined, including:
[0102] Based on the case that the personality of the target role is biased to the positive personality polarity of the personality dimension and the case that the personality of the target role is biased to the negative personality polarity of the personality dimension, an initial polarity numerical value is determined.
[0103] In this step, for each personality dimension, if the personality of the target role is biased to the positive personality polarity, the mean of the polarity probability value of each word in the first target corpus representing the positive personality polarity can be used, and if the personality of the target role is biased to the negative personality polarity, the mean of the polarity probability value of each word in the first target corpus representing the negative personality polarity can be used, but not limited to this. In addition, the Bayesian mean of the polarity probability value of each word in the first target corpus representing the positive personality polarity can also be used for the case that the personality of the target role is biased to the positive personality polarity, and the Bayesian mean of the polarity probability value of each word in the first target corpus representing the negative personality polarity can also be used for the case that the personality of the target role is biased to the negative personality polarity. The process of calculating the Bayesian mean is not described here.
[0104] Based on the case that the personality of the target role in the target literary work is biased to the positive personality polarity and the case that the personality of the target role is biased to the negative personality polarity in the personality dimension, the average polarity value of the plurality of roles is determined.
[0105] It should be noted that the description type or style of the literary work has a great influence on the personality of the characters. For example, the characters in a "nervous" script are generally nervous. In order to eliminate this influence, the average polarity value of the plurality of roles in the target literary work is designed, so as to quantify the writing style of the target literary work. The preset condition is a preset screening condition for screening the plurality of roles in the target literary work. Based on the average case of the plurality of roles screened, the writing style of the target literary work is determined.
[0106] Specifically, the average polarity value can be calculated by using Formula Two. Formula Two:
[0107]
[0108] wherein, V work represents the average polarity value, represents the Bayesian mean of the polarity probability value based on the positive personality polarity corresponding to the role, represents the sum of the polarity probability value of each word in the corpus of the role representing the positive personality polarity, n src represents the number of words in the corpus of the role, represents the Bayesian mean of the polarity probability value based on the negative personality polarity corresponding to the role, represents the sum of the polarity probability value of each word in the corpus of the role representing the negative personality polarity.
[0109] Based on the initial polarity value, the average polarity value, and the role quantity distribution of different personality dimensions and different personality polarities in the literary work, the polarity value of the personality of the target role in the personality dimension is determined.
[0110] It should be noted that, for each personality dimension, the distribution of the number of characters in literary works across different personality polarities is usually a normal or near-normal distribution. To ensure that the polarity values of the target character conform to the overall distribution or near-normal distribution, a model formula can be pre-created based on the distribution of the number of characters with different personality polarities across different personality dimensions in literary works, such as Formula 3:
[0111]
[0112] Among them, V tar V represents the polarity value. src V represents the initial polarity value. work This represents the average polarity value, where 'a' is a preset target parameter. The preset target parameter can be 0.5, but is not limited to this. The specific value of the preset target parameter can be determined based on the distribution of polarity values versus the number of characters. For example... Figure 2 The image shown is a line graph illustrating the relationship between the polarity value and the number of characters in a target literary work, determined using the data analysis method for characters in literary works provided in this embodiment of the invention. Here, the target literary work can be a film script or a television script. Figure 2 The line graph shown is based on the movie script. Figure 3 The line graph shown is based on a TV drama script. Among them, Figure 2 and Figure 3 Five lines are plotted, each corresponding to a personality dimension. Both distributions conform to a normal or near-normal distribution.
[0113] In this embodiment of the invention, by quantifying the average personality polarity of characters in the target literary work, the influence of writing style on the target characters can be eliminated, thereby improving the comparability between literary works of different styles.
[0114] Optionally, multiple roles that meet the preset conditions include: all roles in the target literary work whose number of descriptive sentences is less than a preset threshold.
[0115] It should be noted that literary works often feature characters with limited screen time; these characters are not representative and can be ignored to reduce data processing volume. As shown in Table 2, statistical analysis of some film scripts reveals that approximately 4.5% of characters have only 0-4 lines of dialogue or behavioral descriptions, and approximately 64.2% have only 5-50 lines of dialogue or behavioral descriptions. Therefore, when determining the preset conditions, a preset threshold of 5 can be set; characters meeting the preset conditions are those whose relevant descriptive statements exceed the preset threshold. Different preset thresholds can be set for television drama scripts, as shown in Table 2, where a preset threshold of 10 can be used. Of course, the preset threshold can be set according to user needs and is not limited to the two cases mentioned above where the preset threshold is equal to 5 or 10.
[0116]
[0117] Table 2
[0118] In the embodiment of the present application, by screening out some characters with less roles in the target literary works, the data processing amount can be reduced, and the time consumption of the entire data analysis process can be further shortened.
[0119] Optionally, according to the first target corpus and the preset character polarity analysis model, a plurality of polarity probability values of the words in the first target corpus are determined, including:
[0120] Each word in the first target corpus is input into the preset character polarity analysis model respectively, and initial probability values of each word representing different character polarities under each character dimension are obtained.
[0121] Based on the influence of different sentence types on the description of the character and the sentence type of the word, the weight of the word is determined.
[0122] It should be noted that in many literary works, the language and behavior of some characters are not consistent, and both have different influences on the character. Therefore, "it is better to see how a person acts than how a person speaks". Different weights can be set for different sentence types in the first target corpus, and the weight is the importance thereof. For example, different weights are set for the first sentence type representing the language of the character and the second sentence type representing the behavior of the character. Specifically, the weight of the first sentence type can be 0.5, and the weight of the second sentence type can be 1.0, but is not limited thereto.
[0123] For each word, the initial probability value of the word is multiplied by the weight of the word to obtain the probability value of the word representing different character polarities under each character dimension.
[0124] In the embodiment of the present application, considering the influence of different sentence types on the character, the corresponding weight of the word is set, so that more accurate polarity probability values can be obtained.
[0125] Optionally, after determining the polarity value of the character of the target character under each character dimension for indicating the polarity degree according to the plurality of polarity probability values of the words in the first target corpus, the method further includes:
[0126] Based on the pre-set numerical interval corresponding to each polarity level under each character dimension, the polarity level of the character of the target character under each character dimension is determined.
[0127] It should be noted that the difference between different polarity degrees can be more intuitive through different polarity levels. Here, the number of polarity levels and their corresponding numerical intervals can be set according to the needs. Preferably, seven different polarity levels can be set to represent different polarity degrees. Each polarity level corresponds to a different numerical interval. Specifically, each personality polarity can be divided into three different polarity levels: “a little”, “more” and “extremely”, and the middle polarity level is set as “neutral”. The following takes the affinity as an example to show the division of personality levels, as shown in Table 3:
[0128]
[0129] Table 3
[0130] Output the polarity level of the target character in each personality dimension.
[0131] It should be noted that for each personality dimension, the number of characters and the polarity level also conform to the normal distribution or the quasi-normal distribution. As shown in Figure 4 and Figure 5 , the bar chart of the number of characters and the polarity level of the target literary work in the affinity is shown. Among them, Figure 4 The bar chart is based on a movie script, Figure 5 The bar chart is based on a TV script.
[0132] Preferably, the polarity level and the polarity value can be output together, as shown in Table 4:
[0133]
[0134]
[0135] Table 4
[0136] In the embodiment of the present application, different polarity levels are divided under each personality dimension, and the polarity level of the target character in each personality dimension is determined according to the polarity value under each personality dimension, so that the demand side can more intuitively understand the personality characteristics of the target character.
[0137] Optionally, after the step of outputting the polarity level of the target character in each personality dimension, the method further comprises:
[0138] According to the polarity level of each personality dimension, each character in the preset character personality database is matched with the target character respectively.
[0139] In this step, the preset character personality database stores the identity information of multiple characters and the polarity level of each character in each personality dimension. When the target role matches the character in the preset character personality database, the polarity levels of both parties in the same personality dimension are matched. For example, when the target role matches the target character in the preset character personality database, the polarity level of the target role in affinity is compared with the polarity level of the target character in affinity, the polarity level of the target role in conscientiousness is compared with the polarity level of the target character in conscientiousness, the polarity level of the target role in extraversion is compared with the polarity level of the target character in extraversion, the polarity level of the target role in openness is compared with the polarity level of the target character in openness, and the polarity level of the target role in neuroticism is compared with the polarity level of the target character in neuroticism. It can be understood that the characters in the preset character personality database are characters in real life, for example, can be actors, and the identity information can include the name and personal image of the actor. Each character in the preset character personality database corresponds to at least one set of polarity levels, and each set of polarity levels is the polarity level of the character successfully played by the character in each personality dimension. For example, the character Zhang San in the preset character personality database has successfully played multiple different roles (role A, role B, and role C) in multiple different movies or TV series, and the preset character personality database corresponding to Zhang San stores the polarity level of role A in each personality dimension, the polarity level of role B in each personality dimension, and the polarity level of role C in each personality dimension.
[0140] Output the identity information of the character matched with the target role successfully.
[0141] It should be noted that the matching success condition can be freely set according to the needs, for example, the polarity levels of both parties in each personality are the same, and it is considered that the matching is successful, but it is not limited thereto.
[0142] In the embodiment of the application, the target role can be matched with the character in the preset character personality database according to the polarity level of the target role, and then the identity information of the character matched successfully is output, which facilitates the selection of actors.
[0143] As shown in FIG. Figure 6 The actual application diagram of the data analysis method of the character in the literary work provided by the embodiment of the application is shown in FIG.
[0144] Step 101: obtaining a first target corpus associated with the target role, wherein the first target corpus is the dialogue of the target role and the sentence describing the behavior / activity of the target role.
[0145] Step 102: data analysis and data cleaning are performed on the first target corpus. Irrelevant sentences or words in the first target corpus are deleted, and the remaining sentences are segmented to obtain a large number of words associated with the target role.
[0146] Step 104: the polarity value of the target role in each personality dimension is calculated. Specifically, the probability value of each word in the large number of words associated with the target role in representing each personality polarity in different personality dimensions is determined by the five personality classifiers, i.e., the above-mentioned preset personality polarity analysis model. The first target ratio is obtained based on the original personality model of the character and the Bayesian average algorithm, and the second target ratio is obtained based on the average personality model of the script character, and then the first target ratio and the second target ratio are calculated by the quotient model to obtain the polarity value of the target role in each personality dimension. Here, the original personality model of the character is a model for obtaining the first target ratio by dividing the sum of the probability values of the words representing the positive personality polarity by the sum of the probability values of the words representing the negative personality polarity. The average personality model of the script character is the model shown in Formula 2, and the quotient model is the model shown in Formula 3.
[0147] Step 105: role personality level division. According to the numerical interval corresponding to the different polarity levels set in advance, the numerical interval to which the polarity value in each personality dimension belongs is determined.
[0148] Step 106: output the role personality result. The target role, the polarity value in each personality dimension, and the polarity level in each personality dimension are taken as the data analysis result of the target role, and the data analysis result is output.
[0149] In the embodiment of the application, the polarity value obtained by calculation more accurately shows the personality characteristics of the target role, so that the analysis result of the target role is more accurate, and the whole process is completed in an automated form, which is time-saving and can enable the demander to quickly obtain the analysis result of the target role.
[0150] The above introduces the data analysis method of the role in the literary work provided by the embodiment of the application, and the data analysis device of the role in the literary work provided by the embodiment of the application will be introduced below with reference to the drawings.
[0151] Referring to Figure 7 The embodiment of the application also provides a data analysis device for a role in a literary work, which comprises:
[0152] The acquisition module 71 is configured to acquire a first target corpus in a target literary work, wherein the first target corpus is a corpus associated with a target role in the target literary work;
[0153] The first determining module 72 is configured to determine a plurality of polarity probability values of the words in the first target corpus according to the first target corpus and a preset personality polarity analysis model, wherein the plurality of polarity probability values represent probabilities of the words representing different personality polarities in each personality dimension.
[0154] The second determining module 73 is configured to determine a personality value of the personality of the target role in each personality dimension according to the plurality of polarity probability values of the words in the first target corpus.
[0155] The output module 74 is configured to output the target role and the personality value of the target role in each personality dimension.
[0156] Optionally, each personality dimension includes a positive personality polarity and a negative personality polarity in opposite directions.
[0157] The second determining module 73 includes:
[0158] The first bias unit is configured to determine, for each personality dimension, a case in which the personality of the target role is biased to the positive personality polarity in the personality dimension based on the polarity probability values of the words in the first target corpus representing the positive personality polarity in the personality dimension.
[0159] The second bias unit is configured to determine, for each personality dimension, a case in which the personality of the target role is biased to the negative personality polarity in the personality dimension based on the polarity probability values of the words in the first target corpus representing the negative personality polarity in the personality dimension.
[0160] The determining unit is configured to determine, for each personality dimension, the personality value of the personality of the target role in the personality dimension based on the case in which the personality of the target role is biased to the positive personality polarity in the personality dimension and the case in which the personality of the target role is biased to the negative personality polarity in the personality dimension.
[0161] Optionally, the first bias unit is specifically configured to calculate a Bayesian mean of the polarity probability values of the words in the first target corpus representing the positive personality polarity in the personality dimension based on a Bayesian average algorithm; and determine the Bayesian mean as a bias value of the personality of the target role biased to the positive personality polarity in the personality dimension.
[0162] Correspondingly, the second bias unit is specifically configured to calculate a Bayesian mean of the polarity probability values of the words in the first target corpus representing the negative personality polarity in the personality dimension based on a Bayesian average algorithm; and determine the Bayesian mean as a bias value of the personality of the target role biased to the negative personality polarity in the personality dimension.
[0163] Optionally, the determining unit is specifically configured to determine an initial polarity value based on whether the character of the target role is biased towards a positive character polarity or a negative character polarity in the character dimension, determine an average polarity value of a plurality of roles based on whether the characters of the plurality of roles meeting the preset condition are biased towards a positive character polarity or a negative character polarity in the character dimension, and determine the polarity value of the character of the target role in the character dimension based on the initial polarity value, the average polarity value, and a pre-determined distribution of the number of roles with different character polarities in different character dimensions in the literary work.
[0164] Optionally, the plurality of roles meeting the preset condition include all roles in the target literary work whose number of description sentences is less than a preset threshold.
[0165] Optionally, the first determining module comprises:
[0166] The model unit is configured to input each word in the first target corpus into a preset character polarity analysis model to obtain an initial probability value of each word representing different character polarities in each character dimension.
[0167] The weight unit is configured to determine a weight of each word based on the influence of different sentence types on the description of the character and the sentence type of the sentence to which the word belongs.
[0168] The calculation unit is configured to multiply the initial probability value of each word by the weight of the word to obtain a probability value of the word representing different character polarities in each character dimension.
[0169] Optionally, the device further comprises:
[0170] The level module is configured to determine a polarity level of the character of the target role in each character dimension based on a pre-set numerical interval corresponding to each polarity level in each character dimension.
[0171] The level output module is configured to output the polarity level of the character of the target role in each character dimension.
[0172] The data analysis device for a role in a literary work provided by the embodiment of the present application can realize Figure 1 and Figure 6 the method embodiment of the data analysis method for a role in a literary work, and thus each process is not repeated here.
[0173] In the embodiment of the present application, the polarity value obtained by calculation more finely shows the character features of the target role, so that the analysis result of the target role is more accurate, and the entire process is completed in an automated form, which is time-saving and can enable the demander to quickly obtain the analysis result of the target role.
[0174] The embodiment of the present application also provides an electronic device, such as Figure 8 As shown in the figure, the electronic device comprises a processor 801, a communication interface 802, a memory 803 and a communication bus 804, wherein the processor 801, the communication interface 802 and the memory 803 complete mutual communication through the communication bus 804;
[0175] The memory 803 is used for storing a computer program.
[0176] The processor 801 is used for executing the program stored in the memory 803, and realizes the following steps:
[0177] Obtaining a first target corpus in a target literary work, wherein the first target corpus is a corpus associated with a target role in the target literary work;
[0178] According to the first target corpus and a preset personality polarity analysis model, a plurality of polarity probability values of a word in the first target corpus are determined, wherein the plurality of polarity probability values are probability values of the word representing different personality polarities under each personality dimension;
[0179] According to the plurality of polarity probability values of the word in the first target corpus, a polarity value of the personality of the target role under each personality dimension is determined, wherein the polarity value is used for indicating a polarity degree.
[0180] The target role and the polarity value of the target role under each personality dimension are output.
[0181] The communication bus mentioned above can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or only one type of bus.
[0182] The communication interface is used for communication between the terminal and other devices.
[0183] The memory can include a random access memory (RAM) and can also include a non-volatile memory, for example, at least one disk memory. Optionally, the memory can also be at least one storage device located away from the aforementioned processor.
[0184] The processor described above can be a general processor, including a central processing unit (CPU), a network processor (NP), etc.; or can be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component.
[0185] In yet another embodiment provided by the present application, a computer readable storage medium is provided, and the computer readable storage medium stores instructions which, when executed on a computer, cause the computer to perform the data analysis method of a character in a literary work according to any one of the above embodiments.
[0186] In yet another embodiment provided by the present application, a computer program product is provided, and the computer program product includes instructions which, when executed on a computer, cause the computer to perform the data analysis method of a character in a literary work according to any one of the above embodiments.
[0187] In the above embodiments, the implementation can be wholly or partially achieved by software, hardware, firmware or any combination thereof. When implemented by software, the implementation can be wholly or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the implementation wholly or partially produces the flow or function according to the embodiments of the present application. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer readable storage medium or transferred from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transferred from one website, computer, server or data center to another website, computer, server or data center through a wired (for example, coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (for example, infrared, wireless, microwave, etc.) manner. The computer readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center, etc. including one or more available media sets. The available medium can be a magnetic medium (for example, floppy disk, hard disk, magnetic tape), an optical medium (for example, DVD) or a semiconductor medium (for example, solid state disk (SSD)) etc.
[0188] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0189] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0190] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.
Claims
1. A method of data analysis of a character in a literary work, characterized by, The method comprises: acquiring a first target corpus in a target literary work, wherein the first target corpus is a corpus associated with a target character in the target literary work, including dialogue of the target character and / or sentences describing behaviors of the target character; determining, according to the first target corpus and a preset personality polarity analysis model, a plurality of polarity probability values of words in the first target corpus, wherein the plurality of polarity probability values are probability values of the words representing different personality polarities in each personality dimension; and the preset personality polarity analysis model is a classifier; determining, according to the plurality of polarity probability values of the words in the first target corpus, a personality value of the target character in each personality dimension for indicating a polarity degree; outputting the target character and the personality value of the target character in each personality dimension; the determining, according to the first target corpus and the preset personality polarity analysis model, the plurality of polarity probability values of the words in the first target corpus comprises: inputting each word in the first target corpus into the preset personality polarity analysis model respectively to obtain initial probability values of each word representing different personality polarities in each personality dimension; determining a weight of the word based on an influence situation of different sentence types on describing a personality and a sentence type of a sentence to which the word belongs, wherein the sentence type includes dialogue of the target character and a description of behaviors of the target character; multiplying the initial probability value of each word by the weight of the word to obtain the probability value of the word representing different personality polarities in each personality dimension.
2. The method of claim 1, wherein, Each personality dimension includes a positive personality polarity and a negative personality polarity in opposite directions; the determining, according to the plurality of polarity probability values of the words in the first target corpus, the personality value of the target character in each personality dimension for indicating a polarity degree comprises: for each personality dimension, determining a case that the personality of the target character is biased to the positive personality polarity in the personality dimension based on the polarity probability values of each word representing the positive personality polarity in the personality dimension in the first target corpus; for each personality dimension, determining a case that the personality of the target character is biased to the negative personality polarity in the personality dimension based on the polarity probability values of each word representing the negative personality polarity in the personality dimension in the first target corpus; for each personality dimension, determining the personality value of the target character in the personality dimension based on the case that the personality of the target character is biased to the positive personality polarity in the personality dimension and the case that the personality of the target character is biased to the negative personality polarity in the personality dimension.
3. The method of claim 2, wherein, the determining, based on the polarity probability values of each word representing the positive personality polarity in the personality dimension in the first target corpus, the case that the personality of the target character is biased to the positive personality polarity in the personality dimension comprises: calculating a Bayesian mean of the polarity probability values of each word representing the positive personality polarity in the personality dimension in the first target corpus based on a Bayesian average algorithm; and the determining, based on the polarity probability values of each word representing the negative personality polarity in the personality dimension in the first target corpus, the case that the personality of the target character is biased to the negative personality polarity in the personality dimension comprises: calculating a Bayesian mean of the polarity probability values of each word representing the negative personality polarity in the personality dimension in the first target corpus based on a Bayesian average algorithm. determine the deflection value of the personality of the target role in the personality dimension to the positive personality polarity as the Bayesian mean value; Correspondingly, the case that the personality of the target role in the personality dimension is deflected to the negative personality polarity is determined based on the polarity probability value of each word in the first target corpus representing the negative personality polarity in the personality dimension, including: calculating, based on a Bayesian average algorithm, a Bayesian mean value of the polarity probability value of each word in the first target corpus representing the negative personality polarity in the personality dimension; determine the deflection value of the personality of the target role in the personality dimension to the negative personality polarity as the Bayesian mean value.
4. The method of claim 2, wherein, The case that the personality of the target role in the personality dimension is deflected to the positive personality polarity and the case that the personality of the target role in the personality dimension is deflected to the negative personality polarity are used to determine the polarity numerical value of the personality of the target role in the personality dimension, including: determine an initial polarity numerical value based on the case that the personality of the target role in the personality dimension is deflected to the positive personality polarity and the case that the personality of the target role in the personality dimension is deflected to the negative personality polarity; determine an average polarity numerical value of a plurality of roles in the target literary work based on the case that the personality of the plurality of roles in the target literary work in the personality dimension is deflected to the positive personality polarity and the case that the personality of the plurality of roles in the target literary work in the personality dimension is deflected to the negative personality polarity; determine the polarity numerical value of the personality of the target role in the personality dimension based on the initial polarity numerical value, the average polarity numerical value, and the distribution of the number of roles with different personality polarities in different personality dimensions in literary works.
5. The method of claim 4, wherein, The plurality of roles satisfying the preset condition includes all roles in the target literary work whose number of description sentences is less than a preset threshold.
6. The method of claim 1, wherein, After determining the polarity numerical value of the personality of the target role in each personality dimension for indicating the degree of polarity based on the plurality of polarity probability values of the words in the first target corpus, the method further includes: determine the polarity level of the personality of the target role in each personality dimension based on the preset numerical interval corresponding to each polarity level in each personality dimension; output the polarity level of the personality of the target role in each personality dimension.
7. A data analysis apparatus for a character in a literary work, characterized by, The device includes: an acquisition module configured to acquire a first target corpus in a target literary work, wherein the first target corpus is a corpus associated with a target role in the target literary work, including dialogue of the target role and / or sentences describing behaviors of the target role; a first determination module configured to determine a plurality of polarity probability values of words in the first target corpus according to the first target corpus and a preset personality polarity analysis model, wherein the plurality of polarity probability values are probability values of the words representing different personality polarities in each personality dimension; and the preset personality polarity analysis model is a classifier; a second determination module configured to determine a polarity numerical value of the personality of the target role in each personality dimension for indicating the degree of polarity based on the plurality of polarity probability values of the words in the first target corpus; an output module configured to output the target role and the polarity numerical value of the target role in each personality dimension. The first determining module comprises: a model unit configured to input each word in the first target corpus into a preset personality polarity analysis model to obtain an initial probability value of each word representing different personality polarities in each personality dimension; a weight unit configured to determine a weight of each word based on an influence of different sentence types on the description of the personality of the character and a sentence type of a sentence to which the word belongs, wherein the sentence type comprises dialogue of the target character and description of behavior of the target character; a calculation unit configured to multiply the initial probability value of each word by the weight of the word to obtain a probability value of the word representing different personality polarities in each personality dimension.
8. An electronic device, comprising: The device comprises a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus. The memory is configured to store a computer program. The processor is configured to execute the program stored in the memory to implement the steps of the data analysis method of the character in the literary work according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer program is stored in the computer readable storage medium and is executed by the processor to implement the steps of the data analysis method of the character in the literary work according to any one of claims 1-6.
Citation Information
Patent Citations
Character personality analysis method based on social data
CN109766452A