Personality dictionary expansion method and device, equipment and storage medium
By using corpus classification model to classify and annotate social media data, filter out high-quality words and add them to personality dictionaries, the problem of insufficient vocabulary coverage in the processing of social media data is solved, and wider vocabulary coverage and more accurate personalized analysis are achieved.
Patent Information
- Application Number
- CN202510248644.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-24
AI Technical Summary
The existing personalized dictionary has the problem of insufficient vocabulary coverage when processing social media data, which is difficult to meet the needs of personalized analysis.
The corpus classification model is used to classify social media corpus, generate labeling results, and filter out the terms that meet the conditions by setting the reliability threshold, and add them to the pre-constructed personality dictionary to expand the coverage of the dictionary.
It has achieved effective expansion of personality dictionary, enhanced the coverage and practicality of the dictionary, and can more accurately identify and classify diverse language expressions.
Smart Images

Figure CN120197616A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of natural language processing, and particularly relates to a method, device, equipment and storage medium for expanding a personality dictionary. Background Art
[0002] The Big Five Personality Traits is a widely used personality trait model in psychology, which divides human personality into five main dimensions: conscientiousness, openness, extraversion, agreeableness, and neuroticism. Each main dimension also includes 6 sub-dimensions to describe the specific manifestations of personality traits. However, existing personality dictionaries usually only contain a limited number of words and cannot comprehensively cover these Big Five personality dimensions and their sub-dimensions, resulting in significant limitations in psychological research and applications, especially in large-scale text data processing and automated analysis tasks.
[0003] Existing dictionary expansion methods mainly rely on manual annotation or simple statistical models, with low efficiency and poor accuracy and consistency of annotation results. BERT (Bidirectional Encoder Representations from Transformers) is a deep learning-based language model that learns the deep semantic relationships of words from the context through a bidirectional encoder and performs well in natural language processing tasks. However, in the task of expanding a personality dictionary, how to effectively use the BERT model for multi-level training and automatically generate words covering the Big Five personality dimensions and their sub-dimensions is still a technical problem to be solved urgently.
[0004] A large amount of unstructured text data generated on social media platforms such as Facebook and Twitter has made it difficult for existing dictionaries and annotation tools to handle diverse and complex language expressions. Existing tools have problems with insufficient vocabulary coverage when processing this data and are difficult to meet the needs of personalized analysis. Summary of the Invention
[0005] The purpose of this application is to provide a method, device, equipment and storage medium for expanding a personality dictionary to overcome the defects of the prior art, aiming to expand the personality dictionary and enhance the coverage and practicality of the personality dictionary.
[0006] The purpose of this application is achieved through the following technical solutions:
[0007] A method for expanding a personality dictionary, the method comprising:
[0008] Obtain social media corpus, classify the social media corpus with a corpus classification model, and generate an annotation result according to the classification result;
[0009] Screen out the second corpus words that meet the preset confidence threshold according to the annotation results, and add the second corpus words to the pre-constructed personality dictionary to expand the pre-constructed personality dictionary.
[0010] Further, the classification of the social media corpus by the corpus classification model includes:
[0011] Perform word segmentation on the social media corpus to obtain the first corpus words, and input the first corpus words into the corpus classification model for dimension and sub-dimension classification.
[0012] Further, the construction method of the corpus classification model includes:
[0013] Randomly select several seed words of personality dimensions and their sub-dimensions from the pre-constructed personality dictionary, and the personality dictionary includes the vocabulary of the Big Five personality dimensions and their sub-dimensions;
[0014] Obtain a preset fixed sentence template, fill the seed words into the fixed sentence template to generate a standard sentence pattern;
[0015] Use the pre-constructed generation model to generate training corpus based on the standard sentence pattern;
[0016] Use a word segmentation tool to perform word segmentation on the training corpus to obtain multiple words, and assign labels to each word to obtain an annotated corpus set, and the labels include the dimensions and sub-dimensions corresponding to the words;
[0017] Input the annotated corpus set into the pre-trained language model to perform dimension-level training and sub-dimension-level training on the pre-trained language model to obtain a corpus classification model.
[0018] Further, the pre-constructed generation model includes a rule-based generator, a neural network generation model, or a natural language processing model;
[0019] The pre-trained language model includes a multi-layer neural network model based on deep learning.
[0020] Further, the multi-layer neural network model based on deep learning includes a bidirectional encoder representation transformer model.
[0021] Further, the method performs dimension-level training through a bidirectional encoder representation transformer model, and the loss function L of the dimension-level training dimension includes:
[0022]
[0023] where y i is the true label, is the predicted label, and N is the number of samples.
[0024] Further, the method uses multiple independent bidirectional encoder representation transformers to train each sub-dimension, and the training loss function L at the sub-dimension level sup-dimension includes:
[0025]
[0026] where z j is the true label of the sub-dimension, is the predicted label, and M is the number of sub-dimension samples. The second aspect of the present invention provides a personality dictionary expansion device, characterized in that the device includes:
[0027] A classification module, which obtains social media corpus, classifies the social media corpus with a corpus classification model, and generates an annotation result according to the classification result;
[0028] An expansion module, which filters out second corpus words that meet a preset confidence threshold according to the annotation result, and adds the second corpus words to a pre-constructed personality dictionary to expand the pre-constructed personality dictionary.
[0029] The third aspect of the present invention provides a personality dictionary expansion device, including: a memory and at least one processor, wherein computer-readable instructions are stored in the memory, and the memory and the at least one processor are interconnected by a line; the at least one processor calls the computer-readable instructions in the memory, so that the personality dictionary expansion device executes each step of the above-mentioned personality dictionary expansion method.
[0030] The fourth aspect of the present invention provides a computer-readable storage medium, in which computer-readable instructions are stored, and when it runs on a computer, it causes the computer to execute each step of the above-mentioned personality dictionary expansion method.
[0031] The beneficial effects of this application are as follows:
[0032] In the technical solution provided by this application, a corpus classification model is used to accurately and efficiently classify the dimensions and sub-dimensions of social media corpus, and high-quality new corpus words are screened out by setting a confidence threshold, thus effectively expanding the pre-constructed personality dictionary and enhancing the coverage and practicability of the personality dictionary; moreover, the method for constructing the corpus classification model cleverly combines the pre-constructed personality dictionary with fixed sentence templates, and a large number of training corpus with standard sentence patterns are efficiently generated through a generative model; subsequently, using a word segmentation tool and label assignment, each word in the corpus is accurately associated with the corresponding personality dimension and sub-dimension to form a high-quality labeled corpus set; finally, by inputting these labeled data into a pre-trained language model for dimension-level training and sub-dimension-level in-depth training, not only the recognition ability of the corpus classification model at the dimension and sub-dimension levels is strengthened, but also the accuracy and practicability of the corpus classification model in the personality trait classification task are ensured. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is a flowchart of the method for expanding the personality dictionary in an embodiment of this application;
[0034] Figure 2 is a block diagram of the structure of the device for expanding the personality dictionary in an embodiment of this application;
[0035] Figure 3 is a schematic structural diagram of the device for expanding the personality dictionary in an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] The following uses specific specific examples to illustrate the implementation manners of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific implementation manners, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0037] Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope of protection of this application.
[0038] A large amount of unstructured text data generated on social media platforms such as Facebook and Twitter has made it difficult for existing dictionaries and annotation tools to handle diverse and complex language expressions. Existing tools have the problem of insufficient vocabulary coverage when processing these data and are difficult to meet the needs of personalized analysis.
[0039] In order to solve the above technical problems, the following embodiments of the personality dictionary expansion method, device, equipment and storage medium of the present application are proposed.
[0040] Reference Figure 1 ,like Figure 1 The figure is a flowchart of a personality dictionary expansion method according to an embodiment of the present application, and the method comprises the following steps:
[0041] S101: Acquire social media corpus, classify the social media corpus using a corpus classification model, and generate annotation results based on the classification results.
[0042] S102: Filter out second corpus words that meet a preset confidence threshold according to the annotation results, and add the second corpus words to a pre-constructed personality dictionary to expand the pre-constructed personality dictionary.
[0043] As an implementation mode, the corpus classification model proposed in this embodiment includes randomly selecting a number of seed words of personality dimensions and their sub-dimensions from a pre-constructed personality dictionary, wherein the personality dictionary includes vocabulary of the Big Five personality dimensions and their sub-dimensions; obtaining a preset fixed sentence template, filling the selected seed words into the fixed sentence template, and generating a standard sentence pattern; using a pre-constructed generation model to generate training corpus based on the standard sentence pattern; using a word segmentation tool to perform word segmentation on the training corpus to obtain a plurality of words, and assigning a label to each word to obtain a labeled corpus set, wherein the label includes a dimension and a sub-dimension corresponding to the word; inputting the labeled corpus set into the pre-trained language model for the pre-trained language model. The model is trained at the dimension level and sub-dimension level to obtain a corpus classification model, which cleverly combines the pre-built personality dictionary with the fixed sentence template, and efficiently generates a large number of standard sentence training corpora through the generative model; then, using the word segmentation tool and label assignment, each word in the corpus is accurately associated with the corresponding personality dimension and sub-dimension to form a high-quality annotated corpus set; finally, by inputting these annotated data into the pre-trained language model for dimension level training and sub-dimension level deep training, it not only enhances the recognition ability of the corpus classification model at the dimension and sub-dimension levels, but also ensures the accuracy and practicality of the corpus classification model in the task of personality trait classification.
[0044] In this embodiment, the Big Five personality dimensions are specifically: Openness, Conscientiousness, Extroversion, Agreeableness, and Neuroticism. For each dimension, its sub-dimensions are further subdivided. For example, Openness may include sub-dimensions such as imagination, aesthetics, and intellectual curiosity; Conscientiousness may include sub-dimensions such as organization, responsibility, and self-discipline. Then, using expert knowledge, words related to each dimension and sub-dimension are collected. These words can be adjectives, nouns, or verbs that can accurately describe the characteristics of the dimension or sub-dimension. Next, an expert team is organized to evaluate the collected words to ensure that they are closely related to the corresponding personality dimension and sub-dimension. Finally, the evaluated words are classified according to the personality dimension and sub-dimension to construct a personality dictionary. The personality dictionary includes words of the Big Five personality dimensions and their sub-dimensions. For example:
[0045] Openness:
[0046] Sub-dimensions: imagination (words such as "creative", "fantasy"), aesthetics (words such as "aesthetic sense", "art"), intellectual curiosity (words such as "curious", "learning"), etc.
[0047] Conscientiousness:
[0048] Sub-dimensions: organization (words such as "orderly", "planning"), responsibility (words such as "reliable", "responsible"), self-discipline (words such as "self-control", "persistence"), etc.
[0049] Extroversion:
[0050] Sub-dimensions: enthusiasm (words such as "lively", "exuberant"), sociability (words such as "outgoing", "gregarious"), optimism (words such as "positive", "optimistic"), etc.
[0051] Agreeableness:
[0052] Sub-dimensions: trust (words such as "trust", "tolerant"), altruism (words such as "generous", "considerate"), humility (words such as "modest", "low-key"), etc.
[0053] Neuroticism:
[0054] Sub-dimensions: anxiety (words such as "nervous", "anxious"), vulnerability (words such as "sensitive", "fragile"), anger (words such as "irritable", "temperamental"), etc.
[0055] In this embodiment, the number of selected seed words varies for each dimension and its sub-dimensions. For example, 1 to 5 seed words are randomly selected.
[0056] In this embodiment, when filling the selected seed words into the fixed sentence template, ensure that each seed word can be embedded into a suitable context to form a complete sentence structure.
[0057] In this embodiment, the preset fixed sentence template can be: In terms of [personality dimension], [subject] shows the characteristics of [sub-dimension seed word 1] and [sub-dimension seed word 2], which is reflected in [specific situation or behavior]. The preset fixed sentence template can also be: [Subject] has strong [sub-dimension seed word 1] and [sub-dimension seed word 2] in [personality dimension].
[0058] Exemplarily, the selected personality dimensions are openness and conscientiousness. The sub-dimension seed words of openness are imagination and aesthetics, and the sub-dimension seed words of conscientiousness are sense of responsibility and achievement orientation.
[0059] Then the generated standard sentence pattern is: In terms of openness, Xiaoming shows the characteristics of imagination and aesthetics, which is reflected in his frequent participation in art creation and appreciation of classical music.
[0060] Xiaohong has strong sense of responsibility and achievement orientation in conscientiousness.
[0061] In this embodiment, the pre-constructed generation model can be a rule-based generator, a neural network generation model, or a natural language processing model based on deep learning technology.
[0062] In this embodiment, the pre-constructed generation model is ChatGPT. ChatGPT is a natural language processing model centered on GPT (Generative Pre-trained Transformer), which is further optimized and fine-tuned. It combines the advantages of deep learning technology and large-scale pre-training, can generate coherent and natural text responses, and demonstrates strong context understanding ability. The generation process of ChatGPT does not rely on preset rules and templates, but generates text through the learned data features and patterns.
[0063] Using ChatGPT to expand the standard sentence pattern to generate a large number of training corpora containing diverse language expressions can provide a data basis for the preliminary training of the corpus classification model.
[0064] For example, the standard sentence pattern is: In terms of openness, Xiaoming shows the characteristics of imagination and aesthetics, which is reflected in his frequent participation in art creation and appreciation of classical music.
[0065] The training corpus is: In terms of openness, with his rich imagination and unique aesthetic vision, Xiaoli not only shows extraordinary talent in the field of painting, but also often shares his unique insights on fashion trends on social media.
[0066] Openness allows Xiao Wang to have infinite creativity. She not only loves to design various novel products but also often seeks inspiration during her travels and integrates elements of different cultures into her works.
[0067] Another example, the standard sentence pattern is: Xiao Hong has a strong sense of responsibility and achievement orientation in conscientiousness.
[0068] The training corpus is as follows:
[0069] Xiao Hong performs particularly well in conscientiousness. She not only is enthusiastic about her work but also can always complete tasks on time. This strong sense of responsibility and achievement orientation has won her the respect of her colleagues in the workplace.
[0070] As a project manager, Xiao Zhang demonstrates extremely high professional qualities in conscientiousness. He can not only accurately assess the risks of the project but also lead the team to overcome various difficulties to ensure the project is delivered on time.
[0071] In this embodiment, the word segmentation tool can be selected from Jieba, THULAC, Stanford Chinese Segmenter, etc. In this embodiment, the Jieba word segmentation tool is used.
[0072] It can be understood that the words obtained after word segmentation of the training corpus generated based on the seed words of the personality dimension and its sub-dimensions should also correspond to the personality dimension and its sub-dimensions. Therefore, when assigning labels to the words after word segmentation, the personality dimension and its sub-dimensions can be referred to.
[0073] For example, for the sentence "Openness allows Xiao Zhang to have infinite imagination space", the possible words and labels obtained after word segmentation are: Openness (label: Openness dimension), allows (no label), Xiao Zhang (no label), have (no label), infinite (no label), imagination (label: Imagination sub-dimension), space (no label).
[0074] In this embodiment, the pre-trained language model is a multi-layer neural network model based on deep learning.
[0075] An example of the labeled corpus is as follows:
[0076] Training corpus: "Xiao Li shows strong curiosity and exploratory desire in openness."
[0077] Annotation: "Openness" (dimension label), "curiosity" (sub-dimension label), "exploratory desire" (sub-dimension label)
[0078] Training corpus: "Xiao Wang pays great attention to details and organization in conscientiousness."
[0079] Annotation: "Conscientiousness" (dimension label), "Sense of responsibility" (sub-dimension label), "Organization" (sub-dimension label)
[0080] Input the annotated corpus into the pre-trained language model for dimension-level training and sub-dimension-level training of the pre-trained language model. First, perform dimension-level training. In this step, the goal is to enable the model to recognize and classify different personality dimensions. This can be achieved by inputting the sentences in the annotated corpus into the model and setting the dimension labels as the training targets. For example, for the sentences in the above-mentioned annotated corpus, "Openness" and "Conscientiousness" can be used as the training targets to let the pre-trained language model learn how to recognize these dimensions.
[0081] Then comes the sub-dimension-level training. In this step, the goal is to enable the pre-trained language model to further recognize and classify different personality sub-dimensions. This can also be achieved by inputting the sentences in the annotated corpus into the model and setting the sub-dimension labels as the training targets. For example, for the sentences in the above-mentioned annotated corpus, "Curiosity" and "Organization" can be used as the training targets to let the model learn how to recognize these sub-dimensions.
[0082] During the training process, the pre-trained language model will continuously learn based on the sentences and corresponding labels in the annotated corpus. By adjusting the parameters inside the model, the model can gradually improve its ability to recognize personality dimensions and their sub-dimensions.
[0083] In this embodiment, the pre-trained language model is a BERT (Bidirectional Encoder Representations from Transformers) model. Input the annotated corpus into the multi-layer neural network model based on BERT for dimension-level training. At this stage, the model converts the input text into a vector representation through the BERT encoder, and then performs classification through the fully connected layer to output the corresponding dimension labels. The training loss function at the personality dimension level can be expressed as:
[0084]
[0085] where y i is the true label, is the predicted label, and N is the number of samples. Through the cross-entropy loss function, the model gradually optimizes the parameters during iteration to improve the prediction accuracy. After completing the dimension-level training, the model further conducts refined training at the sub-dimension level.
[0086] The sub-dimension-level training is completed through five independent BERT models, and each model focuses on all sub-dimensions of one dimension. The training loss function at the sub-dimension level can be expressed as:
[0087]
[0088] where zj is the true label of the sub-dimension, is the predicted label, and M is the number of sub-dimension samples; finally, through multi-task learning, the training at the dimension and sub-dimension levels is integrated to gradually refine the prediction ability of the model.
[0089] In this embodiment, the specific steps of tokenizing the social media corpus to obtain new corpus words include:
[0090] Preprocess the social media corpus, including removing punctuation marks, extra spaces, and irrelevant characters. Tokenize the preprocessed social media corpus to obtain new corpus words. Here, the new corpus words are the first corpus words.
[0091] In this embodiment, a confidence threshold of 0.8 is preset, and only the prediction results with a confidence higher than this threshold are retained to ensure the reliability and accuracy of label annotation. The results screened by this preset confidence threshold are the second corpus words.
[0092] In this embodiment, the social media corpus includes 50,000 posts crawled from Facebook and Twitter.
[0093] In this embodiment, for the extended personality dictionary, the Weibo-Sentiment Analysis dataset is used, and the improvement of the coverage rate is tested on five dimensions respectively. The results show that the coverage rate of the extraversion dimension has increased by 9.69%, the conscientiousness dimension has increased by 6.51%, the agreeableness dimension has increased by 6.64%, the neuroticism dimension has increased by 6.37%, and the openness dimension has increased by 21.67%. These results indicate that through dictionary expansion, the coverage rate of the personality dictionary on these personality dimensions has been significantly improved, which helps to more accurately identify and classify diverse information in daily expressions.
[0094] This embodiment provides a personality dictionary expansion method, which uses a corpus classification model to accurately and efficiently classify social media corpus into dimensions and sub-dimensions, and selects high-quality new corpus words by setting a confidence threshold, thereby achieving effective expansion of the pre-constructed personality dictionary and enhancing the coverage and practicality of the personality dictionary. Moreover, the corpus classification model construction method cleverly combines the pre-constructed personality dictionary with a fixed sentence template, and efficiently generates a large number of standard sentence training corpora through the generation model; then, using word segmentation tools and label allocation, each word in the corpus is accurately associated with the corresponding personality dimension and sub-dimension to form a high-quality annotated corpus set; finally, by inputting these annotated data into the pre-trained language model for dimension-level training and sub-dimension-level deep training, not only the recognition ability of the corpus classification model at the dimension and sub-dimension levels is enhanced, but also the accuracy and practicality of the corpus classification model in the personality trait classification task is ensured.
[0095] This embodiment also provides a personality dictionary expansion device, referring to Figure 2 ,like Figure 2 The figure is a structural block diagram of a personality dictionary expansion device according to an embodiment of the present application, and the device includes the following structures:
[0096] The classification module 201 is used to obtain social media corpus, perform word segmentation processing on the social media corpus to obtain new corpus words, input the new corpus words into the corpus classification model to perform dimension and sub-dimension classification, and generate a labeling result according to the classification result;
[0097] The expansion module 202 is used to filter out new corpus words that meet a preset confidence threshold according to the annotation results, and add the filtered new corpus words to the pre-constructed personality dictionary to expand the pre-constructed personality dictionary.
[0098] As an embodiment, the device also includes a construction module, including: a selection unit, which is used to randomly select a number of seed words of personality dimensions and their sub-dimensions from a pre-constructed personality dictionary, and the personality dictionary includes vocabulary of the Big Five personality dimensions and their sub-dimensions; a first generation unit, which is used to obtain a preset fixed sentence template, fill the selected seed words into the fixed sentence template, and generate a standard sentence pattern; a second generation unit, which is used to use a pre-constructed generation model to generate training corpus based on the standard sentence pattern; a word segmentation unit, which is used to use a word segmentation tool to perform word segmentation processing on the training corpus to obtain multiple words, and assign a label to each word to obtain a labeled corpus set, and the label includes a dimension and a sub-dimension corresponding to the word; a construction unit, which is used to input the labeled corpus set into a pre-trained language model to perform dimension-level training and sub-dimension-level training on the pre-trained language model to obtain a corpus classification model.
[0099] In this embodiment, a corpus classification model is used to accurately and efficiently classify the dimensions and sub-dimensions of the social media corpus, and high-quality new corpus words are screened out by setting a confidence threshold, thereby achieving an effective expansion of the pre-constructed personality dictionary and enhancing the coverage and practicality of the personality dictionary; moreover, the construction of the corpus classification model cleverly combines the pre-constructed personality dictionary with a fixed sentence template, and efficiently generates a large number of training corpora of standard sentences through a generative model; then, using word segmentation tools and label allocation, each word in the corpus is accurately associated with the corresponding personality dimension and sub-dimension to form a high-quality annotated corpus set; finally, by inputting these annotated data into the pre-trained language model for dimension-level training and sub-dimension-level deep training, not only the recognition ability of the corpus classification model at the dimension and sub-dimension levels is enhanced, but also the accuracy and practicality of the corpus classification model in the personality trait classification task is ensured.
[0100] Reference Figure 3 ,like Figure 3 The figure is a schematic diagram of the structure of the personality dictionary expansion device in the embodiment of the present application. It should be noted that: Figure 3 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0101] The device 300 may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) 310 (for example, one or more processors) and memory 320, one or more storage media 330 (for example, one or more mass storage devices) storing application programs 333 or data 332. Among them, the memory 320 and the storage medium 330 can be short-term storage or permanent storage. The program stored in the storage medium 330 may include one or more modules (not shown), each module may include a series of instruction operations in the device 300. Furthermore, the processor 310 can be configured to communicate with the storage medium 330 and execute a series of instruction operations in the storage medium on the device 300.
[0102] The device 300 may also include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input and output interfaces 360, and / or one or more operating systems 331, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, etc.
[0103] An embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is caused to execute the steps of the personality dictionary expansion method.
[0104] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system, device, or unit can refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.
[0105] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.
[0106] The foregoing is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A personality dictionary expansion method, characterized in that: The method comprises: Acquire social media corpus, classify the social media corpus using a corpus classification model, and generate annotation results based on the classification results; Second corpus words that meet a preset confidence threshold are screened out according to the annotation results, and the second corpus words are added to a pre-constructed personality dictionary to expand the pre-constructed personality dictionary.
2. The personality dictionary expansion method according to claim 1, characterized in that: The classifying the social media corpus using the corpus classification model includes: The social media corpus is segmented to obtain first corpus words, and the first corpus words are input into a corpus classification model for dimension and sub-dimension classification.
3. The personality dictionary expansion method according to claim 1, characterized in that: The method for constructing the corpus classification model includes: Randomly selecting a number of seed words of personality dimensions and their sub-dimensions from a pre-constructed personality dictionary, wherein the personality dictionary includes vocabulary of the Big Five personality dimensions and their sub-dimensions; Obtaining a preset fixed sentence template, filling the seed word into the fixed sentence template, and generating a standard sentence pattern; Using a pre-built generative model, generating training corpus based on the standard sentence pattern; Using a word segmentation tool to perform word segmentation processing on the training corpus to obtain a plurality of words, and assigning a label to each word to obtain a labeled corpus set, wherein the label includes dimensions and sub-dimensions corresponding to the words; The annotated corpus is input into a pre-trained language model, and dimension-level training and sub-dimension-level training are performed on the pre-trained language model to obtain a corpus classification model.
4. The personality dictionary expansion method according to claim 1, characterized in that: The pre-built generative model includes a rule-based generator, a neural network generative model, or a natural language processing model; The pre-trained language model includes a multi-layer neural network model based on deep learning.
5. The personality dictionary expansion method according to claim 3, characterized in that: The deep learning-based multi-layer neural network model includes a bidirectional encoder-representation converter model.
6. The personality dictionary expansion method according to claim 5, characterized in that: The method uses a bidirectional encoder to represent the converter model for dimension-level training, and the loss function L for dimension-level training is dimension include: Among them, y i is the true label, is the predicted label and N is the number of samples.
7. The personality dictionary expansion method according to claim 6, characterized in that: The method uses multiple independent bidirectional encoder representation converter models to train each sub-dimension, and the training loss function L at the sub-dimension level is sup-dimension include: Among them, z j is the true label of the sub-dimension, is the predicted label, and M is the number of sub-dimension samples.
8. A personality dictionary expansion device, characterized in that: The device comprises: A classification module, wherein the classification module obtains social media corpus, classifies the social media corpus using a corpus classification model, and generates a labeling result according to the classification result; An expansion module is configured to select second corpus words that meet a preset confidence threshold according to the annotation results, and add the second corpus words to a pre-constructed personality dictionary to expand the pre-constructed personality dictionary.
9. A personality dictionary expansion device, characterized in that: comprising a memory and at least one processor, wherein the memory has computer-readable instructions stored therein; The at least one processor calls the computer-readable instructions in the memory to execute each step of the personality dictionary expansion method as claimed in any one of claims 1 to 7.
10. A computer-readable storage medium having computer-readable instructions stored thereon, characterized in that: When the computer-readable instructions are executed by a processor, the steps of the personality dictionary expansion method according to any one of claims 1 to 7 are implemented.