A hashtag upgrading method and device, a storage medium and an electronic device
By generating a phrase library and using a co-occurrence matrix to calculate the weight values of words and phrases, the topic tags are upgraded, solving the problem that topic tags are not specific enough in the existing technology, and achieving a higher recognition rate and more accurate text meaning recognition.
Patent Information
- Application Number
- CN202210857348.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-20
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2042-07-20
AI Technical Summary
In existing technologies, topic tags are not specific enough when they represent topics using words, resulting in low recognition rates and an inability to accurately identify the meaning of the text.
By generating a phrase library composed of words and phrases, and using a co-occurrence matrix to calculate the weight values of words and phrases, phrases that appear frequently in historical real dialogues are selected and output as upgraded topic tags.
It improves the recognition rate of topic tags, making them more specific and able to more accurately reflect the meaning of the text, thus helping chatbots to perform intelligent dialogue analysis.
Smart Images

Figure CN115329750B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of natural language processing, and in particular to a topic label upgrading method and device, a storage medium and an electronic device. BACKGROUND
[0002] With the development of artificial intelligence technology, dialogue robots have been applied in various fields such as catering, education, finance, etc. Among them, natural language processing technology is an important direction in the field of artificial intelligence. In natural language processing technology, a large amount of text data needs to be processed and analyzed first to determine the meaning expressed by the text.
[0003] At present, the main method is to determine some important keywords in the text to represent the main meaning of the text. The set composed of these keywords is called a topic label. For example, the existing LDA document topic generation model contains a three-layer structure of words, topics and documents, determines a topic through the probability of word occurrence, and determines the main meaning of the document through the probability of topic occurrence.
[0004] However, in the prior art, the topic is represented by words, and the topic label composed of these topics is not specific enough and has a low recognition rate. That is, the presented keywords are scattered and cannot be identified by the combination of keywords at a glance. Therefore, how to improve the recognition rate of the topic label has become a technical problem to be solved. SUMMARY
[0005] The embodiments of the present application provide a topic label upgrading method and device, a storage medium and an electronic device to solve the technical problem of the prior art that the topic label is not specific enough and has a low recognition rate.
[0006] In a first aspect, a topic label upgrading method is provided, comprising:
[0007] obtaining a target text and inputting the target text into a pre-established topic label generation model to output a topic label of the target text, wherein the topic label generation model is a model formed according to a classification algorithm;
[0008] generating a phrase library according to the output topic label of the target text, wherein the phrase library includes a plurality of words and phrases composed of words;
[0009] constructing a co-occurrence matrix according to the plurality of words in the phrase library, wherein the co-occurrence matrix is used to calculate the weight values of the plurality of words and phrases in the phrase library;
[0010] upgrading the topic label according to the phrase library and the co-occurrence matrix, wherein the phrase library and the co-occurrence matrix are used to generate high-frequency phrases.
[0011] In a possible implementation, the establishment of the topic label generation model comprises:
[0012] a predefined number of topics in the topic label generation model;
[0013] obtaining a first text database of the topic label generation model according to the predefined number of topics;
[0014] preprocessing the first text database and generating a vocabulary of the first text database, wherein the vocabulary comprises a plurality of words;
[0015] training and constructing feature information of each topic according to a classification algorithm based on the vocabulary.
[0016] In a possible implementation, the topic label comprises a plurality of topics, and the generating of a phrase library according to the output topic label of the target text comprises:
[0017] obtaining a second text database for topic label upgrading according to the plurality of topics in the topic label;
[0018] preprocessing the second text database and generating a phrase library of the second text database, wherein the phrase library comprises a plurality of words and phrases composed of the words.
[0019] In a possible implementation, the upgrading of the topic label according to the phrase library and the co-occurrence matrix comprises:
[0020] calculating weight values of the plurality of words in the phrase library according to the phrase library and the co-occurrence matrix;
[0021] calculating weight values of a plurality of phrases in the phrase library according to the weight values of the plurality of words in the phrase library;
[0022] generating the upgraded topic label according to the weight values of the plurality of words and phrases in the phrase library.
[0023] In a possible implementation, the calculating of the weight values of the plurality of words in the phrase library according to the phrase library and the co-occurrence matrix comprises:
[0024] sequentially calculating a number of times each word in the co-occurrence matrix co-occurs;
[0025] sequentially judging whether each word in the co-occurrence matrix is a word in a phrase, and if so, obtaining a length of the phrase;
[0026] calculating the weight values of the plurality of words in the phrase library according to the number of times each word in the co-occurrence matrix co-occurs and the length of the phrase.
[0027] In a possible implementation, the second text database for the topic label upgrade is derived from the first text database of the topic label generation model, and is used to generate a phrase library with high matching degree with the topic label.
[0028] In a possible implementation, the generating the upgraded topic label according to the weight values of the words and phrases in the phrase library comprises:
[0029] a predefined topic label upgrade threshold;
[0030] According to the topic label upgrade threshold, it is determined whether the phrases and words in the phrase library meet the conditions of the topic label upgrade threshold in sequence, and if the conditions are met, the upgraded topic label is determined.
[0031] In a second aspect, a topic label upgrade device is provided, comprising:
[0032] a topic label output module configured to obtain a target text and input the target text into a pre-established topic label generation model to output a topic label of the target text, wherein the topic label generation model is a model formed according to a classification algorithm;
[0033] a phrase library generation module configured to generate a phrase library according to the output topic label of the target text, wherein the phrase library comprises a plurality of words and phrases composed of the words;
[0034] a co-occurrence matrix construction module configured to construct a co-occurrence matrix according to the plurality of words in the phrase library, wherein the co-occurrence matrix is used to calculate weight values of the plurality of words and phrases in the phrase library;
[0035] a topic label upgrade module configured to upgrade the topic label according to the phrase library and the co-occurrence matrix, wherein the phrase library and the co-occurrence matrix are used to generate high-frequency phrases.
[0036] In a possible implementation, the topic label output module is further configured to establish the topic label generation model, comprising:
[0037] a predefined number of topics in the topic label generation model;
[0038] obtaining a first text database of the topic label generation model according to the predefined number of topics;
[0039] preprocessing the first text database and generating a vocabulary library of the first text database, wherein the vocabulary library comprises a plurality of words;
[0040] training and constructing feature information of each topic according to a classification algorithm based on the vocabulary library.
[0041] In a possible implementation, the topic label includes a plurality of topics, and the phrase library generation module is further configured to:
[0042] According to the plurality of topics in the topic label, a second text database for topic label upgrading is obtained;
[0043] The second text database is preprocessed, and a phrase library of the second text database is generated, where the phrase library includes a plurality of words and phrases composed of the words.
[0044] In a possible implementation, the topic label upgrading module is further configured to:
[0045] According to the phrase library and the co-occurrence matrix, weight values of the plurality of words in the phrase library are calculated;
[0046] According to the weight values of the plurality of words in the phrase library, weight values of a plurality of phrases in the phrase library are calculated;
[0047] According to the weight values of the plurality of words and phrases in the phrase library, an upgraded topic label is generated.
[0048] In a possible implementation, the topic label upgrading module is further configured to:
[0049] The number of times each word co-occurs in the co-occurrence matrix is sequentially calculated;
[0050] It is sequentially determined whether each word in the co-occurrence matrix is a word in a phrase, and if so, the length of the phrase is obtained;
[0051] According to the number of times each word co-occurs in the co-occurrence matrix and the length of the phrase, weight values of the plurality of words in the phrase library are calculated.
[0052] In a possible implementation, the topic label upgrading module is further configured to:
[0053] A topic label upgrading threshold is predefined;
[0054] According to the topic label upgrading threshold, it is sequentially determined whether the phrases and words in the phrase library satisfy the conditions of the topic label upgrading threshold, and if so, the phrases and words are determined as the upgraded topic label.
[0055] In a third aspect, a computer readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the method for upgrading a topic label are implemented.
[0056] In a fourth aspect, an electronic device is provided, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the upgrading method of the subject tag when executing the computer program.
[0057] The upgrading method of the subject tag, the device, the storage medium, and the electronic device improve the recognition rate of the subject tag, and solve the problem that the subject tag is not specific and the recognition rate is not high when the subject is determined by only words in the prior art. BRIEF DESCRIPTION OF DRAWINGS
[0058] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0059] Figure 1 is a flowchart of the upgrading method of the subject tag in an embodiment of the present application;
[0060] Figure 2 is a flowchart of the method for establishing the subject tag generation model in an embodiment of the present application;
[0061] Figure 3 is Figure 1 is a flowchart of a specific embodiment of step S40 in the method;
[0062] Figure 4 is a structural diagram of the upgrading device of the subject tag in an embodiment of the present application;
[0063] Figure 5 is a structural diagram of the computer device in an embodiment of the present application. DETAILED DESCRIPTION
[0064] The technical solutions of the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0065] Before the embodiments of the present application are explained in detail, the application scenarios involved in the embodiments of the present application are introduced.
[0066] With the development of artificial intelligence technology, the dialogue robot is applicable to the fields of catering, education, finance and the like. For example, in the field of finance, a large amount of dialogue data is stored in the server of the dialogue robot. For the above dialogue data, the meaning of each sentence can be determined in the form of a series of short words through processing and analysis of the data. For example, in the dialogue data, a dialogue is: “Please click my policy from the lower left corner of the public number and view the detailed content”. If a series of keywords are used as the topic labels of the sentence, the topic labels are the set of four keywords “policy”, “public number”, “content” and “lower left corner”.
[0067] However, only by using the combination of the above keywords as labels, the topic labels are still not specific enough, and the main meaning of the sentence cannot be recognized by the human eye according to the combination of the above keywords. The embodiments of the present application are implemented based on the above application scenarios. The embodiments of the present application generate a phrase library composed of words and phrases by using a text database formed by historical real dialogue records, construct a co-occurrence matrix by using each word in the phrase library, calculate the weight values of each word and phrase in the phrase library through the co-occurrence matrix, and then filter out phrases with higher occurrence frequency in the historical real dialogue, and finally output these phrases as upgraded topics, thereby improving the recognition rate of the topic labels. The topic labels of the sentence “Please click my policy from the lower left corner of the public number and view the detailed content” are upgraded to the set of two key phrases “public number lower left corner” and “click policy”.
[0068] The present application will be described in detail through specific embodiments.
[0069] Please refer to Figure 1 shown, Figure 1 is a flowchart of a topic label upgrading method in an embodiment of the present application, which includes the following steps:
[0070] S10: obtaining a target text and inputting the target text into a pre-established topic label generation model to output the topic labels of the target text, wherein the topic label generation model is a model formed according to a classification algorithm;
[0071] In the embodiment of the present application, the target text is from the text in the dialogue database of the dialogue robot, and for the convenience of understanding, the target text can be assumed to be a sentence, which is: "Please click my policy from the lower left corner of the public number and view the detailed content". After obtaining the target text, the data in the target text needs to be preprocessed, and the preprocessing can include word segmentation and stop word removal. Assuming that the word segmentation and stop word removal are performed on the sentence "Please click my policy from the lower left corner of the public number and view the detailed content", the result is "public number", "lower left corner", "click", "policy", "view", and "content". The set of the above words is input into the pre-established topic label generation model to output the topic label of the target text. The topic label generation model is a model formed according to a classification algorithm. Specifically, the establishment of the topic label generation model is described with reference to Figure 2 Figure 2 FIG. 1 is a flowchart of a method for establishing a topic label generation model in an embodiment of the present application, which includes the following steps.
[0072] S201: Predefine the number of topics in the topic label generation model.
[0073] In the embodiment of the present application, the topic label generation model is a model formed according to a classification algorithm. For example, the LDA (Latent Dirichlet Allocation) topic model can be used. In the LDA model, it is considered that a topic can be represented by a word distribution, and an article can be represented by a topic distribution. Therefore, to generate the topic label of an article, the probability of occurrence of some words can be used to determine the above-mentioned topic, and the probability of occurrence of some topics can be used to determine the topic label of the article. In the embodiment, to determine the topic label of the target text, the number of topics needs to be predefined in the topic label generation model. The number of topics can be 50 or 100, which can be flexibly defined according to the business situation, and the embodiment of the present application does not limit this.
[0074] S202: Obtain a first text database of the topic label generation model according to the predefined number of topics.
[0075] After the number of topics is predefined, a large amount of text data needs to be obtained, and then the text data is analyzed to construct the feature information of each topic. In the embodiment of the present application, the first text database of the topic label generation model can be from the data in the dialogue database of the dialogue robot.
[0076] S203: Preprocess the first text database and generate a vocabulary of the first text database, wherein the vocabulary includes a plurality of words.
[0077] Since the LDA topic is determined according to the frequency of word occurrence, the database needs to be preprocessed to form a vocabulary. The preprocessing of the database includes tokenization and stop word removal. Tokenization is the process of dividing a sentence into individual words, and stop word removal is the process of removing unnecessary words from a sentence without affecting the overall semantic understanding of the sentence. Stop words can include function words, pronouns, or verbs and nouns without specific meaning. For example, "my" in "Please click on my policy from the lower left corner of the public number and check the details" can be identified as a stop word.
[0078] S204: Based on the vocabulary, the feature information of each topic is trained and constructed in sequence according to the classification algorithm.
[0079] After the above-mentioned topic label generation model is established, the target text is input into the pre-established topic label generation model to output the topic label of the target text. Assuming that the input target text is "Please click on my policy from the lower left corner of the public number and check the details", the output topic label of the target text is the set of four keywords "policy", "public number", "content", and "lower left corner". After the topic label of the target text is output, the topic label needs to be upgraded, which requires generating a phrase library first. The steps are as follows:
[0080] S20: Generating a phrase library according to the output topic label of the target text, wherein the phrase library includes a plurality of words and phrases composed of words;
[0081] In order to filter out high-frequency phrases from the dialogue database of the dialogue robot, a phrase library needs to be generated according to the output topic label of the target text. In one embodiment, generating a phrase library according to the output topic label of the target text includes the following steps:
[0082] A1: According to the plurality of topics in the topic label, a second text database for topic label upgrading is obtained;
[0083] A2: Preprocessing the second text database and generating a phrase library of the second text database, wherein the phrase library includes a plurality of words and phrases composed of words.
[0084] It should be noted that in the process of establishing the topic label generation model, the first text database processed vocabulary library needs to be trained to build the feature information of each topic. That is, to determine the feature information of each topic, a series of text data needs to be trained. In this embodiment, it is assumed that the first text database stores 10,000 pieces of dialogue text data, and the second text database derived from the first text database stores 2,000 pieces of dialogue text data. At the same time, the 2,000 pieces of dialogue text database are dialogue text databases for "policy", "public number", "content", and "lower left corner" topic labels. That is, in the topic label generation model, the "policy", "public number", "content", and "lower left corner" topic labels can only be trained and generated based on the 2,000 pieces of dialogue text stored in the second text database.
[0085] Optionally, in order to generate a phrase library with high matching degree with the topic label, a second text database is extracted from the first text database of the topic label generation model for topic label upgrading.
[0086] After obtaining the second text database, the second text database needs to be preprocessed to generate a phrase library of the second text database. For ease of understanding, it is assumed that three pieces of dialogue text are mapped in the second text database, which are "Hello, in the lower left corner of the public number", "Yes, click on the policy, there is a detail", and "Yes, please click on the WeChat public number lower left corner". The preprocessing of the second text database includes tokenization and stop word removal of all dialogue texts mapped in the second text database. It should be noted that in this embodiment, for ease of understanding, it is assumed that three pieces of dialogue text are mapped in the second text database. In actual operation, there are thousands of dialogue texts mapped in the second text database.
[0087] Then, a phrase library of the second text database is generated according to the preprocessed words, wherein the phrase library includes a plurality of words and phrases composed of words. For example, after preprocessing, "Hello, in the lower left corner of the public number" includes the words "hello", "public number", and "lower left corner". In the phrase library, in addition to including single words, it also includes the phrase "public number lower left corner"; after preprocessing, "Yes, click on the policy, there is a detail" includes the words "click", "policy", and "detail". In the phrase library, it also includes the phrase "click policy"; after preprocessing, "Yes, please click on the WeChat public number lower left corner" includes the words "click", "WeChat", "public number", and "lower left corner". In the phrase library, it also includes the phrase "WeChat public number lower left corner".
[0088] S30: Construct a co-occurrence matrix according to the plurality of words in the phrase library, wherein the co-occurrence matrix is used to calculate the weight values of the plurality of words and phrases in the phrase library;
[0089] Assuming that the above three sentences of dialogue text are mapped in the second text database, after the three sentences are segmented and stop words are removed, a co-occurrence matrix is constructed according to a plurality of words in the phrase library. Please refer to Table 1, which is the constructed co-occurrence matrix. It should be noted that the co-occurrence matrix is constructed by each word in the phrase library. In order to facilitate understanding, only the preprocessed words of the three sentences of dialogue text are listed in this embodiment. In actual operation, the co-occurrence matrix is constructed by thousands of words.
[0090] Hello WeChat Public account Bottom left corner Click Policy Details Hello 1 WeChat 1 1 1 Public account 1 2 2 Bottom left corner 1 2 2 Click 2 1 Policy 1 1 Details 1
[0091] Table 1 Co-occurrence matrix
[0092] S40: upgrading the topic label according to the phrase library and the co-occurrence matrix, wherein the phrase library and the co-occurrence matrix are used to generate high-frequency phrases.
[0093] Please refer to Figure 3 , as shown in Figure 3 is Figure 1 A flowchart of step S40 in the first embodiment, comprising:
[0094] S41: calculating the weight values of the plurality of words in the phrase library according to the phrase library and the co-occurrence matrix.
[0095] In an embodiment, S41 can include the following steps:
[0096] B1: sequentially calculating the number of times each word in the co-occurrence matrix appears together;
[0097] For example, among the preprocessed words of the above three sentences of dialogue text, the number of times "Hello" appears is 1, the number of times "WeChat" appears is 1, the number of times "public number" appears is 2, the number of times "bottom left corner" appears is 2, the number of times "click" appears is 2, the number of times "policy" appears is 1, and the number of times "details" appears is 1.
[0098] B2: sequentially judging whether each word in the co-occurrence matrix is a word in a phrase, and if so, obtaining the length of the phrase;
[0099] For example, "public number bottom left corner" and "WeChat public number bottom left corner" are phrases in the phrase library that contain "public number". When calculating the weight value of "public number", the length of "public number bottom left corner" and "WeChat public number bottom left corner" needs to be obtained. By default, the length of each word is 1, so the lengths of "public number bottom left corner" and "WeChat public number bottom left corner" are 2 and 3 respectively.
[0100] B3: calculating the weight values of the plurality of words in the phrase library according to the number of times each word in the co-occurrence matrix appears together and the length of the phrase.
[0101] In one implementation, the weight of each word can be defined as follows: the total length of the phrase containing the word or the length of the word itself divided by the number of times the word appears. For example, when calculating the weight of "official account," since "official account" appears 2 times and the total length of the phrase containing "official account" is 2+3, the weight of "official account" is 5 / 2. Similarly, the weight of "hello" is 1 / 1, "WeChat" is 3 / 1, "bottom left corner" is 5 / 2, "click" is 3 / 2, "policy" is 2 / 1, and "details" is 1 / 1.
[0102] S42: Calculate the weight values of multiple phrases in the phrase library based on the weight values of multiple words in the phrase library;
[0103] In one implementation, the weight value of a phrase in the phrase library can be obtained by summing the weight values of each word in the phrase. For example, the weight value of "bottom left corner of WeChat official account" is 5 / 2 + 5 / 2; the weight value of "click on insurance policy" is 3 / 2 + 2 / 1; and the weight value of "bottom left corner of WeChat official account" is 5 / 2 + 5 / 2 + 3 / 1.
[0104] S43: Generate upgraded topic tags based on the weight values of multiple words and phrases in the phrase library.
[0105] In one implementation, after obtaining the weight values of multiple words and phrases in the phrase library, upgraded topic tags can be generated through the following steps:
[0106] C1: Predefined topic tag upgrade threshold;
[0107] C2: Based on the topic tag upgrade threshold, determine in turn whether the phrases and words in the phrase library meet the conditions of the topic tag upgrade threshold. If the conditions are met, they are determined as upgraded topic tags.
[0108] In one implementation, for example, a predefined threshold condition for upgrading topic tags is defined as follows: when the weight value of a word or phrase in the phrase library ranks in the top third, the word or phrase that meets the condition is determined to be the upgraded topic tag. Therefore, "bottom left corner of WeChat Official Account," "bottom left corner of Official Account," and "click on the insurance policy" are determined to be upgraded topic tags because their weight values rank in the top third.
[0109] So, the topic label of the above-mentioned sentence "Please click on my policy from the lower left corner of the public number and check the detailed content" is upgraded from the set of four keywords "policy", "public number", "content", and "lower left corner" to the set of two key phrases "public number lower left corner" and "click on policy". The upgraded topic label is more specific and has more practical meaning, which can help further analysis of the intelligent dialogue of the dialogue robot.
[0110] As can be seen, in the above scheme, by using the text database formed by the historical real dialogue record, a phrase library composed of words and phrases is generated, and a co-occurrence matrix is constructed using each word in the phrase library. The weight values of each word and phrase in the phrase library are calculated through the co-occurrence matrix, and then the phrases with higher frequency of occurrence in the historical real dialogue are screened out. Finally, these phrases are output as upgraded topics, which improves the recognition rate of the topic label and solves the problem that the topic label is not specific enough and the recognition rate is not high when only words are used to determine the topic in the prior art.
[0111] It should be understood that the size of the serial number of each step in the above-mentioned embodiments does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. In addition, the term "comprising" and its variants are to be interpreted as an open term of "including but not limited to".
[0112] In an embodiment, a topic label upgrading device is provided, which corresponds to the topic label upgrading method in the above-mentioned embodiments. As shown in the figure, the upgrading device includes a topic label output module 301, a phrase library generation module 302, a co-occurrence matrix construction module 303, and a topic label upgrading module 304. The detailed description of each functional module is as follows: Figure 4
[0113] The topic label output module 301 is used to obtain a target text and input the target text into a pre-established topic label generation model to output the topic label of the target text, wherein the topic label generation model is a model formed according to a classification algorithm.
[0114] The phrase library generation module 302 is used to generate a phrase library according to the output topic label of the target text, wherein the phrase library includes a plurality of words and phrases composed of words.
[0115] The co-occurrence matrix construction module 303 is used to construct a co-occurrence matrix according to the plurality of words in the phrase library, wherein the co-occurrence matrix is used to calculate the weight values of the plurality of words and phrases in the phrase library.
[0116] The topic label upgrading module 304 is configured to upgrade the topic label according to the phrase library and the co-occurrence matrix, wherein the phrase library and the co-occurrence matrix are used to generate high-frequency phrases.
[0117] In an embodiment, the topic label output module 301 is further configured to establish a topic label generation model, including:
[0118] predefining the number of topics in the topic label generation model;
[0119] obtaining a first text database of the topic label generation model according to the predefined number of topics;
[0120] preprocessing the first text database and generating a vocabulary library of the first text database, wherein the vocabulary library includes a plurality of words;
[0121] based on the vocabulary library, training and constructing feature information of each topic in sequence according to a classification algorithm.
[0122] In an embodiment, the topic label includes a plurality of topics, and the phrase library generation module 302 is further configured to:
[0123] obtaining a second text database for topic label upgrading according to the plurality of topics in the topic label;
[0124] preprocessing the second text database and generating a phrase library of the second text database, wherein the phrase library includes a plurality of words and phrases composed of the words.
[0125] In an embodiment, the topic label upgrading module 304 is further configured to:
[0126] calculating weight values of the plurality of words in the phrase library according to the phrase library and the co-occurrence matrix;
[0127] calculating weight values of a plurality of phrases in the phrase library according to the weight values of the plurality of words in the phrase library;
[0128] generating an upgraded topic label according to the weight values of the plurality of words and phrases in the phrase library.
[0129] In an embodiment, the topic label upgrading module 304 is further configured to:
[0130] calculating the number of co-occurrences of each word in the co-occurrence matrix in sequence;
[0131] judging whether each word in the co-occurrence matrix is a word in a phrase, and if so, obtaining the length of the phrase;
[0132] calculating the weight values of the plurality of words in the phrase library according to the number of co-occurrences of each word in the co-occurrence matrix and the length of the phrase.
[0133] In an embodiment, the topic label upgrading module 304 is further configured to:
[0134] predefine a topic label upgrading threshold;
[0135] determine whether the phrases and words in the phrase library meet the conditions of the topic label upgrading threshold in sequence according to the topic label upgrading threshold, and if the conditions are met, determine the phrases as upgraded topic labels.
[0136] The application provides a topic label upgrading device. The device generates a phrase library composed of words and phrases by using a text database formed by historical real conversation records, constructs a co-occurrence matrix by using each word in the phrase library, calculates the weight values of each word and phrase in the phrase library through the co-occurrence matrix, further screens out phrases with a higher frequency of occurrence in the historical real conversation, and finally outputs the phrases as upgraded topics, thereby improving the recognition rate of the topic label and solving the problem that the topic label is not specific and the recognition rate is not high when only words are used to determine the topic in the prior art.
[0137] The specific limitations of the topic label upgrading device can be referred to the limitations of the topic label upgrading method in the foregoing description, and will not be described herein. Each module in the topic label upgrading device can be realized by software, hardware or a combination thereof in whole or in part. Each module can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory in the computer device in the form of software, so as to be called and executed by the processor to perform the operations corresponding to each module.
[0138] In an embodiment, a computer device is provided, and an internal structure diagram of the computer device can be as shown in Figure 5 The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is configured to communicate with an external server through a network connection. The computer program is executed by the processor to implement the functions or steps of the topic label upgrading method.
[0139] In an embodiment, a computer device is provided, and includes a memory, a processor and a computer program stored in the memory and executable on the processor. The processor implements the following steps when executing the computer program:
[0140] acquire a target text, and input the target text into a pre-established topic label generation model to output a topic label of the target text, wherein the topic label generation model is a model formed according to a classification algorithm;
[0141] generate a phrase library according to the output topic label of the target text, wherein the phrase library includes a plurality of words and phrases composed of the words;
[0142] construct a co-occurrence matrix according to the plurality of words in the phrase library, wherein the co-occurrence matrix is used to calculate weight values of the plurality of words and the phrases in the phrase library;
[0143] upgrade the topic label according to the phrase library and the co-occurrence matrix, wherein the phrase library and the co-occurrence matrix are used to generate high-frequency phrases.
[0144] In one embodiment, a computer readable storage medium is provided, which stores a computer program. The computer program is executed by a processor to implement the following steps:
[0145] acquire a target text, and input the target text into a pre-established topic label generation model to output a topic label of the target text, wherein the topic label generation model is a model formed according to a classification algorithm;
[0146] generate a phrase library according to the output topic label of the target text, wherein the phrase library includes a plurality of words and phrases composed of the words;
[0147] construct a co-occurrence matrix according to the plurality of words in the phrase library, wherein the co-occurrence matrix is used to calculate weight values of the plurality of words and the phrases in the phrase library;
[0148] upgrade the topic label according to the phrase library and the co-occurrence matrix, wherein the phrase library and the co-occurrence matrix are used to generate high-frequency phrases.
[0149] It should be noted that the functions or steps that the computer readable storage medium or the computer device can implement are described above with reference to the related description in the foregoing method embodiments, and will not be described here again to avoid repetition.
[0150] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, storage, database or other medium used in the embodiments provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0151] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of functional units and modules is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the above-described functions.
[0152] The above-mentioned embodiments are only used to illustrate the technical solutions of the present application, but not limit it. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features. Such modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. A method for upgrading topic tags, characterized in that, include: The target text is obtained and input into a pre-established topic tag generation model, which outputs topic tags for the target text. The topic tag generation model is a model formed based on a classification algorithm, and the topic tags include multiple topics. Based on multiple topics in the topic tags, a second text database for topic tag upgrade is obtained; the second text database is preprocessed, and a phrase library for the second text database is generated, wherein the phrase library includes: multiple words and phrases composed of words, and the second text database is extracted from the first text database of the topic tag generation model; A co-occurrence matrix is constructed based on multiple words in the phrase library, wherein the co-occurrence matrix is used to calculate the weight values of multiple words and phrases in the phrase library; The frequency of each word in the co-occurrence matrix is calculated sequentially; each word in the co-occurrence matrix is then determined to be a word in a phrase, and if so, the length of the phrase is obtained; based on the frequency of each word in the co-occurrence matrix and the length of the phrase, the weight values of multiple words in the phrase library are calculated; based on the weight values of multiple words in the phrase library, the weight values of multiple phrases in the phrase library are calculated; and based on the weight values of multiple words and phrases in the phrase library, upgraded topic tags are generated, wherein the phrase library and the co-occurrence matrix are used to generate high-frequency phrases.
2. The method according to claim 1, characterized in that, The establishment of the topic tag generation model includes: The number of topics in the predefined topic tag generation model; Based on the predefined number of topics, obtain the first text database of the topic tag generation model; The first text database is preprocessed to generate a vocabulary of the first text database, wherein the vocabulary includes: multiple words; Based on the vocabulary, feature information for each topic is constructed by sequentially training according to the classification algorithm.
3. The method according to claim 1, characterized in that, The second text database for the topic tag upgrade is derived from the first text database of the topic tag generation model and is used to generate a phrase library with a high degree of matching with the topic tags.
4. The method according to claim 1, characterized in that, The step of generating upgraded topic tags based on the weight values of multiple words and phrases in the phrase library includes: Predefined topic tag upgrade thresholds; Based on the topic tag upgrade threshold, the phrases and words in the phrase library are sequentially determined to meet the conditions for the topic tag upgrade threshold. If the conditions are met, they are determined to be upgraded topic tags.
5. A device for upgrading topic tags, characterized in that, include: Topic Tag Output Module: Used to acquire target text, input the target text into a pre-established topic tag generation model, and output topic tags for the target text. The topic tag generation model is a model formed based on a classification algorithm, and the topic tags include multiple topics. The phrase library generation module is used to obtain a second text database of topic tags based on multiple topics in the topic tags; preprocess the second text database and generate a phrase library of the second text database, wherein the phrase library includes multiple words and phrases composed of words, and the second text database is extracted from the first text database of the topic tag generation model; Co-occurrence matrix construction module: used to construct a co-occurrence matrix based on multiple words in the phrase library, wherein the co-occurrence matrix is used to calculate the weight values of multiple words and phrases in the phrase library; The topic tag upgrade module is used to: sequentially calculate the frequency of each word in the co-occurrence matrix; sequentially determine whether each word in the co-occurrence matrix belongs to a phrase, and if so, obtain the length of the phrase; calculate the weight values of multiple words in the phrase library based on the frequency of each word in the co-occurrence matrix and the length of the phrase; calculate the weight values of multiple phrases in the phrase library based on the weight values of multiple words in the phrase library; and generate upgraded topic tags based on the weight values of multiple words and phrases in the phrase library, wherein the phrase library and co-occurrence matrix are used to generate high-frequency phrases.
6. The apparatus according to claim 5, characterized in that, The topic tag output module is also used to establish a topic tag generation model, including: The number of topics in the predefined topic tag generation model; Based on the predefined number of topics, obtain the first text database of the topic tag generation model; The first text database is preprocessed to generate a vocabulary of the first text database, wherein the vocabulary includes: multiple words; Based on the vocabulary, feature information for each topic is constructed by sequentially training according to the classification algorithm.
7. The apparatus according to claim 5, characterized in that, The topic tag upgrade module is also used for: Predefined topic tag upgrade thresholds; Based on the topic tag upgrade threshold, the phrases and words in the phrase library are sequentially determined to meet the conditions for the topic tag upgrade threshold. If the conditions are met, they are determined to be upgraded topic tags.
8. A storage medium, characterized in that, The storage medium stores a computer program, wherein the computer program is configured to execute the method of any one of claims 1 to 4 when it is run.
9. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the method of any one of claims 1 to 4.
Citation Information
Patent Citations
Topic phrase extraction method
CN108920454A
Domain-independent automated processing of free-form text
WO2020131004A1