Method, apparatus, electronic device, and readable medium for generating text topic information
By generating entity phrases and descriptive phrases, a multi-level text theme model was established, and the problem of homogeneity of text theme information in the existing technology was solved, the fine-grained expression of text theme information and the accurate grasp of core semantics were achieved, and the accuracy of analysis and release was improved.
Patent Information
- Application Number
- CN202011145852.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-23
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2040-10-23
AI Technical Summary
The text theme information generated by the prior art is homogeneous, and each topic contains the same common vocabulary, which leads to the inconspicuous semantics of the topic, making it difficult to reflect text information at different levels in a fine-grained manner, which in turn affects the accuracy of network public opinion analysis or message release.
By obtaining a word set with part-of-speech, generating entity phrases and descriptive phrases, removing common high-frequency words, retaining special words, establishing a text theme model at two levels of entity words and descriptive words, generating entity topic information sets and descriptive topic information sets, and combining them to generate distinct text semantics.
The problem of homogeneity of text theme information is solved, and the generated text theme information is more distinct, which can reflect text information at different levels in a fine-grained manner, improving the accuracy of network public opinion analysis and message release.
Smart Images

Figure CN114492388B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technologies, and particularly to methods, apparatuses, electronic devices, and readable media for generating text topic information. Background Art
[0002] Text topic information refers to a type of topic information that can express the semantics of a text after quantifying each word in the text. Currently, common methods for generating text topic information often represent the text as probability values of several text topics under corresponding mathematical distributions.
[0003] However, when generating text topic information in the above manner, the following technical problems often exist:
[0004] First, the topics are homogeneous, and each topic will have the same common words, resulting in the semantics of the topics not being distinct enough.
[0005] Second, the generated text topics are general topics, and it is difficult to reflect text information at different levels in a fine-grained manner. Furthermore, it is difficult to accurately grasp the core content expressed by the text. As a result, the accuracy of network public opinion analysis or message publishing is relatively low. Summary of the Invention
[0006] The content part of the present disclosure is used to briefly introduce concepts, which will be described in detail in the specific implementation part later. The content part of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0007] Some embodiments of the present disclosure propose methods, apparatuses, electronic devices, and readable media for generating text topic information to solve one or more of the technical problems mentioned in the above background art part.
[0008] In a first aspect, some embodiments of the present disclosure provide a method for generating text topic information, the method including: obtaining a set of words with part-of-speech tags. Generating entity phrase groups and description phrase groups based on the set of words with part-of-speech tags. Generating an entity topic information set and a description topic information set based on the entity phrase groups and the description phrase groups. Generating a first entity topic information set and a first description topic information set based on the entity topic information set and the description topic information set. Combining the first entity topic information set and the first description topic information set to generate a first word segment set.
[0009] Second aspect, some embodiments of the present disclosure provide an apparatus for generating text theme information. The apparatus includes: an acquisition unit configured to acquire a set of words with part-of-speech tags; a first generation unit configured to generate entity phrase groups and description phrase groups based on the set of words with part-of-speech tags; a second generation unit configured to generate an entity theme information set and a description theme information set based on the entity phrase groups and the description phrase groups; a third generation unit configured to generate a first entity theme information set and a first description theme information set based on the entity theme information set and the description theme information set; and a combination unit configured to combine the first entity theme information set and the first description theme information set to generate a first word segment set.
[0010] Third aspect, some embodiments of the present disclosure provide an electronic device, including: one or more processors; a storage device storing one or more programs thereon, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method described in any implementation manner of the first aspect above.
[0011] Fourth aspect, some embodiments of the present disclosure provide a computer-readable medium storing a computer program thereon, wherein when the program is executed by a processor, the method described in any implementation manner of the first aspect above is implemented.
[0012] The above various embodiments of the present disclosure have the following beneficial effects: First, acquire a set of words with part-of-speech tags. By acquiring the set of words with part-of-speech tags, data support is provided for subsequent generation of the first word segment set. Second, based on the set of words with part-of-speech tags, generate entity phrase groups and description phrase groups. While generating the entity phrase groups and description phrase groups, common high-frequency words are removed from the set of words with part-of-speech tags, and characteristic words that can reflect the text theme are retained, providing data support for subsequent generation of a text theme with distinct features. Then, based on the entity phrase groups and the description phrase groups, generate an entity theme information set and a description theme information set. The characteristic words that can reflect the text theme are input into the entity theme and description theme models to generate distinct entity themes and description themes. Then, based on the entity theme information set and the description theme information set, generate a first entity theme information set and a first description theme information set. By selecting the more core and key themes from the distinct entity themes and description themes, data support is provided for subsequent generation of core and distinct text semantics. Finally, combine the first entity theme information set and the first description theme information set to generate a first word segment set. The distinct and core entity themes and description themes are combined to generate a first word segment as the core and distinct semantics to be expressed by the text. Thus, the problem that the themes are homogeneous and there are the same common words in each theme, resulting in the semantics of the themes not being distinct enough, is solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and the elements and elements are not necessarily drawn to scale.
[0014] Figure 1 is a schematic diagram of an application scenario of a method for generating text theme information according to some embodiments of the present disclosure;
[0015] Figure 2 is a flowchart of some embodiments of a method for generating text theme information according to the present disclosure;
[0016] Figure 3 is a schematic structural diagram of some embodiments of an apparatus for generating text theme information according to the present disclosure;
[0017] Figure 4 is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. Specific Embodiments
[0018] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0019] In addition, it should be noted that for the sake of convenience of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.
[0020] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence relationship of the functions performed by these devices, modules or units.
[0021] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0022] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0023] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0024] Figure 1 It is a schematic diagram of an application scenario of a method for generating text topic information according to some embodiments of the present disclosure.
[0025] In Figure 1 In the application scenario diagram, first, the computing device 101 can obtain a word set 102 with part-of-speech tags. Second, the computing device 101 can generate an entity phrase set 103 and a description phrase set 104 based on the above word set 102 with part-of-speech tags. Then, the computing device 101 can generate an entity topic information set 105 and a description topic information set 106 based on the above entity phrase set 103 and the above description phrase set 104. Then, the computing device 101 can generate a first entity topic information set 107 and a first description topic information set 108 based on the above entity topic information set 105 and the above description topic information 106. Finally, the computing device 101 can combine the above first entity topic information set 107 and the above first description topic information set 108 to generate a first word segment set 109. Optionally, the computing device 101 can send the above first word segment set 109 to the target display device 110 and display the above first word segment set 109 on the above target display device 110.
[0026] It should be noted that the above computing device 101 can be hardware or software. When the computing device is hardware, it can be implemented as a distributed cluster composed of multiple servers or terminal devices, or can be implemented as a single server or a single terminal device. When the computing device is embodied as software, it can be installed in the above-listed hardware devices. It can be implemented as, for example, multiple software and software modules for providing distributed services, or can also be implemented as a single software or software module. No specific limitation is made here.
[0027] It should be understood that Figure 1 the number of computing devices in
[0028] Continuing to refer to Figure 2 , a flowchart 200 of some embodiments of a method for generating text topic information according to the present disclosure is shown. This method can be executed by the computing device 101 in Figure 1 . The method for generating text topic information includes the following steps:
[0029] Step 201, obtain a word set with part-of-speech tags.
[0030] In some embodiments, the execution subject of the method for generating text topic information (such as Figure 1The computing device 101 shown can obtain a set of words with part-of-speech through a wired connection or a wireless connection. Among them, the words with part-of-speech can be descriptive words, verbs, nouns, or words of other parts of speech.
[0031] As an example, the set of words with part-of-speech can be ["weather (noun)", "Qingdao (noun)", "newly added (verb)", "asymptomatic (descriptive word)", "rising (verb)", "reducing (verb)"].
[0032] Step 202: Generate entity phrases and descriptive phrases based on the set of words with part-of-speech.
[0033] In some embodiments, the execution subject can generate entity phrases and descriptive phrases based on the set of words with part-of-speech in various ways.
[0034] In some optional implementation manners of some embodiments, the execution subject generates entity phrases and descriptive phrases based on the set of words with part-of-speech, which may include the following steps:
[0035] The first step: Select the words with part-of-speech that meet the first preset condition from the set of words with part-of-speech as the first entity words to obtain the first entity phrase.
[0036] Among them, the first preset condition can be that the word with part-of-speech is a noun.
[0037] As an example, the set of words with part-of-speech can be ["weather (noun)", "Qingdao (noun)", "newly added (verb)", "asymptomatic (descriptive word)", "rising (verb)", "reducing (verb)"]. Select the nouns from the set of words with part-of-speech as the first entity words, and the first entity phrase obtained can be ["Huawei", "Qingdao"].
[0038] The second step: Select the words with part-of-speech that meet the second preset condition from the set of words with part-of-speech as the first descriptive words to obtain the first descriptive phrase.
[0039] Among them, the second preset condition can be that the word with part-of-speech is an adjective or a verb.
[0040] As an example, the set of words with part-of-speech can be ["weather (noun)", "Qingdao (noun)", "newly added (verb)", "asymptomatic (descriptive word)", "rising (verb)", "reducing (verb)"]. Select the adjectives or verbs from the set of words with part-of-speech as the first descriptive words, and the first descriptive phrase obtained can be ["newly added", "asymptomatic", "rising", "reducing"].
[0041] The third step is to generate a first entity word probability value group and a first description word probability value group respectively based on the above-mentioned first entity word group and the above-mentioned first description word group, which may include the following sub-steps:
[0042] The first sub-step is to respectively count the word frequencies of each first entity word in the above-mentioned first entity word group and the word frequencies of each first description word in the above-mentioned first description word group based on the above-mentioned first entity word group and the above-mentioned first description word group.
[0043] The second sub-step is to input the word frequencies of each first entity word and the word frequencies of each first description word into the following formula to generate a first entity word probability value group and a first description word probability value group:
[0044]
[0045] Where i represents the serial number. p represents the word frequency of the first entity word in the word frequencies of each first entity word. pi represents the word frequency of the i-th first entity word in the word frequencies of each first entity word. j represents the serial number. q represents the word frequency of the first description word in the word frequencies of each first description word. qj represents the word frequency of the j-th first description word in the word frequencies of each first description word. m represents the number of the word frequencies of each first description word. m represents the number of the word frequencies of each first entity word.
[0046] As an example, the word frequencies of each first entity word may be [2, 7, 5, 2, 3, 5]. The word frequencies of each first description word may be [3, 9, 2, 6, 4, 8]. Then the first entity word probability value group may be [2 / 27, 7 / 27, 5 / 27, 2 / 27, 3 / 27, 5 / 27]. The first description word probability value group may be [3 / 32, 9 / 32, 1 / 16, 3 / 16, 1 / 8, 1 / 4].
[0047] The fourth step is to select a predetermined number of first entity words from the above-mentioned first entity word group in ascending order of the first entity word probabilities in the above-mentioned first entity word probability values to form an entity word group.
[0048] As an example, the above-mentioned first entity word group may be ["influenza", "Qingdao", "doctor", "epidemic situation", "vaccine", "nurse"]. The first entity word probability value group corresponding to the above-mentioned first entity word group may be [3%, 10%, 4%, 80%, 2%, 1%]. The predetermined number may be 5. Selecting 5 first entity words from the above-mentioned first entity word group in ascending order of the first entity word probabilities in the above-mentioned first entity word probability values to form an entity word group may be [1%, 2%, 3%, 4%, 10%].
[0049] Step 5: Select a predetermined number of first descriptive words from the above first descriptive word group in ascending order of the first descriptive word probability values in the above first descriptive word probability value group to form a descriptive word group.
[0050] As an example, the above first descriptive word group can be ["prevention", "strict control", "severe", "major", "asymptomatic", "infection"]. The corresponding first descriptive word probability value group of the above first descriptive word group can be [3%, 10%, 4%, 80%, 2%, 1%]. The above predetermined number can be 5. Selecting 5 first descriptive words from the above first descriptive word group in ascending order of the first descriptive word probability values in the above first descriptive word probability value group to form a descriptive word group can be [1%, 2%, 3%, 4%, 10%].
[0051] Step 203: Generate an entity topic information set and a descriptive topic information set based on the entity word group and the descriptive word group.
[0052] In some embodiments, the above execution subject can generate an entity topic information set and a descriptive topic information set based on the above entity word group and the above descriptive word group in various ways.
[0053] In some optional implementation manners of some embodiments, the above execution subject can generate an entity topic information set and a descriptive topic information set based on the above entity word group and the above descriptive word group, including the following steps:
[0054] First step: Input the above entity word group into an entity word topic model to generate an entity topic information set.
[0055] Among them, the above entity topic information includes at least one of the following: the name of the entity topic, the probability value corresponding to the name of the entity topic. The above entity word topic model can be an LDA (Latent Dirichlet Allocation) topic model.
[0056] As an example, the above entity word group can be ["university", "teacher", "course", "market", "enterprise", "finance", "high-speed rail", "automobile", "airplane"]. Set the number of topics of the above entity word topic model to 3. Then input the above entity word group into the entity word topic model, and the generated entity topic information set can be [[entity topic 1, 10%], [entity topic 2, 20%], [entity topic 3, 70%]].
[0057] Second step: Input the above descriptive word group into a descriptive word topic model to generate a descriptive topic information set.
[0058] Among them, the above-described topic information includes at least one of the following: the name of the description topic, and the probability value corresponding to the name of the description topic. The above-described descriptor topic model may be a PLSA (Probabilistic Latent Semantic Analysis) topic model.
[0059] As an example, the above-described descriptor group may be [abundant, substantial, positive, rise, surge, competitive, take off, build, convenient]. Set the number of topics of the above-described descriptor topic model to 3. Then input the above-described descriptor group into the descriptor topic model, and the generated description topic information set may be [[description topic 1, 25%], [description topic 2, 25%], [description topic 3, 50%]].
[0060] Step 204, generate a first entity topic information set and a first description topic information set based on the entity topic information set and the description topic information set.
[0061] In some embodiments, the above execution subject may generate a first entity topic information set and a first description topic information set based on the entity topic information set and the description topic information set in various ways.
[0062] In some optional implementation manners of some embodiments, the above execution subject may generate a first entity topic information set and a first description topic information set based on the entity topic information set and the description topic information set, including the following steps:
[0063] First step, for each entity topic information in the above entity topic information set, in response to determining that the probability value corresponding to the name of the entity topic included in the entity topic information is within a first preset range, generate first entity topic information to obtain a first entity topic information set.
[0064] Among them, the above first preset range may be [10%, 100%].
[0065] As an example, the above entity topic information set may be [entity topic 1: 30%, entity topic 2: 20%, entity topic 3: 10%, entity topic 4: 3%, entity topic 5: 2%, entity topic 6: 5%, entity topic 7: 10%, entity topic 8: 20%]. Then the first entity topic information set may be [entity topic 1: 30%, entity topic 2: 20%, entity topic 3: 10%, entity topic 7: 10%, entity topic 8: 20%].
[0066] Step 2, for each piece of description topic information in the above-described set of description topic information, in response to determining that the probability value corresponding to the name of the description topic included in the above-described description topic information is within the above first preset range to generate first description word information, a first set of description word information is obtained.
[0067] Among them, the above first preset range can be [10%, 100%].
[0068] As an example, the above set of description topic information can be [Description Topic 1: 30%, Description Topic 2: 20%, Description Topic 3: 10%, Description Topic 4: 3%, Description Topic 5: 2%, Description Topic 6: 5%, Description Topic 7: 10%, Description Topic 8: 20%]. Then the first set of description topic information can be [Description Topic 1: 30%, Description Topic 2: 20%, Description Topic 3: 10%, Description Topic 7: 10%, Description Topic 8: 20%].
[0069] Step 205, combine the first set of entity topic information and the first set of description topic information to generate a first set of word segments.
[0070] In some embodiments, the above execution subject can generate a first set of word segments based on the above first set of entity topic information and the above first set of description topic information in various ways.
[0071] In some alternative implementation manners of some embodiments, the above execution subject combines the above first set of entity topic information and the above first set of description topic information to generate a first set of word segments, which may include the following steps:
[0072] First step, merge the above first set of entity topic information with the above first set of description topic information to obtain a set of combined topics.
[0073] Among them, each piece of first entity topic information in the above first set of entity topic information can be pairwise merged with each piece of first description topic information in the above first set of description topic information to generate a combined topic, and a set of combined topics is obtained.
[0074] As an example, the above first set of entity topic information can be [A, B, C]. The above first set of description topic information can be [D, E, F]. Pairwise merge each piece of first entity topic information in the above first set of entity topic information with each piece of first description topic information in the above first set of description topic information to generate a combined topic, and the set of combined topics obtained can be [AD, AE, AF, BD, BE, BF, CD, CE, CF].
[0075] Second step, determine the frequency and probability value of each combined topic in the above set of combined topics.
[0076] As an example, the frequencies of the respective topic combinations in the above-mentioned topic combination set can be [2, 5, 3, 1, 3, 5, 2, 1, 4]. And the probability values of the respective topic combinations in the above-mentioned topic combination set can be [2 / 26, 5 / 26, 3 / 26, 1 / 26, 3 / 26, 5 / 26, 2 / 26, 1 / 26, 4 / 26].
[0077] In the third step, based on the frequencies and probability values of the respective topic combinations in the above-mentioned topic combination set, the following formula is used to generate a topic co-occurrence coefficient set:
[0078]
[0079] where m represents the serial number. w represents the topic co-occurrence coefficient corresponding to the topic combination in the above-mentioned topic combination set. w m represents the topic co-occurrence coefficient corresponding to the m-th topic combination in the above-mentioned topic combination set. Z represents the probability value of the topic combination in the above-mentioned topic combination set. Z m represents the probability value of the m-th topic combination in the above-mentioned topic combination set. f represents the frequency of the topic combination in the above-mentioned topic combination set. f m represents the frequency of the m-th topic combination in the above-mentioned topic combination set. M represents the number of topic combinations in the above-mentioned topic combination set.
[0080] As an example, the above-mentioned topic combination set can be [AD, AE, AF]. The frequencies of the above-mentioned topic combination set can be [1, 2, 5]. The probability values of the above-mentioned topic combination set can be [1 / 8, 1 / 4, 5 / 8]. Then the generated topic co-occurrence coefficient set can be [5.4%, 18.3%, 76.3%].
[0081] The formulas and related content in the above steps 202-205 are an inventive point of the present disclosure, which solves the second technical problem mentioned in the background art: "The generated text theme is a general theme, making it difficult to reflect text information at different levels in a fine-grained manner. Consequently, it is relatively difficult to accurately grasp the core content expressed by the text. As a result, the accuracy of online public opinion analysis or message dissemination is relatively low." The factors that often make it difficult to accurately grasp the core content expressed by the text are generally as follows: The text themes generated by general text theme models are general text themes and do not reflect the text themes at different levels. If the above factors are addressed, text information at different levels can be reflected in a fine-grained manner, and then the core content expressed by the text can be accurately grasped. To achieve this effect, the present disclosure uses nouns in the word set as entity words and adjectives or verbs as descriptive words. A text theme model is established based on two levels of entity words and descriptive words respectively, so that the text expression is more fine-grained, and then the core content expressed by the text can be accurately grasped. It solves the problem that the generated text theme is a general theme and it is difficult to reflect text information at different levels in a fine-grained manner. Consequently, it is relatively difficult to accurately grasp the core content expressed by the text.
[0082] Fourth, based on the above-mentioned co-occurrence coefficient set of themes, select the theme combinations that meet the third preset condition from the above-mentioned set of theme combinations as the first word segments to obtain the first word segment set.
[0083] Among them, the above-mentioned third preset condition can be a theme combination whose corresponding co-occurrence coefficient of themes is within a preset coefficient range. The above-mentioned preset range can be [10%, 100%].
[0084] As an example, the above-mentioned set of theme combinations can be [AD, AE, AF]. The corresponding co-occurrence coefficient set of themes for the above-mentioned set of theme combinations can be [5.4%, 18.3%, 76.3%]. The theme combinations whose corresponding co-occurrence coefficient of themes is within the preset coefficient range can be AE. It can also be AF.
[0085] In some optional implementation manners of some embodiments, the above-mentioned execution subject can send the above-mentioned first word segment set to the target display device and display the above-mentioned first word segment set on the target display device.
[0086] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: First, a word set with part-of-speech is obtained. By obtaining the word set with part-of-speech, it provides data support for subsequent generation of the first word segment set. Second, based on the above-mentioned word set with part-of-speech, entity phrases and description phrases are generated. While generating entity phrases and description phrases, common high-frequency words are removed from the above-mentioned word set with part-of-speech, and characteristic words that can reflect the text theme are retained, providing data support for subsequent generation of a text theme with distinct characteristics. Then, based on the above-mentioned entity phrases and the above-mentioned description phrases, an entity theme information set and a description theme information set are generated. The characteristic words that can reflect the text theme are input into the entity theme and description theme models to generate distinct entity themes and description themes. Then again, based on the above-mentioned entity theme information set and the above-mentioned description theme information set, a first entity theme information set and a first description theme information set are generated. By selecting the more core and key themes among the distinct entity themes and description themes, it provides data support for subsequent generation of core and distinct text semantics. Finally, the above-mentioned first entity theme information set and the above-mentioned first description theme information set are combined to generate a first word segment set. The distinct and core entity themes and description themes are combined to generate a first word segment as the core and distinct semantics to be expressed by the text. Furthermore, it solves the problem that the themes are homogeneous and there will be the same common vocabulary in each theme, resulting in the lack of distinct semantics of the theme.
[0087] Further reference Figure 3 , as an implementation of the above methods for the above-mentioned various figures, the present disclosure provides some embodiments of an apparatus for generating text theme information. These apparatus embodiments correspond to Figure 2 the above-mentioned method embodiments and the apparatus can be specifically applied to various electronic devices.
[0088] As Figure 3 shown, for an apparatus 300 for generating text theme information in some embodiments, the apparatus includes: an acquisition unit 301, a first generation unit 302, a second generation unit 303, a third generation unit 304, and a combination unit 305. Among them, the acquisition unit 301 is configured to acquire a word set with part-of-speech. The first generation unit 302 is configured to generate entity phrases and description phrases based on the above-mentioned word set with part-of-speech. The second generation unit 303 is configured to generate an entity theme information set and a description theme information set based on the above-mentioned entity phrases and the above-mentioned description phrases. The third generation unit 304 is configured to generate a first entity theme information set and a first description theme information set based on the above-mentioned entity theme information set and the above-mentioned description theme information set. The combination unit 305 is configured to combine the above-mentioned first entity theme information set and the above-mentioned first description theme information set to generate a first word segment set.
[0089] It can be understood that the various units described in the apparatus 300 correspond to the respective steps in the method described in the reference Figure 2 description. Thus, the operations, features, and beneficial effects described above for the method also apply to the apparatus 300 and the units included therein, and will not be elaborated herein.
[0090] Next, with reference to Figure 4 , which shows a schematic structural diagram of an electronic device (such as the computing device 101 in Figure 1 ) 400 suitable for implementing some embodiments of the present disclosure. Figure 4 The server shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0091] As Figure 4 shown, the electronic device 400 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 401, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the electronic device 400 are also stored. The processing device 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An input / output (I / O) interface 404 is also connected to the bus 404.
[0092] Generally, the following devices may be connected to the I / O interface 404: an input device 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc. An output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc. A storage device 408 including, for example, a magnetic tape, a hard disk, etc. And a communication device 409. The communication device 409 may allow the electronic device 400 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 4 shows the electronic device 400 having various devices, it should be understood that it is not required to implement or include all the shown devices. Instead, more or fewer devices may be implemented or included. Figure 4 Each block shown in
[0093] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such some embodiments, the computer program can be downloaded and installed from the network through the communication device 409, or installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above functions defined in the methods of some embodiments of the present disclosure are executed.
[0094] It should be noted that the computer-readable medium in some embodiments of the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program codes are carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program codes contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0095] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0096] The computer-readable medium described above can be included in the above-mentioned device. It can also exist separately without being assembled into the electronic device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device is caused to: obtain a set of words with part-of-speech tags. Based on the set of words with part-of-speech tags, generate entity phrases and description phrases. Based on the entity phrases and the description phrases, generate an entity topic information set and a description topic information set. Based on the entity topic information set and the description topic information set, generate a first entity topic information set and a first description topic information set. Combine the first entity topic information set and the first description topic information set to generate a first word segment set.
[0097] Computer program code for performing the operations of some embodiments of the present disclosure can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).
[0098] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0099] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes an acquisition unit, a first generation unit, a second generation unit, a third generation unit, and a combination unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases. For example, the acquisition unit can also be described as "the unit for acquiring a set of words with part-of-speech".
[0100] The functions described above can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0101] The above description is only some preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, technical methods formed by mutually replacing the above features with (but not limited to) technical features having similar functions disclosed in the embodiments of the present disclosure.
Claims
1. A method for generating text topic information, comprising: obtaining a set of words with part-of-speech; generating entity phrase groups and description phrase groups based on the set of words with part-of-speech; generating an entity topic information set and a description topic information set based on the entity phrase groups and the description phrase groups; generating a first entity topic information set and a first description topic information set based on the entity topic information set and the description topic information set; combining the first entity topic information set and the first description topic information set to generate a first word segment set.
2. The method according to claim 1, wherein, the method further comprises: sending the first word segment set to a target display device and displaying the first word segment set on the target display device.
3. The method according to claim 2, wherein, the generating entity phrase groups and description phrase groups based on the set of words with part-of-speech comprises: selecting words with part-of-speech that meet a first preset condition from the set of words with part-of-speech as first entity words to obtain a first entity phrase group; selecting words with part-of-speech that meet a second preset condition from the set of words with part-of-speech as first description words to obtain a first description phrase group.
4. The method according to claim 3, wherein, the method further comprises: generating a first entity word probability value group and a first description word probability value group based on the first entity phrase group and the first description phrase group respectively; selecting a predetermined number of first entity words from the first entity phrase group in ascending order of the first entity word probability values in the first entity word probability value group to form an entity phrase group; selecting a predetermined number of first description words from the first description phrase group in ascending order of the first description word probability values in the first description word probability value group to form a description phrase group.
5. The method according to claim 4, wherein, the generating an entity topic information set and a description topic information set based on the entity phrase groups and the description phrase groups comprises: inputting the entity phrase group into an entity word topic model to generate an entity topic information set, wherein the entity topic information includes at least one of the following: the name of the entity topic, the probability value corresponding to the name of the entity topic; inputting the description phrase group into a description word topic model to generate a description topic information set, wherein the description topic information includes at least one of the following: the name of the description topic, the probability value corresponding to the name of the description topic.
6. The method according to claim 5, wherein, the generating a first entity topic information set and a first description topic information set based on the entity topic information set and the description topic information set comprises: for each entity topic information in the entity topic information set, in response to determining that the probability value corresponding to the name of the entity topic included in the entity topic information is within a first preset range, generating first entity topic information to obtain a first entity topic information set; for each description topic information in the description topic information set, in response to determining that the probability value corresponding to the name of the description topic included in the description topic information is within the first preset range to generate first description word information to obtain a first description word information set.
7. The method according to claim 6, wherein, the combining the first entity topic information set and the first description topic information set to generate a first word segment set includes: merging the first entity topic information set and the first description topic information set to obtain a topic combination set; determining the frequency and probability value of each topic combination in the topic combination set; generating a topic co-occurrence coefficient set based on the frequency and probability value of each topic combination in the topic combination set through the following formula: Among them, m represents the serial number, w represents the co-occurrence coefficient of the themes corresponding to the theme combination in the theme combination set, w m represents the co-occurrence coefficient of the themes corresponding to the m-th theme combination in the theme combination set, Z represents the probability value of the theme combination in the theme combination set, z m represents the probability value of the m-th theme combination in the theme combination set, f represents the frequency of the theme combination in the theme combination set, f m represents the frequency of the m-th theme combination in the theme combination set, M represents the number of theme combinations in the theme combination set; selecting, from the topic combination set, topic combinations that meet the third preset condition as the first word segments based on the topic co-occurrence coefficient set to obtain a first word segment set.
8. An apparatus for generating text topic information, comprising: an acquisition unit configured to acquire a word set with part-of-speech; a first generation unit configured to generate entity word groups and description word groups based on the word set with part-of-speech; a second generation unit configured to generate an entity topic information set and a description topic information set based on the entity word groups and the description word groups; a third generation unit configured to generate a first entity topic information set and a first description topic information set based on the entity topic information set and the description topic information set; a combination unit configured to combine the first entity topic information set and the first description topic information set to generate a first word segment set.
9. An electronic device, comprising: one or more processors; a storage device storing one or more programs thereon; when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the method according to any one of claims 1-7.
10. A computer-readable medium having a computer program stored thereon, wherein, the program, when executed by a processor, implements the method according to any one of claims 1-7.
Citation Information
Patent Citations
Short text clustering analysis method, device and terminal device
CN109299280A
Block chain project hot word generation method and device
CN110334268A