Hierarchical text topic analysis method, terminal device
Through the hierarchical text theme analysis method, the title information in the text chapter structure is used to generate a theme weight pair set, which solves the problem of low accuracy of topic analysis caused by ignoring title information in the prior art, and achieves more accurate text theme analysis.
Patent Information
- Application Number
- CN202110140866.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-02
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-02-02
AI Technical Summary
When performing text topic analysis, the prior art ignores the title information in the text chapter structure, resulting in the inability to effectively utilize multi-level text topic information, affecting the accuracy of the topic analysis results.
A hierarchical text theme analysis method is proposed. By obtaining the target text, generating a chapter text collection, determining the primary theme sequence and the secondary theme sequence, generating a mixed theme sequence, and generating a theme weight pair collection based on the chapter text collection and the mixed theme sequence.
The hierarchical title information of the text structure is used to generate a theme weight set, which improves the accuracy of the topic analysis results and can better express the theme information in the target text.
Smart Images

Figure CN114840630B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of natural language processing, and in particular, to methods and terminal devices for text topic analysis. Background Art
[0002] Topic analysis is one of the most common forms of qualitative research. It emphasizes precisely locating, examining, and recording themes or patterns in data. A theme is a pattern across datasets that is important for describing a phenomenon and is associated with a specific research question. A topic analysis model is a statistical model that clusters the implicit semantic structure of a corpus in an unsupervised learning manner. In natural language processing, topic models are used to reduce the dimensionality of text representations, cluster texts by topic, and form a text recommendation system based on user preferences.
[0003] However, when performing topic analysis in the above manner, the following technical problems often exist:
[0004] First, when there is a complex chapter structure in the text, the title of the chapter structure usually contains important text information. Existing methods directly extract the text content of the document topic for topic analysis and often ignore the information in the title, failing to effectively utilize multi-level text topic information.
[0005] Second, when the text contains a large amount of information content and the topic overlap is low, topics with low relevance to the topic are usually generated based on a large amount of text unrelated to the topic. The expression accuracy of the topic model generated under the influence of such interference factors is poor. Summary of the Invention
[0006] This part of the content of the present disclosure is used to briefly introduce concepts, which will be described in detail in the following detailed implementation part. This part of the content of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0007] Embodiments of the present disclosure propose a hierarchical text topic analysis method to solve one or more of the technical problems mentioned in the above background art part.
[0008] In a first aspect, embodiments of the present disclosure provide a hierarchical text topic analysis method, the method comprising: obtaining a target text to be processed; generating a chapter text set based on the target text; determining a first-level topic sequence and a second-level topic sequence; for each first-level topic in the first-level topic sequence, generating a mixed topic based on the first-level topic and the second-level topic sequence to obtain a mixed topic sequence; generating a set of topic weight pairs based on the chapter text set and the mixed topic sequence.
[0009] Second aspect, some embodiments of the present disclosure provide a terminal device, including: one or more processors; a storage device storing one or more programs thereon, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of the first aspect.
[0010] The above embodiments of the present disclosure have the following beneficial effects: The hierarchical text topic analysis method according to some embodiments of the present disclosure can generate a set of topic weight pairs by using the hierarchical title information of the text structure in the text, improving the accuracy of the topic analysis result. Specifically, the inventors found that the main reason for the poor accuracy of the topic analysis result in topic analysis is that: directly extracting the text content of the document topic for topic analysis, ignoring the title information in the text structure, and being unable to effectively utilize the multi-level text topic information, thus affecting the accuracy of the topic analysis result. In addition, a large amount of irrelevant text content interferes with the accuracy of the topic model expression. Based on this, first, some embodiments of the present disclosure generate a set of paragraph text word vectors based on the target text. Secondly, for each paragraph text in the paragraph text set of the target text, a set of complete title text word vectors corresponding to the paragraph text is generated to obtain a set of complete title text word vectors. The set of complete title text word vectors includes multi-level title information. Thirdly, a set of chapter text is determined according to the set of paragraph text word vectors and the set of complete title text word vectors. Then, a mixed topic sequence for topic analysis is determined. The mixed topic sequence takes into account both coarse-grained topics and fine-grained topics, and can fully utilize the text structure to mine hierarchical information. Finally, based on the set of chapter text and the mixed topic sequence, a set of topic weight pairs is generated. The topic weight pair includes a topic text and a weight. Generating a set of chapter text using the target text and introducing a mixed topic sequence can mine text topics by using the hierarchical information of the text structure in the target text, improving the accuracy of the generated set of topic weight pairs, and the topic weight pair can better express the target text. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Other features, objects, and advantages of the present disclosure will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings:
[0012] Figure 1 is an architectural diagram of an exemplary system to which some embodiments of the present disclosure can be applied;
[0013] Figure 2 is a flowchart of some embodiments of the hierarchical text topic analysis method according to the present disclosure;
[0014] Figure 3 is an exemplary authorization prompt box;
[0015] Figure 4It is a schematic structural diagram of a terminal device suitable for implementing some embodiments of the present disclosure. Detailed implementation manners
[0016] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0017] In addition, it should be noted that for the sake of convenience of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.
[0018] It should be noted that the modifiers "a" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0019] The present disclosure will be described in detail below with reference to the drawings and in combination with embodiments.
[0020] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0021] In addition, it should be noted that for the sake of convenience of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.
[0022] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or mutual dependence relationship of the functions performed by these devices, modules or units.
[0023] It should be noted that the modifiers "a" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0024] The present disclosure will be described in detail below with reference to the drawings and in combination with embodiments.
[0025] Figure 1 FIG. 100 shows an exemplary system architecture to which the hierarchical text topic analysis method of the present disclosure can be applied.
[0026] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0027] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as information generation applications, text processing applications, information extraction applications, etc.
[0028] The terminal devices 101, 102, 103 may be hardware or software. When the terminal devices 101, 102, 103 are hardware, they may be various terminal devices with a display screen, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc. When the terminal devices 101, 102, 103 are software, they may be installed in the above-listed terminal devices. It may be implemented as multiple software or software modules (such as used to provide target text input, etc.), or it may be implemented as a single software or software module. No specific limitation is made here.
[0029] The server 105 may be a server that provides various services, such as a server for storing the target data input by the terminal devices 101, 102, 103, etc. The server may process the received target work order sequence and feedback the processing result (such as the set of theme weight pairs) to the terminal device.
[0030] It should be noted that the hierarchical text topic analysis method provided in the embodiments of the present disclosure may be executed by the server 105 or by the terminal device.
[0031] It should be pointed out that the target text may also be directly stored locally in the server 105. The server 105 may directly extract the target text locally and obtain the set of theme weight pairs through processing. At this time, the exemplary system architecture 100 may not include the terminal devices 101, 102, 103 and the network 104.
[0032] It should also be noted that the hierarchical text topic analysis method application can also be installed in the terminal devices 101, 102, and 103. In this case, the processing method can also be executed by the terminal devices 101, 102, and 103. In this case, the exemplary system architecture 100 may not include the server 105 and the network 104.
[0033] It should be noted that the server 105 can be hardware or software. When the server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software or software modules (for example, used to provide the hierarchical text topic analysis method service), or as a single software or software module. No specific limitation is made here.
[0034] It should be understood that Figure 1 the numbers of the terminal devices, the network, and the server in
[0035] Continue to refer to Figure 2 , which shows a flow 200 of some embodiments of the hierarchical text topic analysis method according to the present disclosure. The hierarchical text topic analysis method includes the following steps:
[0036] Step 201, obtain the target text to be processed.
[0037] In some embodiments, the execution subject of the hierarchical text topic analysis method (such as Figure 1 the terminal device shown) can obtain the target text to be processed in response to receiving a target authorization signal. Among them, the target text includes a title text set and a paragraph text set. Specifically, the target text to be processed may contain a discourse structure. The title text set of the target text may include "I. Background Knowledge", "II. Method Introduction", "III. Experimental Analysis", "IV. Article Summary". The paragraph text set of the target text may include text paragraphs describing background knowledge, introducing methods, introducing experimental results, and making summaries.
[0038] Specifically, the above-mentioned target authorization signal can be a signal generated by the user of the above-mentioned target text performing a target operation on a target control. The above-mentioned target control may be included in an authorization prompt box. The above-mentioned authorization prompt box can be displayed on the target terminal device. The above-mentioned target terminal device may be a terminal device logged in with the user's corresponding account. The above-mentioned terminal device can be a "mobile phone" or a "computer". The above-mentioned target operation can be a "click operation" or a "swipe operation". The above-mentioned target control can be an "OK button".
[0039] As an example, the above-mentioned authorization prompt box can be as Figure 3As shown in the figure. The above-mentioned authorization prompt box may include: a prompt message display part 301 and a control 302. Among them, the above-mentioned prompt message display part 301 may be used to display a prompt message. The above-mentioned prompt message may be "Whether to allow obtaining the target text". The above-mentioned control 302 may be an "OK button" or a "Cancel button".
[0040] Step 202: Generate a chapter text set based on the target text.
[0041] In some embodiments, the above-mentioned execution subject generates a chapter text set based on the target text. Optionally, for each paragraph text in the paragraph text set of the target text, a paragraph text word vector set of the paragraph text is generated. Specifically, a word vector generation method can be used to convert the paragraph text into a paragraph text word vector set. For each paragraph text in the paragraph text set of the target text, a complete title text word vector set corresponding to the paragraph text is generated to obtain a complete title text word vector set. Optionally, for each paragraph text in the paragraph text set of the target text, a title text set of the paragraph text is determined. Specifically, for the paragraph text of the technical background content in the target text, the corresponding title text set includes "I. Background Knowledge" and "1.1 Technical Background Knowledge". Specifically, a word vector generation method can be used to convert the title text set of the paragraph text into a title text word vector set. Concatenate all the title text vectors in the title text vector set corresponding to the paragraph text to obtain a complete title text vector set.
[0042] Optionally, for each paragraph text in the paragraph text set of the target text, a chapter text is generated based on the complete title text vector set corresponding to the paragraph text and the paragraph text vector set to obtain a chapter text set. Specifically, concatenate the complete title text vector set corresponding to the paragraph text and the paragraph text vector set to generate a chapter text. The elements in the chapter text are word vectors.
[0043] Step 203: Determine the first-level theme sequence and the second-level theme sequence.
[0044] In some embodiments, the above-mentioned execution subject determines the first-level theme sequence and the second-level theme sequence. Among them, the first-level theme sequence includes the second number of first-level themes. The first-level theme corresponds to the third number of second-level themes. Specifically, the first-level theme sequence and the second-level theme sequence can be used to represent different coarse-grained themes. Specifically, the first-level theme may be "Example Company Name", and the second-level theme sequence may include "Company Business Model", "Company Financial Data", "Company Risks", "Company Prospect". "Example Company Name" as the first-level theme can correspond to 4 second-level themes.
[0045] Step 204: For each first-level theme in the first-level theme sequence, generate a mixed theme based on this first-level theme and the second-level theme sequence to obtain a mixed theme sequence.
[0046] In some embodiments, the above-mentioned execution entity generates a mixed theme based on each first-level theme in the first-level theme sequence and the second-level theme sequence to obtain a mixed theme sequence. Optionally, for each second-level theme in the second-level theme sequence, splice this first-level theme and this second-level theme to obtain a mixed theme.
[0047] Specifically, the first-level theme sequence may include "Example Company A Name" and "Example Company B Name", and the second-level theme sequence may include "Company Business Model", "Company Financial Data", "Company Risks", and "Company Prospect". For the first-level theme sequence "Example Company A Name", after splicing with the second-level theme sequence, mixed themes "Example Company A Business Model", "Example Company A Financial Data", "Example Company A Risks", and "Example Company A Prospect" can be obtained. For the first-level theme sequence "Example Company B Name", after splicing with the second-level theme sequence, mixed themes "Example Company B Business Model", "Example Company B Financial Data", "Example Company B Risks", and "Example Company B Prospect" can be obtained. Finally, the obtained mixed theme sequence includes: "Example Company A Business Model", "Example Company A Financial Data", "Example Company A Risks", "Example Company A Prospect", "Example Company B Business Model", "Example Company B Financial Data", "Example Company B Risks", and "Example Company B Prospect".
[0048] Step 205: Generate a set of theme-weight pairs based on the chapter text set and the mixed theme sequence.
[0049] In some embodiments, the above-mentioned execution entity generates a set of theme-weight pairs based on the chapter text set and the mixed theme sequence. Optionally, the above-mentioned execution entity generates a target text sequence based on the chapter text set and the mixed theme sequence. Specifically, use the word vector generation method to convert the mixed theme sequence into a mixed theme word vector sequence. Splice the chapter text set and the mixed theme word vector sequence to generate a target text sequence. The target text can be a word vector.
[0050] Optionally, input the target text sequence into a pre-trained theme analysis model to generate a set of theme-weight pairs. Among them, the theme-weight pair includes a theme text and a weight. The pre-trained theme analysis model includes a first distribution module, a second distribution module, and a generation module.
[0051] Optionally, generate a filtered text sequence based on the target text sequence. For each target text in the target text sequence, use the following formula to generate the weight value of this target text to obtain a weight value sequence:
[0052]
[0053] Among them, i is the target text count, word represents the word frequency, and word i represents the word frequency of the target text, weight represents the weight, and weight i represents the weight of the target text.
[0054] Use the following formula to calculate the overall word vector:
[0055]
[0056] Among them, i is the target text count, word represents the word frequency, and word i represents the word frequency of the i-th target text. weight represents the weight, and weight i represents the weight of the i-th target text. M represents the total number of target texts in the target text sequence, and B represents the overall word vector.
[0057] Based on the overall word vector, use the following formula to determine the divergence sequence:
[0058]
[0059] Among them, w represents the target text in the target text sequence, B represents the overall word vector, P(w|B) represents the frequency of w appearing in B, and T represents the target text sequence. P(w|T) represents the frequency of w appearing in T, and D is the divergence sequence.
[0060] Determine the target texts corresponding to the first three divergences in the divergence sequence as the screened texts to obtain the screened text sequence.
[0061] Optionally, input the screened text sequence into the first distribution module to determine the set of topic distribution probabilities. Specifically, for each screened text in the screened text sequence, input it into the first distribution module and use the following formula to calculate the topic distribution probability to obtain the set of topic distribution probabilities:
[0062]
[0063] Among them, b represents the mixed topic in the mixed topic sequence, and n(b) represents the number of times b appears in the target text sequence. represents the total number of times all mixed topics appear in the target text sequence. d represents the target text sequence, and P(b|d) represents the topic distribution probability corresponding to the screened text.
[0064] Input the screened text sequence into the second distribution module to determine the set of word distribution probabilities. Specifically, for each screened text in the screened text sequence, input it into the second distribution module, and use the following formula to calculate the word distribution probability to obtain the set of word distribution probabilities:
[0065]
[0066] Among them, b represents the mixed topic in the mixed topic sequence. c represents the chapter text. P(z) represents the number of occurrences of the chapter text c in the combination of chapter texts. represents the total number of occurrences of all chapter texts in the target text sequence. P(z|b) represents the word distribution probability.
[0067] Input the set of topic distribution probabilities and the set of word distribution probabilities into the generation module to determine the set of topic weight pairs. Specifically, for the topic distribution probability and the word distribution probability, use the following formula to calculate the weight in the topic weight pair to obtain the set of weights:
[0068]
[0069] Among them, d represents the target text sequence, and b represents the mixed topic in the mixed topic sequence. P(z|b) represents the word distribution probability. P(b|d) represents the topic distribution probability. P(z|d) represents the weight. Specifically, for each weight in the set of weights, determine the weight and the set of its corresponding topic texts b as the topic weight pair to obtain the set of topic weight pairs.
[0070] Optionally, push the set of topic weight pairs to the target device with a display function, and control the target device to display the set of topic weight pairs. Among them, the target device with a display function can be a device communicatively connected to the above-mentioned execution entity, and can display according to the received set of topic weight pairs. For example, the above-mentioned execution entity can display the set of topic weight pairs to display the main topic information involved in the target text. Applied in products such as social analysis, public opinion monitoring, and information distribution, it can improve the accuracy of subsequent text analysis and mining tasks.
[0071] The optional content in step 205 above, that is, "the technical content of generating a filtered text sequence" is an inventive point of an embodiment of the present disclosure, which solves the second technical problem mentioned in the background art: "When the information content contained in the text is large and the theme overlap is low, topics with low relevance to the theme are usually generated based on a large amount of text irrelevant to the theme. Under the influence of such interference factors, the expression accuracy of the generated topic model is poor." The factors that lead to poor expression accuracy of the topic model are often as follows: A large amount of text irrelevant to the theme in the text will interfere with the accuracy of the finally generated theme expression. If the above factors are solved, the effect of improving the accuracy of theme expression can be achieved. To achieve this effect, the present disclosure introduces a divergence sequence to calculate the divergence of each target text in the target text sequence, and determines the target texts corresponding to the first third number of divergences as the filtered texts to obtain a filtered text sequence. First, generate the weights of the target texts to obtain a weight sequence. Secondly, determine the overall word vector of the target text sequence. Then, determine the divergence sequence, and determine the target texts corresponding to the first third number of divergences as the filtered texts to obtain a filtered text sequence. Divergence is an index for measuring the difference between different distributions. Introducing divergence for filtering can eliminate interference information irrelevant to the theme and improve the accuracy of theme expression, thus solving the second technical problem.
[0072] Figure 2 An embodiment given has the following beneficial effects: First, some embodiments of the present disclosure generate a set of paragraph text word vectors based on the target text. Secondly, for each paragraph text in the paragraph text set of the target text, generate a set of complete title text word vectors corresponding to the paragraph text to obtain a set of complete title text word vectors. What is included in the set of complete title text word vectors is multi-layer title information. Thirdly, determine the chapter text set according to the set of paragraph text word vectors and the set of complete title text word vectors. Then, determine the mixed topic sequence for topic analysis. The mixed topic sequence takes into account both coarse-grained topics and fine-grained topics, and can make full use of the discourse structure to mine hierarchical information. Finally, based on the chapter text set and the mixed topic sequence, generate a set of topic weight pairs. The topic weight pair includes a topic text and a weight. Based on the chapter text set and the mixed topic sequence, generate a target text sequence. After performing filtering processing, eliminate information with low relevance and generate a filtered text sequence. Introducing the mixed topic sequence can mine the text topic using the hierarchical information of the discourse structure in the target text, improve the accuracy of the generated set of topic weight pairs, and the topic weight pair can better express the target text.
[0073] Next, refer to Figure 4 , which shows a schematic structural diagram of a computer system 400 of a terminal device suitable for implementing the embodiments of the present disclosure. Figure 4The terminal device shown is only an example and should not impose any limitations on the functions and scope of use of the embodiments of the present disclosure.
[0074] As Figure 4 shown, the computer system 400 includes a central processing unit (CPU) 401 that can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage section 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the system 400 are also stored. The CPU 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0075] The following components are connected to the I / O interface 405: a storage section 406 including a hard disk, etc.; and a communication section 407 including a network interface card such as a LAN (local area network) card, a modem, etc. The communication section 407 performs communication processing via a network such as the Internet. A drive 408 is also connected to the I / O interface 405 as needed. A removable medium 409, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 408 as needed so that a computer program read from it can be installed into the storage section 406 as needed.
[0076] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 407, and / or installed from the removable medium 409. When the computer program is executed by the central processing unit (CPU) 401, the above-described functions defined in the method of the present disclosure are performed. It should be noted that the computer-readable medium described in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0077] Computer program code for performing the operations of this disclosure may be written in one or more programming languages or combinations thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and also including conventional procedural programming languages such as the C language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0078] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0079] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, technical solutions formed by mutually replacing the above features with (but not limited to) technical features having similar functions disclosed in the present disclosure.
Claims
1. A hierarchical text topic analysis method, including: Obtaining a target text to be processed; Generating a chapter text set based on the target text; Determining a first-level topic sequence and a second-level topic sequence, where the first-level topic sequence includes a second number of first-level topics, and each first-level topic corresponds to a third number of second-level topics; For each first-level topic in the first-level topic sequence, generating a mixed topic based on the first-level topic and the second-level topic sequence to obtain a mixed topic sequence; Generating a target text sequence based on the chapter text set and the mixed topic sequence; Inputting the target text sequence into a pre-trained topic analysis model to generate the set of topic weight pairs, where a topic weight pair includes a topic text and a weight, and the pre-trained topic analysis model includes a first distribution module, a second distribution module, and a generation module. Inputting the target text sequence into the pre-trained topic analysis model to generate the set of topic weight pairs includes: Generating a filtered text sequence based on the target text sequence; Inputting the filtered text sequence into the first distribution module to determine a set of topic distribution probabilities; Inputting the filtered text sequence into the second distribution module to determine a set of word distribution probabilities; Inputting the set of topic distribution probabilities and the set of word distribution probabilities into the generation module to determine the set of topic weight pairs.
2. The method according to claim 1, wherein, the target text includes a title text set and a paragraph text set.
3. The method according to claim 2, wherein, generating the chapter text set based on the target text includes: For each paragraph text in the paragraph text set of the target text, generating a set of paragraph text word vectors for the paragraph text; For each paragraph text in the paragraph text set of the target text, generating a set of complete title text word vectors corresponding to the paragraph text to obtain the set of complete title text word vectors; For each paragraph text in the paragraph text set of the target text, generating a chapter text based on the set of complete title text word vectors corresponding to the paragraph text and the set of paragraph text word vectors to obtain the chapter text set.
4. The method according to claim 3, wherein, generating the set of complete title text word vectors corresponding to the paragraph text further includes: Concatenating the set of all title text word vectors corresponding to the paragraph text to obtain the set of complete title text word vectors.
5. The method according to claim 4, wherein, generating the mixed topic based on the first-level topic and the second-level topic sequence further includes: For each second-level topic in the second-level topic sequence, concatenating the first-level topic and the second-level topic to obtain the mixed topic.
6. The method according to claim 5, wherein, generating the filtered text sequence based on the target text sequence includes: For each target text in the target text sequence, using the following formula to generate a weight for the target text to obtain a weight sequence: Among them, is the target text count, represents the word frequency, represents the word frequency of the target text, represents the weight, represents the weight of the target text; Using the following formula to calculate the overall word vector: Among them, is the target text count, represents the word frequency, represents the word frequency of the th target text, represents the weight, represents the weight of the th target text, represents the overall word vector; Based on the overall word vector, using the following formula to determine the divergence sequence: Among them, represents the target text in the target text sequence, represents the overall word vector, represents in the frequency of occurrence, represents the target text sequence, represents in the frequency of occurrence, is the divergence sequence; Determine the target text corresponding to the first third number of divergences in the divergence sequence as the screening text to obtain the screening text sequence.
7. The method according to any one of claims 1-6, wherein, the method further includes: Pushing the set of topic weight pairs to a target device with a display function, and controlling the target device to display the set of topic weight pairs.
8. A first terminal device, comprising: One or more processors; A storage device storing one or more programs thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-7.
Citation Information
Patent Citations
Text event abstract generation method and device, electronic equipment and storage medium
CN111324728A
Data processing method and device
CN112100396A