A method, apparatus, and electronic device for generating a technical topic tree
By using a large language model to extract and merge themes based on semantic similarity, the method addresses the challenge of inaccurate and redundant theme identification in large text volumes, producing a clearer and more logical theme tree structure.
Patent Information
- Application Number
- CN202510197271.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-02-21
AI Technical Summary
In the prior art, the technical topic extraction of text is inaccurate, the hierarchical relationships between various technical topics are fuzzy, and the redundant structure is large, which makes it take longer for users to understand the text.
The topic information of the text is extracted through a large language model, an initial weighted directed graph is constructed, nodes are merged based on semantic similarity, target weighted directed graph is generated, technical topic trees are constructed, redundant nodes are reduced, and hierarchical relationships are clear.
It improves the efficiency and accuracy of obtaining text theme information, and the built technical theme tree is clear at the level and has strong semantic logic, reducing redundant nodes, and improving the convenience of users to understand text.
Smart Images

Figure CN120068849B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of text processing. Specifically, this application relates to a method, apparatus, and electronic device for generating a technical topic tree. Background Art
[0002] When the number of texts is large and the length of each text is long, users need to spend a lot of time understanding the texts. Currently, by extracting the technical topics of the texts, such as information about the technical fields and technical means of the texts, it is convenient for users to understand the core content of the texts and saves the time for users to understand the texts.
[0003] There may also be a hierarchical relationship between various technical topics, but the technical topics of the texts extracted by the related art are still inaccurate, and the hierarchical relationship between various technical topics also needs to be optimized. Summary of the Invention
[0004] Embodiments of this application provide a method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating a technical topic tree, which can solve the problem that the technical topics of the texts extracted by the related art are still inaccurate, and the hierarchical relationship between various technical topics also needs to be optimized. The technical solutions provided by this application are as follows:
[0005] According to one aspect of the embodiments of this application, a method for generating a technical topic tree is provided. The method includes:
[0006] Obtain a target text sequence, where the target text sequence includes a plurality of target texts;
[0007] Extract the topic information of each of the target texts through a large language model. Wherein, for each of the target texts, the topic information of the target text includes at least one technical topic of the target text. When there are multiple technical topics of the target text, the topic information of the target text further includes the hierarchical relationship between the multiple technical topics;
[0008] Determine an initial weighted directed graph corresponding to the target text sequence according to the topic information of each of the target texts. Wherein, each node in the initial weighted directed graph represents a technical topic, the edge between two nodes represents the hierarchical relationship between the technical topics corresponding to the two nodes, and the weight of an edge represents the occurrence frequency of the hierarchical relationship between the two nodes connected by the edge in the hierarchical relationship extracted from the multiple target texts;
[0009] Determine the semantic similarity between the technical topics represented by each node in the initial weighted directed graph, determine the semantic similarity that meets the preset conditions as the target similarity, and determine each node corresponding to each target similarity as a node to be merged;
[0010] Merge the nodes to be merged, the edges associated with the nodes to be merged, and the weights of the edges in the initial weighted directed graph to obtain a target weighted directed graph;
[0011] Construct a technical topic tree for the target text sequence according to the target weighted directed graph;
[0012] Among them, the preset conditions include at least one of the following:
[0013] The semantic similarity is not less than a preset similarity threshold;
[0014] In the order of sorting the semantic similarities from large to small, the top preset number of semantic similarities are sorted forward.
[0015] According to another aspect of the embodiments of the present application, there is provided a device for generating a technical topic tree, and the device includes:
[0016] A target text sequence acquisition module, configured to acquire a target text sequence, where the target text sequence includes multiple target texts;
[0017] A theme information extraction module, configured to respectively extract the theme information of each of the target texts through a large language model. Among them, for each target text, the theme information of the target text includes at least one technical theme of the target text. When there are multiple technical themes of the target text, the theme information of the target text further includes the hierarchical relationship between the multiple technical themes;
[0018] An initial weighted directed graph determination module, configured to determine an initial weighted directed graph corresponding to the target text sequence according to the theme information of each of the target texts. Among them, each node in the initial weighted directed graph represents a technical theme, the edge between two nodes represents the hierarchical relationship between the technical themes corresponding to the two nodes, and the weight of an edge represents the occurrence frequency of the hierarchical relationship between the two nodes connected by the edge in the hierarchical relationships extracted from the multiple target texts;
[0019] A to-be-merged information determination module, configured to determine the semantic similarity between the technical themes represented by each node in the initial weighted directed graph, determine the semantic similarity between the technical themes that meet the preset conditions as the target similarity, and determine each node corresponding to each target similarity as a node to be merged;
[0020] A merging module, configured to merge the nodes to be merged, the edges associated with the nodes to be merged, and the weights of the edges in the initial weighted directed graph to obtain a target weighted directed graph;
[0021] A theme tree construction module, configured to construct a technical theme tree for the target text sequence according to the target weighted directed graph;
[0022] Among them, the preset conditions include at least one of the following:
[0023] The semantic similarity is not less than a preset similarity threshold;
[0024] In the order of sorting the semantic similarities from large to small, the semantic similarities with the top ranking are the preset number.
[0025] Optionally, the topic information extraction module can be used to obtain a preset first prompt information, which is used to prompt the large language model to output the topic information of the target text;
[0026] For the first target text in the target text sequence, based on the first prompt information and this target text, through the large language model, extract the topic information of this target text;
[0027] For each target text other than the first target text in the target text sequence, based on the extracted technical topic and the first prompt information, determine a second prompt information, which is used to prompt the large language model to refer to the extracted technical topic and output the topic information of this target text. Based on this target text and the second prompt information, through the large language model, extract the topic information of this target text. The extracted technical topic is the technical topic of each target text before this target text in the target text sequence.
[0028] Optionally, the initial weighted directed graph determination module can be used to construct a first weighted directed graph according to the topic information of the first target text;
[0029] For each target text other than the first target text, every time the topic information of a target text is extracted, then update the constructed weighted directed graph according to the topic information of this target text until the weighted directed graph updated based on the topic information of the last target text is obtained, and determine this weighted directed graph as the initial weighted directed graph corresponding to the target text sequence.
[0030] Optionally, the initial weighted directed graph determination module can be used to count the first quantity of the extracted technical topics and determine the number of times each technical topic appears to obtain the counting result of each technical topic;
[0031] If the first quantity is greater than a preset value, then delete the topic information corresponding to the technical topics ranked after the preset value in the order of the counting results of the extracted technical topics from large to small;
[0032] Based on the topic information of the remaining target texts, determine the initial weighted directed graph corresponding to the target text sequence.
[0033] Optionally, the technical topic tree construction module may be used to delete self-loop edges in the target weighted directed graph;
[0034] For each node in the target weighted directed graph after deleting self-loop edges, determine the importance of each incoming edge of the node according to the weights of the incoming edges of the node, so as to determine the parent-child relationship in the technical topic tree to be constructed according to the determined importance of each incoming edge;
[0035] Construct the technical topic tree of the target text sequence according to each node in the target weighted directed graph and the importance of each incoming edge of each node.
[0036] Optionally, the technical topic tree construction module may be used to initialize the stack structure;
[0037] By continuously performing the target operation until all nodes in the target weighted directed graph are included in the constructed technical topic tree, determine the technical topic tree constructed by each target operation as the technical topic tree of the target text sequence;
[0038] Among them, the target operation includes:
[0039] Determine the node with the smallest out-degree and / or the smallest sum of weights of outgoing edges in the current directed graph as the leaf node; where the current directed graph of the first target operation is the target weighted directed graph;
[0040] Use the leaf node as the first node to be processed in this target operation, and continuously perform the following operations on the node to be processed until the stop pushing condition is met:
[0041] Push the node to be processed onto the stack; determine the incoming edge with the highest importance of the node to be processed, and determine the other node connected by the incoming edge except the node to be processed as the parent node of the node to be processed;
[0042] If the stop pushing condition is not met, use the parent node of the node to be processed as the next node to be processed in this target operation;
[0043] If the stop pushing condition is met, pop the nodes that have been pushed onto the stack, construct the technical topic tree corresponding to this target operation based on the popped nodes, delete the leaf node of this target operation from the current directed graph, and subtract 1 from the out-degree of the parent node of the leaf node to obtain the current directed graph of the next target operation.
[0044] Optionally, the technical topic tree construction module may be used to determine the other node as the parent node of the node to be processed if there is no other node connected by the incoming edge in the stack except the node to be processed;
[0045] The subject tree construction module can be used to, if there is another node other than the to-be-processed node connected by the incoming edge in the stack, determine the other node other than the to-be-processed node connected by the target incoming edge as the parent node of the to-be-processed node, where the target incoming edge is the incoming edge with the highest importance among the incoming edges in which the other node connected by the incoming edge is not in the stack, and the other incoming edges are the incoming edges of the to-be-processed node other than the incoming edge with the highest importance.
[0046] Optionally, the stop pushing condition includes at least one of the following:
[0047] The in-degree of the to-be-processed node is 0;
[0048] The weight of the incoming edge of the to-be-processed node is less than the fusing threshold;
[0049] When the weight of the incoming edge of the to-be-processed node is less than the fusing threshold, pop out the nodes that have been pushed into the stack, and construct a technical subject tree corresponding to the target operation this time based on the popped-out nodes, including:
[0050] When the weight of the incoming edge of the to-be-processed node is less than the fusing threshold, the subject tree construction module can be used to take the node adjacent to the last pushed-in node in the stack as the root, and construct a technical subject tree based on the nodes pushed into the stack before the last pushed-in node;
[0051] Take the last pushed-in node as the root, construct a technical subject tree, and obtain the technical subject tree corresponding to the target operation this time.
[0052] Optionally, the to-be-merged information determination module can be used to input the technical subject represented by each node in the initial weighted directed graph into the semantic extraction model, and obtain the semantic features of the technical subject represented by this node output by the semantic extraction model;
[0053] According to the semantic features of the technical subject represented by this node, determine the semantic similarity between the technical subject represented by this node and the technical subjects represented by other nodes in the initial weighted directed graph.
[0054] Optionally, the merging module can be used to, for each target similarity, if the node pair corresponding to this target similarity and the node pairs corresponding to other target similarities do not have the same nodes, take the node pair corresponding to this target similarity as the to-be-merged set, and if the node pair corresponding to this target similarity and the node pairs corresponding to other target similarities have the same nodes, take the nodes in the node pairs with the same nodes as the to-be-merged set;
[0055] For each set to be merged, input the technical themes represented by the nodes in the set to be merged and the preset task prompt data into the theme optimization model to obtain the target nodes of the set to be merged. The task prompt data is used to prompt the theme optimization model to screen out the target nodes from the input nodes.
[0056] For non-target nodes in each set to be merged, update the edges and edge weights between other nodes and the non-target nodes except the nodes in the set to be merged to the edges and edge weights between other nodes and the target nodes, and delete the edges, edge weights and non-target nodes between the target nodes and non-target nodes in the set to be merged to obtain a target weighted directed graph.
[0057] According to another aspect of the embodiments of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the method provided in any optional embodiment of the present application.
[0058] According to still another aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method provided in any optional embodiment of the present application are implemented.
[0059] According to one aspect of the embodiments of the present application, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps of the method provided in any optional embodiment of the present application are implemented.
[0060] The beneficial effects brought by the technical solutions provided in the embodiments of the present application are as follows: The present application obtains a target text sequence, which includes a plurality of target texts. Based on the large language model, the theme information of the target text is automatically extracted, improving the efficiency of obtaining the theme information of the target text. Based on the theme information, an initial weighted directed graph is constructed. The initial weighted directed graph includes the hierarchical relationship and weights between technical themes, which is easy to determine the technical theme tree subsequently and improves the efficiency of constructing the technical theme tree. The semantic similarity of the technical themes represented by the nodes is determined, and the nodes can be merged according to the semantic similarity to obtain a target weighted directed graph, reducing redundant nodes, making the extracted technical themes more accurate, and making the hierarchy between the technical themes obtained from the target weighted directed graph clearer and the semantic logic stronger. Description of the Drawings
[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for the description of the embodiments of the present application will be briefly introduced below.
[0062] Figure 1 Schematic diagram of a system for generating a technical theme tree provided by an embodiment of the present application;
[0063] Figure 2 A flowchart showing the method for generating a technical theme tree provided by an embodiment of the present application;
[0064] Figure 3 A schematic diagram of an initial weighted directed graph provided by an embodiment of the present application;
[0065] Figure 4 The unmerged initial weighted directed graph provided by an embodiment of the present application;
[0066] Figure 5 A schematic diagram of a target weighted directed graph provided by an embodiment of the present application;
[0067] Figure 6 Another target weighted directed graph provided by an embodiment of the present application;
[0068] Figure 7 A flowchart showing the method for determining an initial weighted directed graph of a target text sequence provided by an embodiment of the present application;
[0069] Figure 8 A flowchart showing the method for constructing a technical theme tree provided by an embodiment of the present application;
[0070] Figure 9 A schematic diagram of the structure of the technical theme tree provided by an embodiment of the present application;
[0071] Figure 10 A schematic diagram of the structure of a generating device for a technical theme tree provided by an embodiment of the present application;
[0072] Figure 11 A schematic diagram of the structure of an electronic device corresponding to the method for generating a technical theme tree provided by an embodiment of the present application. Detailed implementation manners
[0073] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions of the embodiments of the present application.
[0074] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the terms "comprising" and "including" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components, and / or their combinations supported by the technical field of the present invention. It should be understood that when we say an element is "connected" or "coupled" to another element, the element can be directly connected or coupled to the other element, or it can mean that the element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein can include wireless connection or wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term. For example, "A and / or B" or "A, B" indicates that it is implemented as "A", or implemented as "B", or implemented as "A and B". When describing multiple (two or more) items, if the relationship between the multiple items is not clearly defined, the multiple items can refer to one, multiple, or all of the multiple items. For example, for the description of "parameter A includes A1, A2, A3", it can be implemented that parameter A includes A1 or A2 or A3, or it can also be implemented that parameter A includes at least two of the three items of parameter A1, A2, and A3.
[0075] Structural extraction of the text is easy for users to understand the text and facilitates subsequent execution of various services. There are still problems with the technical themes extracted by the related technology. These include inaccurate extraction of technical themes, ambiguous hierarchical relationships between technical themes, and a relatively large number of redundant structures.
[0076] The embodiments of the present application provide a method for generating a technical theme tree. The embodiments of the present application can automatically extract unoptimized technical themes through a large language model, and then generate a directed graph, which is easy to construct a technical theme tree subsequently. Through semantic similarity, node merging is performed to complete the optimization of technical themes and obtain more accurate technical themes. Redundant nodes are reduced to obtain a technical theme tree with clearer levels and stronger semantic logic.
[0077] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0078] The method in the embodiments of the present application can be implemented by a technical theme tree generation system. For example, Figure 1Schematic diagram of a system for generating a technical theme tree provided by an embodiment of the present application. The system includes a terminal device 101 and a server 102. The terminal device 101 can send a target text sequence to the server 102 through a network, so that the server 102 extracts the theme information of the target text sequence and constructs a technical theme tree. The server 102 is deployed with a large language model. The server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (Content Delivery Network, CDN), and big data and artificial intelligence platforms, but is not limited thereto. The terminal device 101 and the server 102 can be directly or indirectly connected through wired or wireless communication methods, which are not limited in the embodiments of the present application.
[0079] Next, through the description of several exemplary embodiments, the technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application are described. It should be noted that the following embodiments can be referenced, learned from, or combined with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.
[0080] Figure 2 Flow chart of a method for implementing the generation of a technical theme tree provided by an embodiment of the present application, as Figure 2 shown, the method includes S201~S206:
[0081] S201: Obtain a target text sequence, which includes multiple target texts.
[0082] The target texts in the target text sequence are arranged based on a preset order. The target text can be any type of text for which technical themes are to be extracted, including academic literature, patent documents, user manuals, etc. Each target text in the target text sequence can belong to the same technical field. For example, each target text in the target text sequence is in the field of electronic design.
[0083] The server can receive the target text sequence sent by the terminal, or obtain the target text sequence through other means, which is not limited in the embodiments of the present application.
[0084] S202: Respectively extract the theme information of each target text through a large language model. Among them, for each target text, the theme information of the target text includes at least one technical theme of the target text. When there are multiple technical themes of the target text, the theme information of the target text also includes the hierarchical relationship between the multiple technical themes.
[0085] Among them, the technical theme can characterize the core content recorded in the target text, including the technical field, technical means, etc. of the target text. For example, the technical theme can include "electronic engineering / chip design / low-power design", indicating that the target text corresponding to this technical theme belongs to the field of electronic engineering technology and specifically involves the realization of low-power design for chips.
[0086] The hierarchical relationship refers to the inclusion relationship between technical themes. The embodiments of the present application do not limit the format of the hierarchical relationship in the theme information. It can be directly recorded in the text, such as directly recording that "theme a includes theme b", or characterized by characters, such as a / b, indicating that theme a includes theme b.
[0087] For example, the technical theme can include "electronic engineering / chip design / low-power design", indicating that electronic engineering includes chip design and chip design includes low-power design. It can be understood that the technical field of the target text is electronic engineering. When performing chip design, the technical means of low-power design used is the low-power design method in the chip design field.
[0088] It should be noted that the method of low-power design can be used not only in chip design but also in the design of other fields, such as software algorithm design. The differences in the upper-level fields of low-power design result in different specific methods for low-power design. For example, in the field of chip design, low-power design can consider changing the architecture of the processor, and in software algorithm design, low-power design can consider changing the data structure of the algorithm.
[0089] For each target text in the target text sequence, the server can input the target text into a large language model to obtain the theme information of the target text output by the large language model. The embodiments of the present application do not limit the specific type of the large language model.
[0090] Through the large language model, the theme information of the target text can be automatically extracted to improve the efficiency of constructing the technical theme tree of the target text sequence. In addition, the large language model has a large number of model parameters and a relatively complex network structure, making the large language model itself have a powerful natural language analysis ability. Therefore, the theme information of each target text output by the large language model is more accurate, and the hierarchy of the technical theme tree of the target text sequence constructed based on the theme information of each target text is also more accurate and clear.
[0091] It should be noted that the technical topic tree of the target text sequence can represent the core content of the target text sequence. The nodes in the technical topic tree represent technical topics, and the edges connecting two nodes represent the hierarchical relationship between the technical topics corresponding to the two nodes. For example, node a is connected to node b, node a is the root node, node b is the leaf node, and the hierarchical relationship is that the technical topic represented by node a includes the technical topic represented by node b.
[0092] S203: Determine the initial weighted directed graph corresponding to the target text sequence according to the topic information of each target text, where each node in the initial weighted directed graph represents a technical topic, the edge between two nodes represents the hierarchical relationship between the technical topics corresponding to the two nodes, and the weight of an edge represents the occurrence frequency of the hierarchical relationship corresponding to the two nodes connected by the edge in the hierarchical relationships extracted from the multiple target texts.
[0093] In the initial weighted directed graph, a technical topic can be represented by a node, and the hierarchical relationship between the two technical topics corresponding to the two nodes connected by the directed edge can be represented by the directed edge. It can be understood that in the initial weighted directed graph, the hierarchical relationship can be represented by the parent-child relationship between nodes. For example, if the technical topic corresponding to node h includes the technical topic corresponding to node i, then node h is the parent node of node i, and in the initial weighted directed graph, node h points to node i.
[0094] For each hierarchical relationship, the hierarchical relationship may be extracted at least once from the multiple target texts, and the number of times the hierarchical relationship is extracted can also represent the importance of the hierarchical relationship in the target text sequence. Therefore, the occurrence frequency of the hierarchical relationship corresponding to the two nodes connected by the edge in the hierarchical relationships extracted from the multiple target texts can be represented by the weight of the edge.
[0095] Figure 3 The following is a schematic diagram of an initial weighted directed graph provided by an embodiment of the present application, as Figure 3 shown. For example, the technical topics included in the topic information of each target text are a, b, c, d, e, where the hierarchical relationships are that a includes b and d, b includes c, c includes d, and d includes e. The occurrence frequencies of the hierarchical relationships between the two nodes in the hierarchical relationships extracted from each target text are 10, 6, 8, 7, and 4 respectively. According to the hierarchical relationships in the topic information, the parent-child relationships between the nodes can be determined. Specifically, a points to b, the weight of the edge is 10, a points to d, the weight of the edge is 6, b points to c, the weight of the edge is 8, c points to d, the weight of the edge is 7, and d points to e, the weight of the edge is 4.
[0096] Construct an initial technical theme structure of the target text sequence in a graphical format according to the theme information of each target text. The technical theme in graphical format is clearer and also helps to merge the technical themes in the initial technical theme structure later, reducing redundant technical themes. Constructing a technical theme tree based on a weighted directed graph also significantly improves the efficiency of constructing the technical theme tree.
[0097] S204: Determine the semantic similarity between the technical themes represented by each node in the initial weighted directed graph, determine the semantic similarity between the technical themes that meet the preset conditions as the target similarity, and determine each node corresponding to each target similarity as the node to be merged.
[0098] Since the initial weighted directed graph is directly obtained according to the theme information of each target text, and there may be highly similar content between the target texts, then the technical themes represented by each node in the obtained initial weighted directed graph may also be highly similar. The representation of too many relatively similar technical themes in the initial weighted directed graph may lead to a blurred structure of the initial weighted directed graph, and further cause the technical theme tree of the final obtained target text sequence to be blurred and the hierarchical relationship to be chaotic. Therefore, the server can merge the technical themes represented by each node in the initial weighted directed graph.
[0099] Specifically, the server can determine the semantic similarity between the technical themes represented by each node in the initial weighted directed graph, and then screen out the nodes to be merged. The server can extract the semantic features of each technical theme, and based on the semantic features of each technical theme, determine the semantic similarity between each technical theme through cosine similarity. It is also possible to determine the semantic similarity between the technical themes represented by each node in the initial weighted directed graph by other methods, such as performing singular value decomposition (SVD, Singular Value Decomposition) on each technical theme through latent semantic analysis (LSA) to determine the potential semantic correlation between each technical theme, and then obtaining the semantic similarity between each technical theme. The embodiments of the present application do not limit this.
[0100] After that, the server can determine that the semantic similarities between the technical themes represented by each node that meet the preset conditions are the target similarities, and determine each node corresponding to each target similarity as the node to be merged. Among them, the preset conditions include at least one of the following: the semantic similarity is not less than the preset similarity threshold; in the order of sorting the similarities from largest to smallest, the top preset number of semantic similarities are selected. For example, if the preset similarity threshold is 70%, then the semantic similarities with the semantic similarities between the technical themes represented by each node not less than 70% are determined as the target similarities to determine the nodes to be merged. It is also possible to sort the semantic similarities between the technical themes represented by each node to obtain the similarity ranking of each semantic similarity, and determine the top three semantic similarities in the similarity ranking as the target similarities.
[0101] By using the semantic similarities between the technical themes to determine the nodes to be merged, multiple technical theme information with relatively similar semantics can be screened out. For multiple technical themes with relatively similar semantics, only one technical theme can be retained, and redundant technical themes can be removed to ensure that the structure of the technical theme tree of the target text sequence is clearer, more accurate and hierarchically reasonable.
[0102] S205: Merge the nodes to be merged, the edges associated with the nodes to be merged, and the weights of the edges in the initial weighted directed graph to obtain a target weighted directed graph.
[0103] The server can determine the nodes to be retained among the nodes to be merged, and merge the edges and the weights of the edges associated with the nodes to be deleted among the nodes to be merged into the nodes to be retained. Among them, there is no unique limitation on how to specifically determine the nodes to be retained among the nodes to be merged in the embodiments of the present application. For example, for each target similarity, any node in a node pair corresponding to the target similarity can be used as the node to be retained, and the other node in the node pair can be used as the node to be deleted. It is also possible to determine the node to be retained and the node to be deleted according to the number of connected edges of each node in the node pair. For example, retain the node with the largest number of connected edges. If there are two nodes with the largest number of connected edges among at least two nodes, the sum of the weights of all the connected edges of these two nodes can also be considered, and the node with the larger sum of weights is retained.
[0104] As another alternative, for each target similarity, if the node pair corresponding to the target similarity does not have the same node as the node pairs corresponding to other target similarities, the node pair corresponding to the target similarity is used as the set to be merged. If the node pair corresponding to the target similarity has the same node as the node pairs corresponding to other target similarities, the nodes in the node pairs with the same node are used as the set to be merged;
[0105] For each set to be merged, input the technical themes represented by the nodes in the set to be merged and the preset task prompt data into the theme optimization model to obtain the target node of the set to be merged. The task prompt data is used to prompt the theme optimization model to screen out the target node from the input nodes.
[0106] For non-target nodes in each set to be merged, update the edges and edge weights between other nodes and the non-target node except the nodes in the set to be merged to the edges and edge weights between other nodes and the target node, and delete the edges, edge weights, and non-target nodes between the target node and the non-target node in the set to be merged to obtain the target weighted directed graph.
[0107] Figure 4 The unmerged initial weighted directed graph provided by the embodiment of the present application is as Figure 4 shown. The nodes include a, b, c, d, and e. The weight of the edge connecting node a and node b is 8, the weight of the edge connecting node c and node b is 5, the weight of the edge connecting node c and node d is 3, and the weight of the edge connecting node e and node b is 2. Select a pair of nodes with the highest semantic similarity, that is, b and node c, for the merging operation. Then, the target similarity is only the semantic similarity corresponding to node b and node c, and there is no other target similarity. Then, there are no identical nodes between the node pair corresponding to the target similarity and the node pairs corresponding to other target similarities. Node b and node c form a set to be merged. Input the technical themes represented by nodes b and c in the set to be merged and the preset task prompt data into the theme optimization model to obtain the target node, that is, node b. The target node is the node to be retained among the nodes to be merged, and the other node of the node pair is the non-target node, that is, the node to be deleted, that is, node c.
[0108] The technical themes represented by nodes b and c respectively and the preset task prompt data can also be combined to obtain the input data. Suppose the technical theme represented by node b is machine learning, and the technical theme represented by node c is deep learning. Table 1 shows the input data for the input theme optimization model, as shown in Table 1:
[0109] Table 1
[0110]
[0111] In Table 1, the input data includes task prompt data, that is, the task description, and also includes the technical themes represented by the embedded nodes b and c respectively, that is, machine learning and deep learning.
[0112] The theme optimization model can not only output the target node but also give reasons. Table 2 shows the output result of the theme optimization model, as shown in Table 2:
[0113] Table 2
[0114]
[0115] According to the target theme output by the theme optimization model, the nodes in the set to be merged corresponding to the target theme are determined as the target nodes. The target theme retained in Table 2 is machine learning, and the node corresponding to machine learning is node b. Therefore, the target node is node b, and the non-target node is node c.
[0116] Figure 5 For an initial weighted directed graph provided by an embodiment of the present application based on Figure 4 A schematic diagram of the target weighted directed graph obtained by merging nodes is as Figure 5 shown.
[0117] During the merging, the edge connecting node d and node c can be changed to connect node d and node b, and the edge weight of the edge connecting node c and node d can be changed to the weight of the edge connecting node b and node d, resulting in the weight of the edge connecting node b and node d being 3. Node c and the edge connecting node c and node b and the edge weight are deleted.
[0118] It should be noted that if there are the same nodes among the node pairs corresponding to the target similarity and the node pairs corresponding to other target similarities, the nodes in each node pair with the same nodes are used as the set to be merged, and then the target nodes in the set to be merged are determined for subsequent merging. There are at least three nodes in this set to be merged. When determining the target nodes, the technical themes represented by all the nodes in the set to be merged and the preset task prompt data can be input into the theme optimization model, or the technical themes represented by the nodes in one node pair corresponding to each target similarity in the set to be merged can be input into the theme optimization model respectively.
[0119] For example, the node pair corresponding to target similarity 1 includes node b and node c, and the node pair corresponding to target similarity 2 includes node c and node d. Then, the server can obtain a set to be merged including node b, node c, and node d. The target similarities corresponding to the nodes in the set to be merged include target similarity 1 and target similarity 2. Therefore, the technical themes represented by the node pair corresponding to target similarity 1, that is, node b and node c, are input into the theme optimization model, and it is determined that the target node in the node pair corresponding to target similarity 1 is node c. The technical themes represented by the node pair corresponding to target similarity 2, that is, node d and node c, are input into the theme optimization model, and it is determined that the target node in the node pair corresponding to target similarity 2 is node d. Then, it is determined that the target node of the set to be merged is node d, and the non-target nodes are node b and node c. When performing node merging, node c and node b can be merged first to obtain an updated node c, and then the updated node d and node c can be merged.
[0120] Figure 6 Another target weighted directed graph provided by the embodiment of the present application is as follows Figure 6 shown.
[0121] The target weighted directed graph includes nodes a, d, and e. Among them, node a includes node d, and node d includes node e. The weight of the edge connecting node a and node d is 8, and the weight of the edge connecting node d and node e is 2.
[0122] The structure of the obtained target weighted directed graph after merging is more concise, and the complexity is reduced. When there are many similar technical themes in the target text sequence, a large amount of computing resources are also required to construct the technical theme tree of the target text sequence subsequently. By merging nodes, there is no need to add redundant nodes to the technical theme tree of the constructed target text sequence subsequently, reducing the computing resources for constructing the technical theme tree of the target text sequence.
[0123] S206: Construct the technical theme tree of the target text sequence according to the target weighted directed graph.
[0124] The server can construct the technical theme tree of the target text sequence according to the parent-child relationship between the nodes in the target weighted directed graph.
[0125] Specifically, the server can determine each leaf node in the target weighted directed graph through a depth-first search algorithm, and then determine each parent node of each leaf node according to the parent-child relationship between the nodes. Determine each parent node of each leaf node as a child node, and then determine each parent node of each child node until there is no parent node for each child node, obtaining the link relationship corresponding to each node in the target weighted directed graph, and constructing a corresponding technical theme tree according to the link relationship.
[0126] The tree structure form can more intuitively display the association relationship between the technical themes of each target text in the target text sequence. Therefore, presenting the technical theme tree of the target text sequence to the user can enable the user to more quickly understand each target text in the target text sequence. After the foregoing steps S202~S205, the levels of the obtained technical theme trees of the target text sequence are also clear and the logic is reasonable, further improving the convenience for the user to understand the target text sequence.
[0127] For step S201, the server can also perform text preprocessing on the obtained target text sequence. The text preprocessing includes integrity check and structured processing. The integrity check includes verifying the integrity of each target text in the target text sequence to ensure that the content is complete and error-free. The structured processing includes parsing unstructured document content, such as paragraph text and technical descriptions, into a structured format for subsequent processing link calls.
[0128] Regarding step S202, when extracting the topic information of each target text, the server inputs each target text into the large language model according to the order of the target texts in the target text sequence, and obtains the topic information of each target text.
[0129] The server can also obtain the first prompt information of the target text sequence. Based on the first prompt information, through the large language model, the topic information of each target text is extracted, and the first prompt information of each target text can be the same. Among them, the first prompt information can include any one of task description, step description, precautions, input and output examples, and can be specifically set according to needs.
[0130] Through the first prompt information, the accuracy of the topic information output by the large language model can be improved, the technical topics in each target text can be extracted more comprehensively, and the technical topic tree of the subsequent constructed target text sequence can be made more reasonable. The first prompt information can also include prompt words in a specific field, etc., so that the large language model can combine specific domain knowledge to hierarchically extract the technical topics of the target text to meet the user's needs for extracting technical topics in a specific field.
[0131] The first prompt information can be the content in Table 3 below:
[0132] Table 3
[0133]
[0134] In the example of the prompt information shown in Table 3, the prompt information includes the tasks of the large language model and the task processing methods (task description, step description, and precautions), and also includes the learning examples for the large language model to refer to and learn (the examples in Table 3). In the example of the prompt information in Table 3, taking the prompt word in a specific field as integrated circuit, the large language model will extract the technical topics in the field of integrated circuits, and can output up to 5 hierarchical relationships between each technical topic, and use a semicolon to separate each hierarchical relationship.
[0135] Specifically, the text received by the large language model includes target text sequence 1 or target text sequence 2. The first prompt information and each target text in target text sequence 1 or target text sequence 2 can be input into the large language model to obtain the topic information of each target text. Of course, each target text in target text sequence 1 or target text sequence 2 can also be embedded into the first prompt information, and then the first prompt information embedded with each target text is input into the large language model. The embodiments of the present application do not limit this.
[0136] The server can also obtain the preset first prompt information, which is used to prompt the large language model to output the topic information of the target text;
[0137] For the first target text in the target text sequence, based on the first hint information and the target text, use a large language model to extract the theme information of the target text;
[0138] For each target text other than the first target text in the target text sequence, based on the extracted technical theme and the first hint information, determine the second hint information. The second hint information is used to prompt the large language model to refer to the extracted technical theme and output the theme information of the target text. Based on the target text and the second hint information, use a large language model to extract the theme information of the target text. The extracted technical theme is the technical theme of each target text before the target text in the target text sequence.
[0139] It can be understood that for each target text other than the first target text in the target text sequence, the server can use the technical themes of each target text before the target text in the target text sequence as reference words, so that the large language model can refer to the extracted technical theme and output the technical theme of the target text.
[0140] The extracted technical theme included in the second hint information enables the large language model to directly use the extracted technical theme without regenerating the technical theme, avoiding repeated generation. When there are more target texts, there may be more repeated technical themes. The existence of the second hint information can improve the extraction efficiency of theme information. In addition, if the target texts in the target text sequence belong to the same field, the extracted technical theme can be used as the hint information of the large language model, which can further improve the accuracy of the technical theme output by the large language model.
[0141] The extracted technical theme in the second hint information can be as shown in Table 4 below:
[0142] Table 4
[0143]
[0144] The technical theme in the reference words in Table 4 can be at least one extracted technical theme. Each target text can be embedded in the second hint information, that is, the position of the target text shown in Table 4, combined with the reference words, to obtain the input data, and the input data is input into the large language model.
[0145] For step S203, when constructing the initial weighted directed graph corresponding to the target text sequence, after the server obtains the theme information of all target texts in the target text sequence, it can then determine the initial weighted directed graph corresponding to the target text sequence. It can also construct the first weighted directed graph according to the theme information of the first target text;
[0146] For each target text except the first one, every time the theme information of a target text is extracted, the constructed weighted directed graph is updated according to the theme information of this target text until a weighted directed graph updated based on the theme information of the last target text is obtained, and this weighted directed graph is determined as the initial weighted directed graph corresponding to this target text sequence.
[0147] Figure 7 The figure is a schematic flowchart of a method for determining the initial weighted directed graph of a target text sequence provided by an embodiment of the present application, as Figure 7 shown.
[0148] i is the order of the target text currently input to the large language model in the target text sequence. It can be understood that after the theme information of the first target text in the target text sequence is extracted, a first weighted directed graph in the target text sequence is constructed according to the theme information of the first target text in the target text sequence. Then, the theme information of the second target text in the target text sequence is obtained, and the first weighted directed graph is updated according to the theme information of the second target text until a weighted directed graph updated based on the theme information of the last target text is obtained as the initial weighted directed graph of this target text sequence.
[0149] It can be understood that after the theme information of the target text is generated each time, the hierarchical relationship between the technical themes of the currently generated theme information can be dynamically added to the initial weighted directed graph to ensure the continuous accumulation and optimization of the results.
[0150] When there are many target texts in the target text sequence, the number of technical themes in the corresponding theme information will also be large. Then, the number of nodes in the initial weighted directed graph corresponding to the target text sequence constructed will also be large, resulting in a relatively complex structure of the initial weighted directed graph corresponding to the target text sequence. Therefore, by retaining a preset number of technical themes, an initial weighted directed graph corresponding to the target text sequence with a relatively simple structure can be obtained.
[0151] Specifically, count the first number of the extracted technical themes; and determine the number of times each technical theme appears to obtain the count result of each technical theme;
[0152] If the first number is greater than the preset value, then according to the order of the count results of the extracted technical themes from large to small, delete the theme information corresponding to the technical themes ranked after the preset value;
[0153] Based on the retained theme information of each target text, determine the initial weighted directed graph corresponding to this target text sequence.
[0154] For example, if the preset value is 50, then 50 technical topics are retained. If the first quantity is 101, then they are sorted in descending order according to the counting results of each technical topic, and the top 50 technical topics corresponding to the counting results are retained, and the remaining 51 technical topics and their hierarchical relationships are deleted, that is, the topic information corresponding to the technical topics sorted after 50 is deleted.
[0155] The number of times a technical topic appears in the topic information of each target text can characterize the importance of the technical topic to the target text sequence. If the number of times is more, it means the technical topic is more important and can better represent the target text sequence. Then, a preset number of technical topics with more occurrences can be retained to construct an initial weighted directed graph corresponding to the target text sequence.
[0156] It should be noted that after the server obtains the topic information of all target texts in the target text sequence, it can then count the first quantity of technical topics in the topic information of all target texts. It can also count the first quantity of technical topics in the topic information of the currently extracted target text each time it obtains the topic information of a target text.
[0157] For the latter case, when performing step S203, the counting result of the technical topic in the currently obtained topic information in the initial weighted directed graph can be updated. For example, the currently obtained topic information includes topic 1, which is represented by node 1 in the initial weighted directed graph, and the counting result of node 1 is 3. Then, the counting result of this node 1 is incremented by 1 to become 4, indicating that after obtaining the topic information of the current target text, the technical topic represented by this node appears 4 times in the topic information of each target text that has been extracted.
[0158] In addition, for the latter case, the server can optimize the efficiency of the technical topic generation task by dynamically storing high-frequency technical topics. By recording the nodes with higher frequencies during the generation process and preferentially retaining the content that contributes more to the task result, redundant calculations and interference from low-frequency topic information can be reduced. If the topic information in the second prompt message is all the technical topics that have been extracted, then, through the foregoing method, the data volume of the second prompt message can be reduced, and the processing efficiency of the large language model can be improved. In addition, if the server updates the retained technical nodes each time it obtains the topic information of a target text, then by dynamically storing high-frequency technical topics, it can ensure that the retained technical topics always match the business requirements, thereby significantly improving the overall performance and response speed of large-scale text processing tasks.
[0159] For step S204, the server can input the technical topic represented by each node in the initial weighted directed graph into the semantic extraction model to obtain the semantic features of the technical topic represented by this node output by the semantic extraction model;
[0160] Determine the semantic similarity between the technical theme represented by this node and the technical themes represented by other nodes in the initial weighted directed graph according to the semantic features of the technical theme represented by this node.
[0161] Follow Figure 4 Example, suppose Figure 4 The technical themes represented by nodes a, b, c, d, and e in are artificial intelligence, machine learning, deep learning, neural network, and computer vision respectively. Input "artificial intelligence, machine learning, deep learning, neural network, computer vision" into the semantic extraction model respectively, and obtain the semantic features of the technical themes represented by each node output by the semantic extraction model. As follows: The semantic feature (semantic vector) of node a is [0.9, 0.7, 0.2], the semantic feature of node b is [0.8, 0.6, 0.3], the semantic feature of node c is [0.7, 0.6, 0.4], the semantic feature of node d is [0.6, 0.5, 0.5], and the semantic feature of node e is [0.4, 0.3, 0.8].
[0162] Calculate the cosine similarity between each pair of nodes, generate a similarity matrix and sort it, and take the top three pairs as examples: The semantic similarity between node b and node c is 0.95, the semantic similarity between node c and node d is 0.90, and the semantic similarity between node a and node b is 0.85.
[0163] Of course, the theme information of the technical themes represented by each node can also be input into the semantic extraction model. The theme information of the technical theme includes the technical theme name in text form and the hierarchical relationship between this technical theme and other technical themes. Following the above example, the technical theme names of each node are artificial intelligence, machine learning, deep learning, neural network, and computer vision respectively.
[0164] Extracting the semantic features of the technical themes represented by each node through the semantic extraction model is efficient, and can extract the deep semantic features of the technical theme, which helps to improve the accuracy of the nodes to be merged determined subsequently, and thus improve the accuracy of the constructed technical theme tree.
[0165] For step S205, when merging nodes, the merging of the counting results of the nodes is also involved.
[0166] Specifically, in this initial weighted directed graph, add the counting results of each node in the nodes to be merged to obtain the merged counting result, and update the counting result of the target node to this merged counting result.
[0167] Follow Figure 4Example, suppose the counting result of node a is 15, the counting result of node b is 12, the counting result of node c is 10, the counting result of node d is 7, and the counting result of node e is 5. Suppose the merged target weighted directed graph is Figure 7 , then in the target weighted directed graph, the counting result of node a is 15, the counting result of node d is 12 + 10 + 7 = 29, and the counting result of node e is 5.
[0168] For step S206, the server can delete the self-loop edges in the target weighted directed graph;
[0169] For each node in the target weighted directed graph after deleting the self-loop edges, determine the importance of each incoming edge of the node according to the weights of the incoming edges of the node, so as to determine the parent-child relationship in the technical topic tree to be constructed according to the determined importance of each incoming edge;
[0170] Construct the technical topic tree of the target text sequence according to the nodes in the target weighted directed graph and the importance of each incoming edge of each node.
[0171] When merging nodes, if two adjacent nodes represent the merging of two technical topics, self-loop edges may be generated. Here, adjacent means that when there is only one edge connecting two nodes between two nodes, then these two nodes are adjacent. Self-loop edges may also appear. For example, if the technical topic is "A / B / B", then an edge of "B pointing to B" will appear in the structure of the target weighted directed graph. Such edges will cause hierarchical ambiguity between technical topics and affect the subsequent construction of the technical topic tree. Therefore, the server can first clean all self-loop edges in the graph to ensure that there is no repeated pointing relationship of the same-name technical topics.
[0172] After that, the server can determine the importance of each incoming edge to each node based on the weights of the incoming edges of each node in the target weighted directed graph after deleting the self-loop edges. The server can directly use the weights of the incoming edges of each node in the target weighted directed graph as the importance of each incoming edge to each node. As another alternative, the importance of a certain incoming edge of the node can also be determined according to the ratio of the weight of the incoming edge of the node to the sum of the weights of all incoming edges of the node.
[0173] Continue to use Figure 3 Example, for example, Figure 3 for node d in 0.54. The importance of the incoming edges of node b is 10 / 10 = 1, the importance of the incoming edges of node c is 8 / 8 = 1, the importance of the incoming edges of node e is 4 / 4 = 1, and node a has no incoming edges.
[0174] The importance of the incoming edges to a node can be used to determine the parent - child relationship in the subsequent technical topic tree, so that the connection relationships of the nodes in the constructed technical topic tree are of relatively high importance, and the constructed technical topic tree can better represent the target text sequence.
[0175] When constructing the technical topic tree of the target text sequence according to the importance of the incoming edges of each node and the target weighted directed graph, the server can initialize the stack structure;
[0176] By continuously performing the target operation until all the nodes in the target weighted directed graph are included in the constructed technical topic tree, the technical topic tree constructed by each target operation is determined as the technical topic tree of the target text sequence;
[0177] Among them, the target operation includes:
[0178] Determine the node with the smallest out - degree and / or the smallest sum of the weights of all outgoing edges in the current directed graph as the leaf node; where the current directed graph of the first target operation is the target weighted directed graph;
[0179] Take the leaf node as the first node to be processed in this target operation, and continuously perform the following operations on the node to be processed until the stop - pushing condition is met:
[0180] Push the node to be processed onto the stack; determine the incoming edge with the highest importance of the node to be processed, and determine the other node except the node to be processed connected by this incoming edge as the parent node of the node to be processed;
[0181] If the stop - pushing condition is not met, take the parent node of the node to be processed as the next node to be processed in this target operation;
[0182] If the stop - pushing condition is met, pop the nodes that have been pushed onto the stack, construct the technical topic tree corresponding to this target operation based on the popped nodes, delete the leaf node of this target operation from the current directed graph, and decrement the out - degree of the parent node of the leaf node by 1 to obtain the current directed graph of the next target operation.
[0183] Figure 8 It is a schematic flow diagram of a process for constructing a technical topic tree provided by an embodiment of the present application, as Figure 8 shown.
[0184] Continue to use Figure 3 the example to Figure 8It is described that the out-degrees of nodes a, b, c, d, and e are 2, 1, 1, 1, and 0 respectively. The node with the smallest out-degree is node e. Therefore, node e is determined as the leaf node. It should be noted that if the smallest out-degree corresponds to multiple nodes, the sum of the weights of all the out-edges included in these nodes can be further compared, and the node with the smallest sum of the weights of the out-edges is selected as the leaf node. After determining the leaf node of this target operation, the leaf node, that is, node e, can be placed into the initialized stack structure as the bottom node of the stack. Then, it is determined that the in-edge with the highest importance degree of node e is the in-edge connecting node e and node d. It can be understood that the other node except node e connected by the in-edge with the highest importance degree of node e is node d. Therefore, node d is determined as the parent node of node e.
[0185] The rule of selecting the parent node according to the importance can ensure that the technical theme tree is gradually constructed from the bottom up.
[0186] Among them, the stack-in stopping conditions include at least one of the following:
[0187] The in-degree of the node to be processed is 0;
[0188] The weight of the in-edge of the node to be processed is less than the fusing threshold.
[0189] Among them, the fusing threshold can also be called the reference weight or the benchmark weight. Optionally, the fusing threshold can be a preset threshold, such as a preset empirical value or experimental value. As another alternative, the fusing threshold can be determined based on the weights of the in-edges of each node in the stack, specifically as follows:
[0190]
[0191] Among them, n is the number of nodes in the stack, is the weight of the directed edge from the (i + 1)-th node to the i-th node in the stack, and this directed edge is the in-edge of the i-th node.
[0192] As an example, the in-degree of node e is 1, and the weight of the in-edge of node e is 4, which is greater than the current fusing threshold , since the condition for stopping pushing onto the stack is not met, the parent node of node e, which is node d, can be used as the node to be processed in the next target operation. Push node d onto the stack, select the other node besides node d connected by the incoming edge with a higher importance, which is node c, and determine node c as the parent node of node d. Still, the condition for stopping pushing onto the stack is not met. Push node c onto the stack, determine the parent node of node c as node b. Still not meeting the condition for stopping pushing onto the stack, push node b onto the stack, determine the parent node of node b as node a. Still not meeting the condition for stopping pushing onto the stack, push node a onto the stack. At this time, node a has no parent node and an in-degree of 0, meeting the condition for stopping pushing onto the stack. The nodes a, b, c, d, and e in the stack can be taken out in the last-in, first-out manner, and then the technical theme tree corresponding to this target operation can be constructed.
[0193] It can be understood that if the in-degree of the node to be processed is not 0 and / or the weight of the incoming edge of the node to be processed is not less than the fusing threshold, the target operation can continue to be executed until the in-degree of the node to be processed is 0 and / or the weight of the incoming edge of the node to be processed is less than the fusing threshold, at which point the stack push stops, and the technical theme tree is constructed based on the nodes in the stack.
[0194] Figure 9 This is a schematic structural diagram of the technical theme tree provided by the embodiment of the present application, as Figure 9 shown.
[0195] Determine the top node of the stack, which is node a, as the root of the tree, and the bottom node of the stack, which is node e, as the leaf node. Based on the parent-child relationship between the nodes, the technical theme tree is constructed.
[0196] All the nodes in the target weighted directed graph are included in this technical theme tree. Therefore, the construction stops.
[0197] It should be noted that if after constructing a technical theme tree, the technical theme tree does not include all the nodes in the target weighted directed graph, then to ensure that a new leaf node can be correctly searched for each time the target operation is executed, before each execution of the target operation, the bottom node of the current stack, that is, the leaf node, is removed from the current directed graph, and the out-degree of the parent node of this leaf node is decreased by 1 to obtain the current directed graph for the next target operation.
[0198] In addition, when the weight of the incoming edge of the node to be processed is less than the fusing threshold, the server can use the node adjacent to the last node pushed onto the stack in the stack as the root, and construct a technical theme tree based on the nodes pushed onto the stack before the last node pushed onto the stack;
[0199] Use the last node pushed onto the stack as the root to construct a technical theme tree to obtain the technical theme tree corresponding to this target operation.
[0200] For example, the stack includes nodes a, b, and c, and node a is the top node of the stack. Therefore, node a can be regarded as a technical topic tree, node b can be regarded as the root of another technical topic tree, and the leaf node is node c.
[0201] Since the large language model may generate some uncommon or accidental hallucinations during the process of generating topic information, resulting in a chaotic hierarchy between technical topics or a hierarchical relationship between irrelevant technical topics. At this time, the weight of the incoming edge of the node to be processed is less than the fusing threshold. It can also be understood that the importance of the hierarchical relationship between the node to be processed and other nodes in the current stack is relatively low, and the relevance is relatively low. Therefore, a technical topic tree for the node to be processed can be constructed separately to avoid forced combination and obtain a technical topic tree with relatively low accuracy.
[0202] For example, "Electronic Engineering / Medical Field / Heart Imaging". The relevance between the medical field and electronic engineering is small, and it has only appeared once in the output of the large language model. Moreover, the medical field is already a topic with rich enough sub-levels and sub-nodes. Therefore, a rule is needed to disconnect it. Suppose the weight of the edge from electronic engineering to the medical field is 1, and the weight of the edge from the medical field to heart imaging is 5, then the fusing threshold is . Fuse and disconnect the edge from electronic engineering to the medical field. Electronic engineering is the root node to form a technical topic tree, and the medical field is determined as the root node of another technical topic tree, with the leaf node being heart imaging.
[0203] It should be noted that when determining the incoming edge with the highest importance of the node to be processed and determining the other node connected by this incoming edge (excluding the node to be processed) as the parent node of the node to be processed, if there is no other node connected by this incoming edge (excluding the node to be processed) in the stack, then this other node is determined as the parent node of the node to be processed.
[0204] If there is another node connected by this incoming edge (excluding the node to be processed) in the stack, then the other node connected by the target incoming edge (excluding the node to be processed) is determined as the parent node of the node to be processed, where the target incoming edge is the incoming edge with the highest importance among the incoming edges whose connected other nodes are not in the stack among other incoming edges, and the other incoming edges are the incoming edges of the node to be processed other than the incoming edge with the highest importance.
[0205] If the parent node of the node to be processed exists in the stack, it means that there is a cyclic structure in the target weighted directed graph, such as a pointing to b, b pointing to c, and c pointing to a. The cyclic structure will cause chaos in the technical topic tree. Therefore, the server can determine other parent nodes that are not in the stack to avoid the occurrence of a cyclic structure.
[0206] For example, currently the node to be processed is node b, and the determined parent node is node a. Among all the incoming edges connecting node a and node b, the importance ranking is the first. However, node a is already in the stack, and node b also has another incoming edge, and the node of the other incoming edge is node d, which is not in the stack and ranks first in importance among other incoming edges. Then, node d is determined as the parent node of node b, and the target operation is continued.
[0207] By calculating the importance of incoming edges, dynamically generating a stack, and multiple technical topic trees, a technical topic tree with a clear hierarchy and no loops can be extracted from the target weighted directed graph. Through noise cleaning, such as removing self-loop edges and circular chain structures, the high quality and consistency of the generated technical topic tree are ensured. During the process of constructing the technical topic tree, the bottom-up construction logic is followed, progressing layer by layer from leaf nodes to root nodes, and finally organizing all nodes into multiple independent technical topic trees.
[0208] It can be understood that the embodiments of the present application adopt a large language model and a semantic extraction model to achieve the full-process automation from unstructured target text to hierarchicalization, reduce manual participation, and greatly improve efficiency and accuracy. Through the importance of incoming edges and noise cleaning, the clarity of the hierarchical structure of the technical topic tree and the rationality of the semantic logic are ensured. The finally obtained technical topic tree of the target text sequence more accurately reflects the semantic associations of technical topics and supports deeper knowledge mining. By extracting semantic features and calculating semantic similarity, technical topics with similar semantics can be more precisely merged, and redundant or noisy nodes can be removed. When processing a large amount of text, high-quality technical topics can still be generated, avoiding problems such as chaotic technical topics or unreasonable hierarchies. The method for determining the retained technical topics based on the counting results enables the embodiments of the present application to exhibit excellent performance when processing large-scale text, can quickly process a large amount of data, and ensure that the generated technical topic tree is complete and acyclic.
[0209] The finally obtained technical topic tree can also be displayed to the user, providing the user with a more intuitive view of the technical topic structure and supporting various application scenarios such as technical intelligence analysis and technical innovation navigation.
[0210] The embodiments of the present application provide a device for generating a technical topic tree, such as Figure 10 shown. The device 100 for generating a technical topic tree may include: a target text sequence acquisition module 1001, a topic information extraction module 1002, an initial weighted directed graph determination module 1003, an information to be merged determination module 1004, a merging module 1005, and a topic tree construction module 1006, where:
[0211] The target text sequence acquisition module 1001 is configured to acquire a target text sequence, and the target text sequence includes multiple target texts;
[0212] The subject information extraction module 1002 is used to extract subject information of each target text respectively through a large language model, wherein, for each target text, the subject information of the target text includes at least one technical subject of the target text, and when the target text has multiple technical subjects, the subject information of the target text also includes a hierarchical relationship between the multiple technical subjects;
[0213] An initial weighted directed graph determination module 1003 is used to determine an initial weighted directed graph corresponding to the target text sequence according to the subject information of each of the target texts, wherein each node in the initial weighted directed graph represents a technical subject, an edge between two nodes represents a hierarchical relationship between the technical subjects corresponding to the two nodes, and the weight of an edge represents the occurrence frequency of the hierarchical relationship corresponding to the two nodes connected by the edge in the hierarchical relationship extracted from the multiple target texts;
[0214] The module 1004 for determining information to be merged is used to determine the semantic similarity between the technical themes represented by the nodes in the initial weighted directed graph, determine the semantic similarity between the technical themes that meet the preset conditions as the target similarity, and determine the nodes corresponding to each target similarity as the nodes to be merged;
[0215] A merging module 1005 is used to merge the nodes to be merged, the edges associated with the nodes to be merged, and the weights of the edges in the initial weighted directed graph to obtain a target weighted directed graph;
[0216] A topic tree construction module 1006 is used to construct a technical topic tree of the target text sequence according to the target weighted directed graph;
[0217] The preset condition includes at least one of the following:
[0218] The semantic similarity is not less than the preset similarity threshold;
[0219] The similarities are sorted in descending order, with the semantic similarities of the preset number at the top.
[0220] Optionally, the topic information extraction module 1002 may be used to obtain preset first prompt information, where the first prompt information is used to prompt the large language model to output topic information of the target text;
[0221] For a first target text in the target text sequence, extracting topic information of the target text through a large language model based on the first prompt information and the target text;
[0222] For each target text except the first target text in the target text sequence, based on the extracted technical themes and the first prompt information, determine second prompt information, where the second prompt information is used to prompt the large language model to output the theme information of the target text with reference to the extracted technical themes. Based on the target text and the second prompt information, extract the theme information of the target text through the large language model. The extracted technical themes are the technical themes of the previous target texts in the target text sequence before this target text.
[0223] Optionally, the initial weighted directed graph determination module 1003 can be used to construct a first weighted directed graph according to the theme information of the first target text;
[0224] For each target text except the first target text, every time the theme information of a target text is extracted, update the constructed weighted directed graph according to the theme information of the target text until a weighted directed graph updated based on the theme information of the last target text is obtained, and determine the weighted directed graph as the initial weighted directed graph corresponding to the target text sequence.
[0225] Optionally, the initial weighted directed graph determination module 1003 can be used to count the first quantity of the extracted technical themes and determine the number of occurrences of each technical theme to obtain the count result of each technical theme;
[0226] If the first quantity is greater than a preset value, delete the theme information corresponding to the technical themes ranked after the preset value according to the descending order of the count results of the extracted technical themes;
[0227] Based on the theme information of the remaining target texts, determine the initial weighted directed graph corresponding to the target text sequence.
[0228] Optionally, the theme tree construction module 1006 can be used to delete the self-loop edges in the target weighted directed graph;
[0229] For each node in the target weighted directed graph after deleting the self-loop edges, determine the importance of each incoming edge of the node according to the weights of the incoming edges of the node, so as to determine the parent-child relationship in the technical theme tree to be constructed according to the determined importance of each incoming edge;
[0230] Construct the technical theme tree of the target text sequence according to the nodes in the target weighted directed graph and the importance of each incoming edge of each node.
[0231] Optionally, the theme tree construction module 1006 can be used to initialize the stack structure;
[0232] By continuously performing the target operation until all nodes in the target weighted directed graph are included in the constructed technical topic tree, the technical topic tree constructed by each target operation is determined as the technical topic tree of the target text sequence;
[0233] Among them, the target operation includes:
[0234] Determine the node with the minimum out-degree and / or the minimum sum of the weights of each out-edge in the current directed graph as the leaf node; among them, the current directed graph of the first target operation is the target weighted directed graph;
[0235] Take the leaf node as the first node to be processed in this target operation, and continuously perform the following operations on the node to be processed until the stop pushing condition is met:
[0236] Push the node to be processed onto the stack; determine the in-edge with the highest importance of the node to be processed, and determine the other node connected by this in-edge except the node to be processed as the parent node of the node to be processed;
[0237] If the stop pushing condition is not met, take the parent node of the node to be processed as the next node to be processed in this target operation;
[0238] If the stop pushing condition is met, pop all the nodes on the stack, construct the technical topic tree corresponding to this target operation based on the popped nodes, delete the leaf node of this target operation from the current directed graph, and decrement the out-degree of the parent node of the leaf node by 1 to obtain the current directed graph of the next target operation.
[0239] Optionally, the topic tree construction module 1006 may be used to determine the other node as the parent node of the node to be processed if there is no other node connected by this in-edge except the node to be processed in the stack;
[0240] The topic tree construction module 1006 may be used to determine the other node connected by the target in-edge except the node to be processed as the parent node of the node to be processed if there is the other node connected by this in-edge except the node to be processed in the stack, where the target in-edge is the in-edge with the highest importance among the other in-edges where the other node is not in the stack, and the other in-edges are the in-edges of the node to be processed except the in-edge with the highest importance.
[0241] Optionally, the stop pushing condition includes at least one of the following:
[0242] The in-degree of the node to be processed is 0;
[0243] The weight of the in-edge of the node to be processed is less than the fusing threshold;
[0244] When the weight of the incoming edge of the node to be processed is less than the fusing threshold, the technical topic tree construction module 1006 can be used to take the node adjacent to the last pushed node in the stack as the root, and construct a technical topic tree based on each node pushed into the stack before the last pushed node;
[0245] Take the last pushed node as the root, construct a technical topic tree, and obtain the technical topic tree corresponding to the target operation this time.
[0246] Optionally, the information to be merged determination module 1004 can be used to input the technical topic represented by each node in the initial weighted directed graph into the semantic extraction model, and obtain the semantic features of the technical topic represented by this node output by the semantic extraction model;
[0247] According to the semantic features of the technical topic represented by this node, determine the semantic similarity between the technical topic represented by this node and the technical topics represented by other nodes in the initial weighted directed graph.
[0248] Optionally, for each target similarity, the merging module 1005 can be used to take the node pair corresponding to this target similarity as the set to be merged if there are no identical nodes in the node pair corresponding to this target similarity and the node pairs corresponding to other target similarities, and take the nodes in the node pairs with identical nodes in the node pairs corresponding to this target similarity and the node pairs corresponding to other target similarities as the set to be merged;
[0249] For each set to be merged, input the technical topics represented by the nodes in this set to be merged and the preset task prompt data into the topic optimization model, and obtain the target node of this set to be merged. The task prompt data is used to prompt the topic optimization model to screen out the target node from the input nodes;
[0250] For non-target nodes in each set to be merged, update the edges and the weights of the edges between the other nodes except the nodes in this set to be merged and this non-target node to the edges and the weights of the edges between the other nodes and the target node, and delete the edges, the weights of the edges, and the non-target nodes between the target node and the non-target nodes in this set to be merged, to obtain the target weighted directed graph.
[0251] The device of the embodiment of the present application can execute the method provided by the embodiment of the present application, and its implementation principle is similar and has corresponding technical effects. The actions performed by each module in the device of each embodiment of the present application correspond to the steps in the method of each embodiment of the present application. For the detailed function descriptions of the modules of the device, reference can specifically be made to the descriptions in the corresponding methods shown above, and details are not described herein again.
[0252] In an embodiment of the present application, an electronic device is provided, which includes a memory, a processor, and a computer program stored in the memory. The processor executes the above computer program to implement the steps of the method provided in any optional embodiment of the present application. Compared with the prior art, the following can be achieved: The present application obtains a target text sequence, which includes multiple target texts. Based on a large language model, it realizes the automatic extraction of the theme information of the target text, improving the efficiency of obtaining the theme information of the target text. Based on the theme information, an initial weighted directed graph is constructed. The initial weighted directed graph includes the hierarchical relationship and weights between technical themes, which is conducive to subsequent determination of the technical theme tree and improves the efficiency of constructing the technical theme tree. The semantic similarity of the technical themes represented by the nodes is determined, and the nodes can be merged according to the semantic similarity to obtain a target weighted directed graph, reducing redundant nodes, making the extracted technical themes more accurate, and the hierarchy between the technical themes obtained from the target weighted directed graph is clearer and the semantic logic is stronger.
[0253] In an optional embodiment, an electronic device is provided, as Figure 11 shown, as Figure 11 shown in Figure 11 in which, the electronic device 2000 mainly includes at least one processor 2001 ( Figure 11 one is shown in
[0254] Among them, the memory 2002 can be used to store the operating system, application programs, etc. The application programs can include computer programs that implement the methods shown in the embodiments of the present application when called by the processor 2001, and can also include programs for implementing other functions or services. The memory 2002 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and computer programs. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0255] The processor 2001 is connected to the memory 2002 through the bus 2005 and realizes corresponding functions by calling the application programs stored in the memory 2002. Among them, the processor 2001 can be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof, which can implement or execute various exemplary logic blocks, modules, and circuits described in connection with the disclosure of the present application. The processor 2001 can also be a combination that realizes computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0256] The electronic device 2000 can be connected to a network through a communication module 2003 (which may include, but is not limited to, components such as a network interface) to communicate with other devices (such as user terminals or servers) through the network and achieve data interaction, such as sending data to other devices or receiving data from other devices. Among them, the communication module 2003 may include a wired network interface and / or a wireless network interface, etc., that is, the communication module may include at least one of a wired communication module or a wireless communication module.
[0257] The electronic device 2000 can be connected to the required input / output devices, such as a keyboard, a display device, etc., through an input / output interface 2004. The electronic device 2000 itself may have a display device and may also externally connect other display devices through the interface 2004. Optionally, a storage device, such as a hard disk, etc., can also be connected through this interface 2004 to store the data in the electronic device 2000 into the storage device, or read the data in the storage device, and the data in the storage device can also be stored in the memory 2002. It can be understood that the input / output interface 2004 can be a wired interface or a wireless interface. According to different actual application scenarios, the devices connected to the input / output interface 2004 can be components of the electronic device 2000 or external devices connected to the electronic device 2000 when needed.
[0258] The bus 2005 for connecting each component may include a path to transfer information between the above components. The bus 2005 can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. According to different functions, the bus 2005 can be divided into an address bus, a data bus, a control bus, etc.
[0259] Optionally, for the solution provided in the embodiments of the present application, the memory 2002 can be used to store a computer program for executing the solution of the present application, and the processor 2001 runs it. When the processor 2001 runs this computer program, it implements the actions of the method or device provided in the embodiments of the present application.
[0260] The embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can implement the steps and corresponding content of the foregoing method embodiments.
[0261] The embodiments of the present application also provide a computer program product, including a computer program. When the computer program is executed by a processor, it can implement the steps and corresponding content of the foregoing method embodiments.
[0262] It should be understood that although the flowcharts of the embodiments of the present application indicate various operation steps by arrows, the execution order of these steps is not limited to the order indicated by the arrows. Unless there is a clear description in this article, in some implementation scenarios of the embodiments of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage among these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present application do not limit this.
[0263] The above are only optional implementation manners of some implementation scenarios of the present application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the technical concept of the solution of the present application, using other similar implementation means based on the technical idea of the present application also belongs to the protection scope of the embodiments of the present application.
Claims
1. A method for generating a technical theme tree, characterized in that include: Acquire a target text sequence, wherein the target text sequence includes a plurality of target texts; By using a large language model, extracting the subject information of each target text respectively, wherein, for each target text, the subject information of the target text includes at least one technical subject of the target text, and when the target text has multiple technical subjects, the subject information of the target text also includes the hierarchical relationship between the multiple technical subjects; Determine an initial weighted directed graph corresponding to the target text sequence according to the subject information of each of the target texts, wherein each node in the initial weighted directed graph represents a technical subject, an edge between two nodes represents a hierarchical relationship between the technical subjects corresponding to the two nodes, and the weight of an edge represents the occurrence frequency of the hierarchical relationship corresponding to the two nodes connected by the edge in the hierarchical relationship extracted from the multiple target texts; Determine the semantic similarity between the technical themes represented by each node in the initial weighted directed graph, determine the semantic similarity between the technical themes that meet the preset conditions as the target similarity, and determine each node corresponding to each target similarity as the node to be merged; Merging the nodes to be merged, the edges associated with the nodes to be merged, and the weights of the edges in the initial weighted directed graph to obtain a target weighted directed graph; Constructing a technical theme tree of the target text sequence according to the target weighted directed graph; The preset condition includes at least one of the following: The semantic similarity is not less than the preset similarity threshold; The similarities are sorted in descending order, with the semantic similarities of the preset number at the top.
2. The method according to claim 1, characterized in that, The extracting the subject information of each target text respectively through the large language model includes: Obtaining preset first prompt information, where the first prompt information is used to prompt the large language model to output topic information of the target text; For a first target text in the target text sequence, extracting topic information of the target text through a large language model based on the first prompt information and the target text; For each target text except the first target text in the target text sequence, second prompt information is determined based on the extracted technical theme and the first prompt information, the second prompt information is used to prompt the large language model to output the theme information of the target text with reference to the extracted technical theme, and the theme information of the target text is extracted through the large language model based on the target text and the second prompt information, the extracted technical theme is the technical theme of each target text before the target text in the target text sequence.
3. The method according to claim 2, characterized in that, Determining an initial weighted directed graph corresponding to the target text sequence according to the subject information of each target text includes: Constructing a first weighted directed graph according to the subject information of the first target text; For each target text except the first one, every time the theme information of a target text is extracted, the constructed weighted directed graph is updated according to the theme information of the target text until the weighted directed graph updated based on the theme information of the last target text is obtained, and this weighted directed graph is determined as the initial weighted directed graph corresponding to the target text sequence.
4. The method according to any one of claims 1 to 3, characterized in that, Determining the initial weighted directed graph corresponding to the target text sequence according to the theme information of each target text includes: Counting the first quantity of the extracted technical themes and determining the number of occurrences of each technical theme to obtain the counting result of each technical theme; If the first quantity is greater than a preset value, delete the theme information corresponding to the technical themes ranked after the preset value according to the order of the counting results of the extracted technical themes from large to small; Based on the theme information of the remaining target texts, determine the initial weighted directed graph corresponding to the target text sequence.
5. The method according to claim 4, characterized in that, Constructing the technical theme tree of the target text sequence according to the target weighted directed graph includes: Deleting the self-loop edges in the target weighted directed graph; For each node in the target weighted directed graph after deleting the self-loop edges, determine the importance of each incoming edge of the node according to the weights of the incoming edges of the node, so as to determine the parent-child relationship in the technical theme tree to be constructed according to the determined importance of each incoming edge; Construct the technical theme tree of the target text sequence according to the nodes in the target weighted directed graph and the importance of the incoming edges of each node.
6. The method according to claim 5, characterized in that, Constructing the technical theme tree of the target text sequence according to the nodes in the target weighted directed graph and the importance of the incoming edges of each node includes: Initializing the stack structure; By continuously performing the target operation until the constructed technical theme tree includes all the nodes in the target weighted directed graph, determine the technical theme tree constructed by each target operation as the technical theme tree of the target text sequence; Among them, the target operation includes: Determine the node with the minimum out-degree and / or the minimum sum of the weights of the out-edges in the current directed graph as the leaf node; where the current directed graph of the first target operation is the target weighted directed graph; Take the leaf node as the first node to be processed in this target operation, and continuously perform the following operations on the node to be processed until the stop condition for pushing into the stack is met: Push the node to be processed onto the stack; determine the incoming edge with the highest importance of the node to be processed, and determine the other node connected by the incoming edge except the node to be processed as the parent node of the node to be processed; If the stop condition for pushing into the stack is not met, take the parent node of the node to be processed as the next node to be processed in this target operation; If the stop condition for pushing into the stack is met, pop the nodes that have been pushed onto the stack, construct the technical theme tree corresponding to this target operation based on the popped nodes, delete the leaf node of this target operation from the current directed graph, and subtract 1 from the out-degree of the parent node of the leaf node to obtain the current directed graph of the next target operation.
7. The method according to claim 6, wherein Determining the incoming edge with the highest importance for the node to be processed, and determining the other node connected by this incoming edge except the node to be processed as the parent node of the node to be processed, includes: If there is no other node connected by this incoming edge in the stack except the node to be processed, then determine this other node as the parent node of the node to be processed; The method further includes: If there is other node connected by this incoming edge in the stack except the node to be processed, then determine the other node connected by the target incoming edge except the node to be processed as the parent node of the node to be processed, where the target incoming edge is the incoming edge with the highest importance among the incoming edges whose connected other nodes are not in the stack, and the other incoming edges are the incoming edges of the node to be processed except the incoming edge with the highest importance.
8. The method according to claim 6, wherein The stack stop condition includes at least one of the following: The in-degree of the node to be processed is 0; The weight of the incoming edge of the node to be processed is less than the fusing threshold; When the weight of the incoming edge of the node to be processed is less than the fusing threshold, pop out each node in the stack, and construct a technical theme tree corresponding to the target operation this time based on the popped-out nodes, including: Taking the node adjacent to the last pushed-in node in the stack as the root, and constructing a technical theme tree based on each node pushed in before the last pushed-in node; Taking the last pushed-in node as the root, and constructing a technical theme tree to obtain the technical theme tree corresponding to the target operation this time.
9. The method according to claim 1, characterized in that, Determining the semantic similarity between the technical themes represented by each node in the initial weighted directed graph, includes: For each node in the initial weighted directed graph, input the technical theme represented by this node into the semantic extraction model, and obtain the semantic feature of the technical theme represented by this node output by the semantic extraction model; According to the semantic feature of the technical theme represented by this node, determine the semantic similarity between the technical theme represented by this node and the technical themes represented by other nodes in the initial weighted directed graph.
10. The method according to claim 1, characterized in that Merging the nodes to be merged, the edges associated with the nodes to be merged, and the weights of the edges in the initial weighted directed graph to obtain a target weighted directed graph, includes: For each target similarity, if the node pair corresponding to this target similarity has no same nodes as the node pairs corresponding to other target similarities, then take the node pair corresponding to this target similarity as the set to be merged, if the node pair corresponding to this target similarity has same nodes as the node pairs corresponding to other target similarities, then take the nodes in the node pairs with same nodes as the set to be merged; For each set to be merged, input the technical themes represented by the nodes in this set to be merged and the preset task prompt data into the theme optimization model, and obtain the target node of this set to be merged, where the task prompt data is used to prompt the theme optimization model to screen out the target node from the input nodes. For each non-target node in the set to be merged, the edges and edge weights between other nodes except the node in the set to be merged and the non-target node are updated to the edges and edge weights between other nodes and the target node, and the edges, edge weights and non-target nodes between the target node and the non-target node in the set to be merged are deleted to obtain the target weighted directed graph.
11. An apparatus for generating a technical topic tree, characterized in that include: A target text sequence acquisition module, used to acquire a target text sequence, wherein the target text sequence includes a plurality of target texts; A topic information extraction module, used to extract topic information of each target text respectively through a large language model, wherein, for each target text, the topic information of the target text includes at least one technical topic of the target text, and when the target text has multiple technical topics, the topic information of the target text also includes a hierarchical relationship between the multiple technical topics; An initial weighted directed graph determination module is used to determine an initial weighted directed graph corresponding to the target text sequence according to the subject information of each of the target texts, wherein each node in the initial weighted directed graph represents a technical subject, an edge between two nodes represents a hierarchical relationship between the technical subjects corresponding to the two nodes, and the weight of an edge represents the occurrence frequency of the hierarchical relationship corresponding to the two nodes connected by the edge in the hierarchical relationship extracted from the multiple target texts; The module for determining information to be merged is used to determine the semantic similarity between the technical themes represented by each node in the initial weighted directed graph, determine the semantic similarity between the technical themes that meet the preset conditions as the target similarity, and determine each node corresponding to each target similarity as the node to be merged; A merging module, used for merging the nodes to be merged, the edges associated with the nodes to be merged, and the weights of the edges in the initial weighted directed graph to obtain a target weighted directed graph; A topic tree construction module, used to construct a technical topic tree of the target text sequence according to the target weighted directed graph; The preset condition includes at least one of the following: The semantic similarity is not less than the preset similarity threshold; The similarities are sorted in descending order, with the semantic similarities of the preset number at the top.
12. An electronic device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 10.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
14. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
Citation Information
Patent Citations
Mining sequential patterns in weighted directed graphs
US20100251210A1
Generating descriptive topic labels
US20170103074A1