Technical theme tree generation method and device and electronic equipment

The text theme is extracted through a large language model and a weighted directed graph is solved, and the problems of inaccurate technical topics and fuzzy hierarchical relationships in the existing technology are achieved, and a more accurate and clear technical topic tree construction is achieved.

CN120068849AActive Publication Date: 2025-05-30INST OF SCI & TECHN INFORMATION OF CHINA
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510197271.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-30
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

When extracting technical topics of text, the topics are inaccurate and the hierarchical relationships are blurred, resulting in a waste of time for users to understand the core content of the text.

Method used

The topic information of the target text is extracted through a large language model, an initial weighted directed graph is constructed, the hierarchical relationship and semantic similarity between technical topics are determined, node merging is performed, and the target weighted directed graph is generated, and a technical topic tree is constructed.

Benefits of technology

It improves the accuracy of technical topic extraction and the clarity of hierarchical relationships, reduces redundant nodes, and the extracted technical topics are more accurate, the hierarchical relationships are clearer, and the semantic logic is stronger.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068849A_ABST
    Figure CN120068849A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a technical topic tree generation method and device and electronic equipment, and relates to the technical field of text processing. The method comprises the following steps: acquiring a target text sequence, wherein the sequence comprises a plurality of target texts; based on the large language model, the subject information of the target text is automatically extracted, and the efficiency of obtaining the subject information of the target text is improved. According to the method, the initial weighted directed graph is constructed based on the topic information, the initial weighted directed graph comprises the hierarchical relationship and the weight between the technical topics, the technical topic tree is easy to subsequently determine, and the efficiency of constructing the technical topic tree is improved. According to the technical scheme, the semantic similarity of the technical themes represented by the nodes is determined, the nodes can be merged according to the semantic similarity, the target weighted directed graph is obtained, redundant nodes are reduced, the extracted technical themes are more accurate, the hierarchy between the technical themes obtained according to the target weighted directed graph is clearer, and the semantic logicality is higher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of text processing. Specifically, the present application relates to a method, apparatus, and electronic device for generating a technical topic tree. Background Art

[0002] When the number of texts is large and the length of each text is long, users need to spend a lot of time understanding the texts. Currently, by extracting the technical topics of the texts, such as information on the technical fields and technical means of the texts, it is convenient for users to understand the core content of the texts and save the time for users to understand the texts.

[0003] There may also be a hierarchical relationship between the technical topics, but the technical topics of the texts extracted by the related technologies are still inaccurate, and the hierarchical relationship between the technical topics also needs to be optimized. Summary of the Invention

[0004] Embodiments of the present application provide a method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating a technical topic tree, which can solve the problem that the technical topics of the texts extracted by the related technologies are still inaccurate, and the hierarchical relationship between the technical topics also needs to be optimized. The technical solutions provided by the present application are as follows: According to one aspect of the embodiments of the present application, a method for generating a technical topic tree is provided. The method includes: Obtain a target text sequence, where the target text sequence includes a plurality of target texts; Extract the topic information of each target text through a large language model. Wherein, for each target text, the topic information of the target text includes at least one technical topic of the target text. When there are multiple technical topics of the target text, the topic information of the target text further includes the hierarchical relationship between the multiple technical topics; Determine an initial weighted directed graph corresponding to the target text sequence according to the topic information of each target text. Wherein, each node in the initial weighted directed graph represents a technical topic, the edge between two nodes represents the hierarchical relationship between the technical topics corresponding to the two nodes, and the weight of an edge represents the occurrence frequency of the hierarchical relationship between the two nodes connected by the edge in the hierarchical relationship extracted from the multiple target texts; Determine the semantic similarity between the technical topics represented by each node in the initial weighted directed graph, determine the semantic similarity between the technical topics that meet the preset conditions as the target similarity, and determine each node corresponding to each target similarity as the node to be merged; Merge the nodes to be merged, the edges associated with the nodes to be merged, and the weights of the edges in the initial weighted directed graph to obtain a target weighted directed graph; Construct a technical theme tree for the target text sequence according to the target weighted directed graph; Wherein, the preset conditions include at least one of the following: The semantic similarity is not less than a preset similarity threshold; In the order of sorting the similarities from large to small, the ones with higher rankings are the semantic similarities of a preset number.

[0005] According to another aspect of the embodiments of the present application, a technical theme tree generation device is provided, and the device includes: A target text sequence acquisition module, configured to acquire a target text sequence, where the target text sequence includes multiple target texts; A theme information extraction module, configured to extract the theme information of each target text through a large language model. Wherein, for each target text, the theme information of the target text includes at least one technical theme of the target text. When there are multiple technical themes of the target text, the theme information of the target text further includes the hierarchical relationship between the multiple technical themes; An initial weighted directed graph determination module, configured to determine an initial weighted directed graph corresponding to the target text sequence according to the theme information of each target text. Wherein, each node in the initial weighted directed graph represents a technical theme, the edge between two nodes represents the hierarchical relationship between the technical themes corresponding to the two nodes, and the weight of an edge represents the occurrence frequency of the hierarchical relationship between the two nodes connected by the edge in the hierarchical relationship extracted from the multiple target texts; A to-be-merged information determination module, configured to determine the semantic similarity between the technical themes represented by each node in the initial weighted directed graph, determine the semantic similarity between the technical themes that meet the preset conditions as the target similarity, and determine each node corresponding to each target similarity as the node to be merged; A merging module, configured to merge the nodes to be merged in the initial weighted directed graph, the edges associated with the nodes to be merged, and the weights of the edges to obtain a target weighted directed graph; A theme tree construction module, configured to construct a technical theme tree for the target text sequence according to the target weighted directed graph; Wherein, the preset conditions include at least one of the following: The semantic similarity is not less than a preset similarity threshold; In the order of sorting the similarities from large to small, the ones with higher rankings are the semantic similarities of a preset number.

[0006] Optionally, the theme information extraction module may be configured to obtain preset first prompt information, and the first prompt information is used to prompt the large language model to output the theme information of the target text; For the first target text in the target text sequence, based on the first prompt information and this target text, use a large language model to extract the theme information of this target text; For each target text except the first target text in the target text sequence, based on the extracted technical theme and the first prompt information, determine second prompt information, where the second prompt information is used to prompt the large language model to output the theme information of this target text with reference to the extracted technical theme. Based on this target text and the second prompt information, use a large language model to extract the theme information of this target text, and the extracted technical theme is the technical theme of each target text before this target text in the target text sequence.

[0007] Optionally, the initial weighted directed graph determination module can be used to construct a first weighted directed graph according to the theme information of the first target text; For each target text except the first target text, every time the theme information of a target text is extracted, update the constructed weighted directed graph according to the theme information of this target text until a weighted directed graph updated based on the theme information of the last target text is obtained, and determine this weighted directed graph as the initial weighted directed graph corresponding to the target text sequence.

[0008] Optionally, the initial weighted directed graph determination module can be used to count the first quantity of the extracted technical themes and determine the number of times each technical theme appears to obtain the count result of each technical theme; If the first quantity is greater than a preset value, delete the theme information corresponding to the technical themes ranked after the preset value according to the order of the count results of the extracted technical themes from largest to smallest; Based on the theme information of the remaining target texts, determine the initial weighted directed graph corresponding to the target text sequence.

[0009] Optionally, the theme tree construction module can be used to delete the self-loop edges in the target weighted directed graph; For each node in the target weighted directed graph after deleting the self-loop edges, determine the importance of each incoming edge of this node according to the weights of each incoming edge of this node, so as to determine the parent-child relationship in the technical theme tree to be constructed according to the determined importance of each incoming edge; Construct the technical theme tree of the target text sequence according to the nodes in the target weighted directed graph and the importance of each incoming edge of each node.

[0010] Optionally, the theme tree construction module can be used to initialize the stack structure; By continuously performing the target operation until all nodes in the target weighted directed graph are included in the constructed technical topic tree, determining the technical topic tree constructed by each target operation as the technical topic tree of the target text sequence; Among them, the target operation includes: Determining the node with the smallest out-degree and / or the smallest sum of the weights of each out-edge in the current directed graph as the leaf node; where the current directed graph of the first target operation is the target weighted directed graph; Taking the leaf node as the first node to be processed in this target operation, and continuously performing the following operations on the node to be processed until the stop pushing condition is met: Pushing the node to be processed onto the stack; determining the in-edge with the highest importance of the node to be processed, and determining the other node connected by this in-edge except the node to be processed as the parent node of the node to be processed; If the stop pushing condition is not met, taking the parent node of the node to be processed as the next node to be processed in this target operation; If the stop pushing condition is met, popping all the nodes on the stack, constructing the technical topic tree corresponding to this target operation based on the popped nodes, deleting the leaf node of this target operation from the current directed graph, and reducing the out-degree of the parent node of the leaf node by 1 to obtain the current directed graph of the next target operation.

[0011] Optionally, the topic tree construction module can be used to determine the other node as the parent node of the node to be processed if there is no other node connected by this in-edge except the node to be processed in the stack; The topic tree construction module can be used to determine the other node connected by the target in-edge except the node to be processed as the parent node of the node to be processed if there is another node connected by this in-edge except the node to be processed in the stack, where the target in-edge is the in-edge with the highest importance among the in-edges where the other node connected by it is not in the stack, and the other in-edges are the in-edges of the node to be processed except the in-edge with the highest importance.

[0012] Optionally, the stop pushing condition includes at least one of the following: The in-degree of the node to be processed is 0; The weight of the in-edge of the node to be processed is less than the melting threshold; When the weight of the in-edge of the node to be processed is less than the melting threshold, popping all the nodes on the stack, and constructing the technical topic tree corresponding to this target operation based on the popped nodes, including: When the weight of the incoming edge of the node to be processed is less than the fusing threshold, the technical topic tree construction module can be used to use the node adjacent to the last pushed node in the stack as the root, and construct a technical topic tree based on each node pushed into the stack before the last pushed node; Use the last pushed node as the root to construct a technical topic tree, and obtain the technical topic tree corresponding to the target operation this time.

[0013] Optionally, the information to be merged determination module can be used to input the technical topic represented by each node in the initial weighted directed graph into the semantic extraction model, and obtain the semantic features of the technical topic represented by this node output by the semantic extraction model; According to the semantic features of the technical topic represented by this node, determine the semantic similarity between the technical topic represented by this node and the technical topics represented by other nodes in the initial weighted directed graph.

[0014] Optionally, the merging module can be used for each target similarity. If the node pair corresponding to this target similarity does not have the same node as the node pairs corresponding to other target similarities, then use the node pair corresponding to this target similarity as the set to be merged. If the node pair corresponding to this target similarity has the same node as the node pairs corresponding to other target similarities, then use the nodes in each node pair with the same node as the set to be merged; For each set to be merged, input the technical topics represented by the nodes in this set to be merged and the preset task prompt data into the topic optimization model, and obtain the target node of this set to be merged. The task prompt data is used to prompt the topic optimization model to screen out the target node from the input nodes; For non-target nodes in each set to be merged, update the edges and edge weights between the other nodes except the nodes in this set to be merged and this non-target node to the edges and edge weights between the other nodes and the target node, and delete the edges, edge weights and non-target nodes between the target node and the non-target node in this set to be merged, and obtain the target weighted directed graph.

[0015] According to another aspect of the embodiments of the present application, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the method provided in any optional embodiment of the present application.

[0016] According to still another aspect of the embodiments of the present application, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method provided in any optional embodiment of the present application are implemented.

[0017] According to one aspect of the embodiments of the present application, there is provided a computer program product including a computer program which, when executed by a processor, implements the steps of the method provided in any optional embodiment of the present application.

[0018] The beneficial effects brought by the technical solutions provided in the embodiments of the present application are as follows: The present application obtains a target text sequence, which includes a plurality of target texts. Based on the large language model, it realizes the automatic extraction of the theme information of the target text, improving the efficiency of obtaining the theme information of the target text. Based on the theme information, an initial weighted directed graph is constructed. The initial weighted directed graph includes the hierarchical relationship and weights between technical themes, which is conducive to subsequent determination of the technical theme tree and improves the efficiency of constructing the technical theme tree. By determining the semantic similarity of the technical themes represented by the nodes, node merging can be performed according to the semantic similarity to obtain a target weighted directed graph, reducing redundant nodes, making the extracted technical themes more accurate, and making the hierarchy between the technical themes obtained from the target weighted directed graph clearer and the semantic logic stronger. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description in the embodiments of the present application.

[0020] Figure 1 Schematic diagram of a system for generating a technical theme tree provided in an embodiment of the present application; Figure 2 Schematic flowchart of a method for implementing the generation of a technical theme tree provided in an embodiment of the present application; Figure 3 Schematic diagram of an initial weighted directed graph provided in an embodiment of the present application; Figure 4 Unmerged initial weighted directed graph provided in an embodiment of the present application; Figure 5 Schematic diagram of a target weighted directed graph provided in an embodiment of the present application; Figure 6 Another target weighted directed graph provided in an embodiment of the present application; Figure 7 Schematic flowchart of a process for determining an initial weighted directed graph of a target text sequence provided in an embodiment of the present application; Figure 8 Schematic flowchart of a process for constructing a technical theme tree provided in an embodiment of the present application; Figure 9 Schematic diagram of the structure of a technical theme tree provided in an embodiment of the present application; Figure 10 Schematic diagram of the structure of a device for generating a technical theme tree provided in an embodiment of the present application; Figure 11A schematic structural diagram of an electronic device corresponding to a method for generating a technical theme tree provided by an embodiment of the present application. Detailed implementation manners

[0021] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the implementation manners described below with reference to the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions of the embodiments of the present application.

[0022] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the terms "comprising" and "including" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude being implemented as other features, information, data, steps, operations, elements, components and / or combinations thereof supported by the art of the present technology. It should be understood that when we say an element is "connected" or "coupled" to another element, this element can be directly connected or coupled to the other element, or it can mean that this element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used here can include wireless connection or wireless coupling. The term "and / or" used here indicates at least one of the items defined by this term. For example, "A and / or B" or "A, B" indicates being implemented as "A", or being implemented as "B", or being implemented as "A and B". When describing multiple (two or more) items, if the relationship between the multiple items is not clearly defined, these multiple items can refer to one, multiple or all of the multiple items. For example, for the description of "parameter A includes A1, A2, A3", it can be implemented that parameter A includes A1 or A2 or A3, and it can also be implemented that parameter A includes at least two of the three items of parameter A1, A2, A3.

[0023] Structural extraction of text is easy for users to understand the text and facilitates subsequent execution of various services. There are still problems with the technical themes extracted by the related technology. These include inaccurate extracted technical themes, fuzzy hierarchical relationships between various technical themes, and many redundant structures.

[0024] The embodiments of the present application provide a method for generating a technical theme tree. The embodiments of the present application can automatically extract unoptimized technical themes through a large language model, and then generate a directed graph, which is easy to construct a technical theme tree subsequently. Through semantic similarity, node merging is performed to complete the optimization of technical themes and obtain more accurate technical themes. Redundant nodes are reduced to obtain a technical theme tree with clearer levels and stronger semantic logic.

[0025] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe the embodiments of this application in detail with reference to the accompanying drawings.

[0026] This method in the embodiment of this application can be implemented by a technical theme tree generation system. For example, Figure 1 FIG. 5 is a schematic diagram of a technical theme tree generation system provided in an embodiment of this application. The system includes a terminal device 101 and a server 102. The terminal device 101 can send a target text sequence to the server 102 through a network, so that the server 102 extracts the theme information of the target text sequence and constructs a technical theme tree. The server 102 is deployed with a large language model. The server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (Content Delivery Network, CDN), and big data and artificial intelligence platforms, but is not limited thereto. The terminal device 101 and the server 102 can be directly or indirectly connected through wired or wireless communication methods, and are not limited in the embodiment of this application.

[0027] The following will illustrate the technical solutions of the embodiments of this application and the technical effects produced by the technical solutions of this application through the description of several exemplary embodiments. It should be noted that the following embodiments can refer to, draw on, or combine with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.

[0028] Figure 2 FIG. 12 is a flowchart of a method for implementing the generation of a technical theme tree provided in an embodiment of this application. As Figure 2 shown, the method includes S201 to S206: S201: Obtain a target text sequence, where the target text sequence includes multiple target texts.

[0029] The target texts in the target text sequence are arranged based on a preset order. The target text can be any type of text for which technical themes need to be extracted, including academic literature, patent documents, user manuals, etc. Each target text in the target text sequence can belong to the same technical field. For example, each target text in the target text sequence is in the field of electronic design.

[0030] The server can receive the target text sequence sent by the terminal, or obtain the target text sequence through other means, which is not limited in the embodiment of this application.

[0031] S202: Extract the theme information of each target text respectively through a large language model. For each target text, the theme information of the target text includes at least one technical theme of the target text. When there are multiple technical themes of the target text, the theme information of the target text further includes the hierarchical relationship between the multiple technical themes.

[0032] Among them, the technical theme can represent the core content recorded in the target text, including the technical field, technical means, etc. of the target text. For example, the technical theme can include "electronic engineering / chip design / low-power design", indicating that the target text corresponding to this technical theme belongs to the field of electronic engineering technology, and specifically involves the realization of low-power design for chips.

[0033] The hierarchical relationship refers to the inclusion relationship between technical themes. The embodiments of the present application do not limit the format of the hierarchical relationship in the theme information. It can be directly recorded in the text, such as directly recording "theme a includes theme b", or characterized by characters, such as a / b, indicating that theme a includes theme b.

[0034] For example, the technical theme can include "electronic engineering / chip design / low-power design", indicating that electronic engineering includes chip design, and chip design includes low-power design. It can be understood that the technical field of the target text is electronic engineering. When performing chip design, the technical means of low-power design used is the low-power design method in the chip design field.

[0035] It should be noted that the method of low-power design can be used not only in chip design, but also in the design of other fields, such as software algorithm design. The difference in the upper-level fields of low-power design makes the specific methods of low-power design different. For example, in the field of chip design, low-power design can consider changing the architecture of the processor. In software algorithm design, low-power design can consider changing the data structure of the algorithm.

[0036] For each target text in the target text sequence, the server can input the target text into the large language model to obtain the theme information of the target text output by the large language model. The embodiments of the present application do not limit the specific type of the large language model.

[0037] The theme information of the target text can be automatically extracted through the large language model to improve the efficiency of constructing the technical theme tree of the target text sequence. In addition, the model parameter order of the large language model is relatively large, and at the same time, it has a relatively complex network structure, making the large language model itself have a powerful natural language analysis ability. Then, the theme information of each target text output by the large language model is more accurate, and the hierarchy of the technical theme tree of the target text sequence constructed based on the theme information of each target text is also more accurate and clear.

[0038] It should be noted that the technical theme tree of the target text sequence can represent the core content of the target text sequence. The nodes in the technical theme tree represent technical themes, and the edges connecting two nodes represent the hierarchical relationship between the technical themes corresponding to the two nodes. For example, node a is connected to node b, node a is the root node, node b is the leaf node, and the hierarchical relationship is that the technical theme represented by node a includes the technical theme represented by node b.

[0039] S203: Determine the initial weighted directed graph corresponding to the target text sequence according to the theme information of each target text, where each node in the initial weighted directed graph represents a technical theme, the edge between two nodes represents the hierarchical relationship between the technical themes corresponding to the two nodes, and the weight of an edge represents the occurrence frequency of the hierarchical relationship corresponding to the two nodes connected by the edge in the hierarchical relationships extracted from the multiple target texts.

[0040] In the initial weighted directed graph, a technical theme can be represented by a node, and the hierarchical relationship between the two technical themes corresponding to the two nodes connected by the directed edge can be represented by the directed edge. It can be understood that in the initial weighted directed graph, the hierarchical relationship can be represented by the parent-child relationship between nodes. For example, if the technical theme corresponding to node h includes the technical theme corresponding to node i, then node h is the parent node of node i, and in the initial weighted directed graph, node h points to node i.

[0041] For each hierarchical relationship, the hierarchical relationship may be extracted at least once from the multiple target texts, and the number of times the hierarchical relationship is extracted can also represent the importance of the hierarchical relationship in the target text sequence. Therefore, the occurrence frequency of the hierarchical relationship corresponding to the two nodes connected by the edge in the hierarchical relationships extracted from the multiple target texts can be represented by the weight of the edge.

[0042] Figure 3 The figure is a schematic diagram of an initial weighted directed graph provided by an embodiment of the present application, as Figure 3 shown. For example, the technical themes included in the theme information of each target text are a, b, c, d, e, where the hierarchical relationship is that a includes b and d, b includes c, c includes d, and d includes e. The occurrence frequencies of the hierarchical relationships between the two nodes in the hierarchical relationships extracted from each target text are 10, 6, 8, 7, and 4 respectively. According to the hierarchical relationships in the theme information, the parent-child relationships between the nodes can be determined. Specifically, a points to b, the weight of the edge is 10, a points to d, the weight of the edge is 6, b points to c, the weight of the edge is 8, c points to d, the weight of the edge is 7, and d points to e, the weight of the edge is 4.

[0043] Construct an initial technical topic structure of the target text sequence in a graphical format according to the theme information of each target text. The graphical technical topic is clearer and also helps to merge the technical topics in the initial technical topic structure later, reducing redundant technical topics. Constructing a technical topic tree based on a weighted directed graph also significantly improves the efficiency of constructing the technical topic tree.

[0044] S204: Determine the semantic similarity between the technical topics represented by each node in the initial weighted directed graph, determine the semantic similarity between the technical topics that meet the preset conditions as the target similarity, and determine each node corresponding to each target similarity as the node to be merged.

[0045] Since the initial weighted directed graph is directly obtained according to the theme information of each target text, and there may be highly similar content between each target text, then the technical topics represented by each node in the obtained initial weighted directed graph may also be highly similar. The representation of too many relatively similar technical topics in the initial weighted directed graph may lead to a blurred structure of the initial weighted directed graph, and further make the technical topic tree of the final obtained target text sequence blurred and the hierarchical relationship chaotic. Therefore, the server can merge each technical topic represented by each node in the initial weighted directed graph.

[0046] Specifically, the server can determine the semantic similarity between the technical topics represented by each node in the initial weighted directed graph, and then screen out the nodes to be merged. The server can extract the semantic features of each technical topic, and based on the semantic features of each technical topic, determine the semantic similarity between each technical topic through the cosine similarity. It is also possible to determine the semantic similarity between the technical topics represented by each node in the initial weighted directed graph by other methods, such as performing singular value decomposition (SVD, Singular Value Decomposition) on each technical topic in latent semantic analysis (LSA) to determine the potential semantic correlation between each technical topic, and then obtaining the semantic similarity between each technical topic. The embodiments of the present application do not limit this.

[0047] After that, the server can determine that the semantic similarities between the technical themes represented by each node meet the preset conditions as the target similarities, and determine each node corresponding to each target similarity as the node to be merged. Among them, the preset conditions include at least one of the following: the semantic similarity is not less than the preset similarity threshold; in the order of sorting the similarities from large to small, the top preset number of semantic similarities are sorted forward. For example, if the preset similarity threshold is 70%, then the semantic similarities with the semantic similarities between the technical themes represented by each node not less than 70% are determined as the target similarities to determine the nodes to be merged. The semantic similarities between the technical themes represented by each node can also be sorted to obtain the similarity ranking of each semantic similarity, and the top three semantic similarities in the similarity ranking are determined as the target similarities.

[0048] By using the semantic similarities between each technical theme to determine the nodes to be merged, multiple technical theme information with relatively similar semantics can be screened out. For multiple technical themes with relatively similar semantics, only one technical theme can be retained, and redundant technical themes can be removed to ensure that the structure of the technical theme tree of the target text sequence is clearer, more accurate, and has a reasonable hierarchy. S205: Merge the nodes to be merged, the edges associated with the nodes to be merged, and the weights of the edges in the initial weighted directed graph to obtain a target weighted directed graph.

[0049] The server can determine the nodes to be retained among the nodes to be merged, and merge the edges and the weights of the edges associated with the nodes to be deleted among the nodes to be merged into the nodes to be retained. Among them, there is no unique limitation on how to specifically determine the nodes to be retained among the nodes to be merged in the embodiments of the present application. For example, for each target similarity, any node in a node pair corresponding to the target similarity can be used as the node to be retained, and the other node in the node pair can be used as the node to be deleted. It is also possible to determine the node to be retained and the node to be deleted according to the number of connected edges of each node in the node pair. For example, the node with the largest number of connected edges is retained. If there are two nodes with the largest number of connected edges among at least two nodes, the sum of the weights of all the connected edges of these two nodes can also be considered, and the node with the larger sum of weights is retained.

[0050] As another alternative, for each target similarity, if the node pair corresponding to the target similarity does not have the same node as the node pair corresponding to other target similarities, the node pair corresponding to the target similarity is used as the set to be merged. If the node pair corresponding to the target similarity has the same node as the node pair corresponding to other target similarities, the nodes in the node pairs with the same node are used as the set to be merged; For each set to be merged, input the technical themes represented by the nodes in the set to be merged and the preset task prompt data into the theme optimization model to obtain the target node of the set to be merged. The task prompt data is used to prompt the theme optimization model to screen out the target node from the input nodes. For each non-target node in the set to be merged, update the edges and the weights of the edges between the other nodes except the nodes in the set to be merged and the non-target node to the edges and the weights of the edges between the other nodes and the target node, and delete the edges, the weights of the edges, and the non-target node between the target node and the non-target node in the set to be merged to obtain the target weighted directed graph.

[0051] Figure 4 This is the unmerged initial weighted directed graph provided by the embodiment of the present application, as Figure 4 shown. The nodes include a, b, c, d, and e. The weight of the edge connecting node a and node b is 8, the weight of the edge connecting node c and node b is 5, the weight of the edge connecting node c and node d is 3, and the weight of the edge connecting node e and node b is 2. Select a node pair with the highest semantic similarity, that is, nodes b and c, for the merging operation. Then, the target similarity is only the semantic similarity corresponding to nodes b and c, and there is no other target similarity. Then, there are no identical nodes between the node pair corresponding to the target similarity and the node pairs corresponding to other target similarities. Nodes b and c form a set to be merged. Input the technical themes represented by nodes b and c in the set to be merged and the preset task prompt data into the theme optimization model to obtain the target node, that is, node b. This target node is the node that needs to be retained among the nodes to be merged, and the other node of the node pair is the non-target node, that is, the node that needs to be deleted, that is, node c.

[0052] The technical themes represented by nodes b and c can also be combined respectively with the preset task prompt data to obtain the input data. For example, if the technical theme represented by node b is machine learning and the technical theme represented by node c is deep learning, Table 1 is the input data for inputting into the theme optimization model, as shown in Table 1: Table 1

[0053] In Table 1, the input data includes the task prompt data, that is, the task description, and also includes the technical themes represented by the embedded nodes b and c respectively, that is, machine learning and deep learning.

[0054] The theme optimization model can not only output the target node but also give reasons. Table 2 is the output result of the theme optimization model, as shown in Table 2: Table 2

[0055] According to the target theme optimized from the model output according to the theme, the nodes in the set to be merged corresponding to the target theme are determined as target nodes. The target theme retained in Table 2 is machine learning, and the node corresponding to machine learning is node b. Therefore, the target node is node b, and the non-target node is node c.

[0056] Figure 5 This is an initial weighted directed graph provided by an embodiment of the present application based on Figure 4 , and a schematic diagram of the target weighted directed graph obtained by node merging is as Figure 5 shown.

[0057] During merging, the edge connecting node d and node c can be changed to connect node d and node b, and the edge weight of the edge connecting node c and node d can be changed to the weight of the edge connecting node b and node d, resulting in the weight of the edge connecting node b and node d being 3. Node c and the edge connecting node c and node b and the edge weight are deleted.

[0058] It should be noted that if there are identical nodes among the node pairs corresponding to the target similarity and the node pairs corresponding to other target similarities, the nodes in each node pair with identical nodes are used as the set to be merged, and then the target nodes in the set to be merged are determined for subsequent merging. There are at least three nodes in this set to be merged. When determining the target nodes, the technical themes represented by all the nodes in the set to be merged and the preset task prompt data can be input into the theme optimization model, or the technical themes represented by the nodes in one node pair corresponding to each target similarity in the set to be merged can be input into the theme optimization model respectively.

[0059] For example, the node pair corresponding to target similarity 1 includes node b and node c, and the node pair corresponding to target similarity 2 includes node c and node d. Then, the server can obtain a set to be merged including node b, node c, and node d. The target similarities corresponding to the nodes in the set to be merged include target similarity 1 and target similarity 2. Therefore, the technical themes represented by the node pair corresponding to target similarity 1, that is, node b and node c, are input into the theme optimization model, and it is determined that the target node in the node pair corresponding to target similarity 1 is node c. The technical themes represented by the node pair corresponding to target similarity 2, that is, node d and node c, are input into the theme optimization model, and it is determined that the target node in the node pair corresponding to target similarity 2 is node d. Then, it is determined that the target node of the set to be merged is node d, and the non-target nodes are node b and node c. When performing node merging, node c and node b can be merged first to obtain an updated node c, and then the updated node d and node c can be merged.

[0060] Figure 6 This is another target weighted directed graph provided by an embodiment of the present application, as Figure 6 shown.

[0061] The target weighted directed graph includes nodes a, d, and e. Among them, node a includes node d, node d includes node e, the weight of the edge connecting node a and node d is 8, and the weight of the edge connecting node d and node e is 2.

[0062] The structure of the obtained target weighted directed graph after merging is more concise, and the complexity is reduced. When there are many similar technical topics in the target text sequence, constructing the technical topic tree of the target text sequence subsequently also requires a large amount of computing resources. By merging nodes, there is no need to add redundant nodes to the technical topic tree of the constructed target text sequence subsequently, reducing the computing resources for constructing the technical topic tree of the target text sequence.

[0063] S206: Construct the technical topic tree of the target text sequence according to the target weighted directed graph.

[0064] The server can construct the technical topic tree of the target text sequence according to the parent-child relationship between the nodes in the target weighted directed graph.

[0065] Specifically, the server can determine each leaf node in the target weighted directed graph through a depth-first search algorithm, and then determine each parent node of each leaf node according to the parent-child relationship between the nodes. The determined parent nodes of each leaf node are used as child nodes, and then each parent node of each child node is determined until there is no parent node for each child node, obtaining the link relationship corresponding to each node in the target weighted directed graph. According to the link relationship, the corresponding technical topic tree is constructed.

[0066] The tree structure form can more intuitively display the association relationship between the technical topics of each target text in the target text sequence. Therefore, presenting the technical topic tree of the target text sequence to the user can enable the user to understand each target text in the target text sequence more quickly. After the foregoing steps S202~S205, the levels of the obtained technical topic trees of the target text sequence are also clear and the logic is reasonable, further improving the convenience for the user to understand the target text sequence.

[0067] For step S201, the server can also perform text preprocessing on the obtained target text sequence. The text preprocessing includes integrity check and structuring. The integrity check includes verifying the integrity of each target text in the target text sequence to ensure that the content is complete and error-free. The structuring includes parsing unstructured document content, such as paragraph text and technical descriptions, into a structured format for subsequent processing link calls.

[0068] Regarding step S202, when extracting the theme information of each target text, the server inputs each target text into the large language model according to the order of the target texts in the target text sequence to obtain the theme information of each target text.

[0069] The server can also obtain the first prompt information of the target text sequence. Based on the first prompt information, through the large language model, the theme information of each target text is extracted, and the first prompt information of each target text can be the same. Among them, the first prompt information can include any one of task descriptions, step descriptions, precautions, input and output examples, and can be specifically set according to needs.

[0070] Through the first prompt information, the accuracy of the theme information output by the large language model can be improved, and the technical themes in each target text can be extracted more comprehensively, making the subsequent constructed technical theme tree of the target text sequence more reasonable. The first prompt information can also include prompt words in a specific field, etc., so that the large language model can combine specific domain knowledge to hierarchically extract the technical themes of the target text to meet the user's needs for extracting technical themes in a specific field.

[0071] The first prompt information can be the content in Table 3 below: Table 3

[0072] In the example of the prompt information shown in Table 3, the prompt information includes the tasks and task processing methods of the large language model (task description, step description, and precautions), and also includes the learning examples for the large language model to refer to and learn (the examples in Table 3). In the example of the prompt information in Table 3, taking the prompt words in the specific field of integrated circuits as an example, the large language model will extract the technical themes for the integrated circuit field, and can output up to 5 hierarchical relationships between each technical theme, and use semicolons to separate each hierarchical relationship.

[0073] Specifically, the text received by the large language model includes target text sequence 1 or target text sequence 2. The first prompt information and each target text in target text sequence 1 or target text sequence 2 can be input into the large language model to obtain the theme information of each target text. Of course, each target text in target text sequence 1 or target text sequence 2 can also be embedded into the first prompt information, and then the first prompt information embedded with each target text is input into the large language model. The embodiments of the present application do not limit this.

[0074] The server can also obtain the preset first prompt information, which is used to prompt the large language model to output the theme information of the target text; For the first target text in the target text sequence, based on the first prompt information and the target text, the theme information of the target text is extracted through a large language model; For each target text other than the first target text in the target text sequence, based on the extracted technical theme and the first prompt information, second prompt information is determined. The second prompt information is used to prompt the large language model to output the theme information of the target text with reference to the extracted technical theme. Based on the target text and the second prompt information, the theme information of the target text is extracted through the large language model. The extracted technical theme is the technical theme of each target text before the target text in the target text sequence.

[0075] It can be understood that for each target text other than the first target text in the target text sequence, the server can use the technical themes of each target text before the target text in the target text sequence as reference words, so that the large language model can refer to the extracted technical theme and output the technical theme of the target text.

[0076] The extracted technical theme included in the second prompt information enables the large language model to directly use the extracted technical theme without regenerating the technical theme, avoiding repeated generation. When there are more target texts, there may be more repeated technical themes. The existence of the second prompt information can improve the extraction efficiency of theme information. In addition, if the fields to which the target texts in the target text sequence belong are the same field, the extracted technical theme can be used as the prompt information of the large language model, which can further improve the accuracy of the technical theme output by the large language model.

[0077] The extracted technical theme in the second prompt information can be as shown in Table 4 below: Table 4

[0078] The technical theme in the reference word in Table 4 can be at least one extracted technical theme. Each target text can be embedded in the second prompt information, that is, the position of the target text shown in Table 4. Combining with the reference word, input data is obtained, and the input data is input into the large language model.

[0079] For step S203, when constructing the initial weighted directed graph corresponding to the target text sequence, after the server obtains the theme information of all target texts in the target text sequence, the initial weighted directed graph corresponding to the target text sequence is determined. Or, according to the theme information of the first target text, a first weighted directed graph can be constructed; For each target text except the first one, every time the topic information of a target text is extracted, the constructed weighted directed graph is updated according to the topic information of the target text until the weighted directed graph updated based on the topic information of the last target text is obtained, and this weighted directed graph is determined as the initial weighted directed graph corresponding to the target text sequence.

[0080] Figure 7 This is a schematic flowchart of a process for determining the initial weighted directed graph of a target text sequence provided by an embodiment of this application, as Figure 7 shown.

[0081] i is the order of the target text currently input to the large language model in the target text sequence. It can be understood that after the topic information of the first target text in the target text sequence is extracted, the first weighted directed graph in the target text sequence is constructed according to the topic information of the first target text in the target text sequence. Then, the topic information of the second target text in the target text sequence is obtained, and the first weighted directed graph is updated according to the topic information of the second target text until the weighted directed graph updated based on the topic information of the last target text is obtained as the initial weighted directed graph of the target text sequence.

[0082] It can be understood that after the topic information of each target text is generated each time, the hierarchical relationship between the technical topics of the currently generated topic information can be dynamically added to the initial weighted directed graph to ensure the continuous accumulation and optimization of the results.

[0083] When there are many target texts in the target text sequence, the number of technical topics in the corresponding topic information will also be large. Then, the number of nodes in the initial weighted directed graph corresponding to the target text sequence constructed will also be large, resulting in a relatively complex structure of the initial weighted directed graph corresponding to the target text sequence. Therefore, by retaining a preset number of technical topics, an initial weighted directed graph corresponding to the target text sequence with a relatively simple structure can be obtained.

[0084] Specifically, count the first number of the extracted technical topics; and determine the number of times each technical topic appears to obtain the counting result of each technical topic; If the first number is greater than the preset value, then according to the order of the counting results of the extracted technical topics from largest to smallest, delete the topic information corresponding to the technical topics ranked after the preset value; Based on the retained topic information of each target text, determine the initial weighted directed graph corresponding to the target text sequence.

[0085] For example, if the preset value is 50, then 50 technical topics are retained. If the first quantity is 101, then they are sorted in descending order according to the counting results of each technical topic, and the top 50 technical topics corresponding to the counting results are retained, and the remaining 51 technical topics and their hierarchical relationships are deleted, that is, the topic information corresponding to the technical topics sorted after 50 is deleted.

[0086] The number of times a technical topic appears in the topic information of each target text can characterize the importance of the technical topic to the target text sequence. The more times it appears, the more important the technical topic is and the more representative it is of the target text sequence. Then, a preset number of technical topics with a large number of occurrences can be retained to construct an initial weighted directed graph corresponding to the target text sequence.

[0087] It should be noted that after the server obtains the topic information of all target texts in the target text sequence, it can then count the first quantity of technical topics in the topic information of all target texts. It can also count the first quantity of technical topics in the topic information of the currently extracted target text each time it obtains the topic information of a target text.

[0088] For the latter case, when performing step S203, the counting result of the technical topic in the currently obtained topic information in the initial weighted directed graph can be updated. For example, the currently obtained topic information includes topic 1, which is represented by node 1 in the initial weighted directed graph, and the counting result of node 1 is 3. Then, the counting result of this node 1 is incremented by 1 to become 4, indicating that after obtaining the topic information of the current target text, the technical topic represented by this node appears 4 times in the topic information of each extracted target text.

[0089] In addition, for the latter case, the server can optimize the efficiency of the technical topic generation task by dynamically storing high-frequency technical topics. By recording the nodes with higher frequencies of occurrence during the generation process and preferentially retaining the content that contributes more to the task result, redundant calculations and interference from low-frequency topic information can be reduced. If the topic information in the second prompt message is all the extracted technical topics, then, through the foregoing method, the data volume of the second prompt message can be reduced, and the processing efficiency of the large language model can be improved. In addition, if the server updates the retained technical nodes each time it obtains the topic information of a target text, then by dynamically storing high-frequency technical topics, it can ensure that the retained technical topics always match the business requirements, thereby significantly improving the overall performance and response speed of large-scale text processing tasks.

[0090] For step S204, the server can input the technical topic represented by each node in the initial weighted directed graph into a semantic extraction model to obtain the semantic features of the technical topic represented by this node output by the semantic extraction model; Determine the semantic similarity between the technical theme represented by this node and the technical themes represented by other nodes in the initial weighted directed graph according to the semantic characteristics of the technical theme represented by this node.

[0091] Follow Figure 4 Example, suppose Figure 4 The technical themes represented by nodes a, b, c, d, and e in are artificial intelligence, machine learning, deep learning, neural network, and computer vision respectively. Input "artificial intelligence, machine learning, deep learning, neural network, computer vision" into the semantic extraction model respectively, and obtain the semantic characteristics of the technical themes represented by each node output by the semantic extraction model. As follows: The semantic characteristics (semantic vectors) of node a are [0.9, 0.7, 0.2], the semantic characteristics of node b are [0.8, 0.6, 0.3], the semantic characteristics of node c are [0.7, 0.6, 0.4], the semantic characteristics of node d are [0.6, 0.5, 0.5], and the semantic characteristics of node e are [0.4, 0.3, 0.8].

[0092] Calculate the cosine similarity between each pair of nodes, generate a similarity matrix and sort it, and take the top three pairs as examples: The semantic similarity between node b and node c is 0.95, the semantic similarity between node c and node d is 0.90, and the semantic similarity between node a and node b is 0.85.

[0093] Of course, the theme information of the technical themes represented by each node can also be input into the semantic extraction model. The theme information of the technical theme includes the technical theme name in text form and the hierarchical relationship between this technical theme and other technical themes. Following the above example, the technical theme names of each node are artificial intelligence, machine learning, deep learning, neural network, and computer vision respectively.

[0094] Extracting the semantic characteristics of the technical themes represented by each node through the semantic extraction model is efficient, and can extract the deep semantic characteristics of the technical themes, which helps to improve the accuracy of the nodes to be merged determined subsequently, and further improves the accuracy of the constructed technical theme tree.

[0095] For step S205, when merging nodes, the merging of the counting results of the nodes is also involved.

[0096] Specifically, in the initial weighted directed graph, add the counting results of each node in the nodes to be merged to obtain the merged counting result, and update the counting result of the target node to the merged counting result.

[0097] Follow Figure 4Example, suppose the counting result of node a is 15, the counting result of node b is 12, the counting result of node c is 10, the counting result of node d is 7, and the counting result of node e is 5. Suppose the target weighted directed graph after merging is Figure 7 , then in the target weighted directed graph, the counting result of node a is 15, the counting result of node d is 12 + 10 + 7 = 29, and the counting result of node e is 5.

[0098] For step S206, the server can delete the self-loop edges in the target weighted directed graph; For each node in the target weighted directed graph after deleting the self-loop edges, determine the importance of each incoming edge of the node according to the weights of the incoming edges of the node, so as to determine the parent-child relationship in the technical topic tree to be constructed according to the determined importance of each incoming edge; Construct the technical topic tree of the target text sequence according to the nodes in the target weighted directed graph and the importance of each incoming edge of each node.

[0099] When merging nodes, if two adjacent nodes represent the merging of two technical topics, self-loop edges may be generated. Here, adjacent means that when there is only one edge connecting two nodes between two nodes, then these two nodes are adjacent. Self-loop edges may also appear. For example, if the technical topic is "A / B / B", then in the structure of the target weighted directed graph, there will be an edge "B pointing to B". This kind of edge will cause hierarchical ambiguity between technical topics and affect the subsequent construction of the technical topic tree. Therefore, the server can first clean all self-loop edges in the graph to ensure that there is no repeated pointing relationship of the same technical topic.

[0100] After that, the server can determine the importance of each incoming edge to each node based on the weights of each incoming edge of each node in the target weighted directed graph after deleting the self-loop edges. The server can directly use the weights of each incoming edge of each node in the target weighted directed graph as the importance of each incoming edge to each node. As another alternative, the importance of a certain incoming edge of the node can also be determined according to the ratio of the weight of a certain incoming edge of the node to the sum of the weights of all incoming edges of the node.

[0101] Continue to use Figure 3 Example, for example, Figure 3 for node d in, node d has two incoming edges, namely the incoming edge from node a to node d, which is incoming edge 1, and the incoming edge from node c to node d, which is incoming edge 2. The importance of incoming edge 1 to node d is 6 / (6 + 7) 0.46, and the importance of incoming edge 2 to node d is 7 / (6 + 7) 0.54. The importance of the incoming edge of node b is 10 / 10 = 1, the importance of the incoming edge of node c is 8 / 8 = 1, the importance of the incoming edge of node e is 4 / 4 = 1, and node a has no incoming edge.

[0102] The importance of incoming edges to nodes can be used to determine the parent-child relationships in the subsequent technical topic tree, so that the connection relationships of each node in the constructed technical topic tree are of relatively high importance, and the constructed technical topic tree can better represent the target text sequence.

[0103] When constructing the technical topic tree of the target text sequence according to the importance of each incoming edge of each node and the target weighted directed graph, the server can initialize the stack structure; By continuously performing the target operation until all nodes in the target weighted directed graph are included in the constructed technical topic tree, the technical topic tree constructed by each target operation is determined as the technical topic tree of the target text sequence; Among them, the target operation includes: Determine the node with the smallest out-degree and / or the smallest sum of the weights of each outgoing edge in the current directed graph as the leaf node; among them, the current directed graph of the first target operation is the target weighted directed graph; Take the leaf node as the first node to be processed in this target operation, and continuously perform the following operations on the node to be processed until the stop stack-in condition is met: Push the node to be processed onto the stack; determine the incoming edge with the highest importance of the node to be processed, and determine the other node connected by the incoming edge except the node to be processed as the parent node of the node to be processed; If the stop stack-in condition is not met, take the parent node of the node to be processed as the next node to be processed in this target operation; If the stop stack-in condition is met, pop the nodes that have been pushed onto the stack, construct the technical topic tree corresponding to this target operation based on the popped nodes, delete the leaf node of this target operation from the current directed graph, and decrement the out-degree of the parent node of the leaf node by 1 to obtain the current directed graph of the next target operation.

[0104] Figure 8 It is a schematic flowchart of a process for constructing a technical topic tree provided by an embodiment of the present application, as Figure 8 shown.

[0105] Continue to use Figure 3 the example to Figure 8For illustration, the out-degrees of nodes a, b, c, d, and e are 2, 1, 1, 1, and 0 respectively. The node with the smallest out-degree is node e. Therefore, node e is determined as the leaf node. It should be noted that if the smallest out-degree corresponds to multiple nodes, the sum of the weights of all the incoming edges included in these nodes can be further compared, and the node with the smallest sum of the weights of the incoming edges is selected as the leaf node. After determining the leaf node of this target operation, the leaf node, i.e., node e, can be placed into the initialized stack structure as the bottom node of the stack. Subsequently, the incoming edge with the highest importance degree of node e is determined to be the incoming edge connecting node e and node d. It can be understood that the other node except node e connected by the incoming edge with the highest importance degree of node e is node d. Therefore, node d is determined as the parent node of node e.

[0106] The rule of selecting the parent node according to importance can ensure the gradual construction of the technical topic tree from the bottom up.

[0107] Among them, the stack-in stopping conditions include at least one of the following: The in-degree of the node to be processed is 0; The weight of the incoming edge of the node to be processed is less than the fusing threshold.

[0108] Among them, the fusing threshold can also be called the reference weight or benchmark weight. Optionally, the fusing threshold can be a preset threshold, such as a preset empirical value or experimental value. As another alternative, the fusing threshold can be determined based on the weights of the incoming edges of each node in the stack, specifically as follows:

[0109] Among them, n is the number of nodes in the stack, is the weight of the directed edge from the (i + 1)-th node to the i-th node in the stack, and this directed edge is the incoming edge of the i-th node.

[0110] As an example, the in-degree of node e is 1, and the weight of the incoming edge of node e is 4, which is greater than the current fusing threshold , not meeting the stack-in stopping condition. The parent node of node e, i.e., node d, can be used as the node to be processed in the next target operation. Push node d onto the stack, select the other node except node d connected by the incoming edge with a relatively high importance degree, i.e., node c, and determine node c as the parent node of node d. It still does not meet the stack-in stopping condition. Push node c onto the stack, determine the parent node of node c as node b. It still does not meet the stack-in stopping condition. Push node b onto the stack, determine the parent node of node b as node a. It still does not meet the stack-in stopping condition. Push node a onto the stack. At this time, node a has no parent node and its in-degree is 0, meeting the stack-in stopping condition. The nodes a, b, c, d, and e in the stack can be taken out in the order of last-in first-out, and then the technical topic tree corresponding to this target operation can be constructed.

[0111] It can be understood that if the in-degree of the node to be processed is not 0 and / or the weight of the incoming edge of the node to be processed is not less than the fusing threshold, the target operation can continue to be executed until the in-degree of the node to be processed is 0 and / or the weight of the incoming edge of the node to be processed is less than the fusing threshold, at which point the stack push stops, and a technical topic tree is constructed based on the nodes in the stack.

[0112] Figure 9 FIG. is a schematic structural diagram of the technical topic tree provided by an embodiment of the present application, as Figure 9 shown.

[0113] The top node of the stack, i.e., node a, is determined as the root of the tree, and the bottom node of the stack, i.e., node e, is the leaf node. A technical topic tree is constructed based on the parent-child relationship between the nodes.

[0114] All nodes in the target weighted directed graph are included in this technical topic tree, so the construction stops.

[0115] It should be noted that if after constructing a technical topic tree, the technical topic tree does not include all nodes in the target weighted directed graph, then in order to ensure that a new leaf node can be correctly searched each time the target operation is executed, before each execution of the target operation, the bottom node of the current stack, i.e., the leaf node, is removed from the current directed graph, and the out-degree of the parent node of this leaf node is decreased by 1 to obtain the current directed graph for the next target operation.

[0116] In addition, when the weight of the incoming edge of the node to be processed is less than the fusing threshold, the server can use the node adjacent to the last node pushed onto the stack in the stack as the root, and construct a technical topic tree based on the nodes pushed onto the stack before the last node pushed onto the stack; Use the last node pushed onto the stack as the root to construct a technical topic tree to obtain the technical topic tree corresponding to this target operation.

[0117] For example, the stack includes nodes a, b, and c, and node a is the top node of the stack. Therefore, node a can be used as one technical topic tree, and node b can be used as the root of another technical topic tree, with node c as the leaf node.

[0118] Since in the process of generating topic information by the large language model, some uncommon or accidental hallucinations may occur, resulting in a chaotic hierarchy between technical topics or a hierarchical relationship between irrelevant technical topics. At this time, the weight of the incoming edge of the node to be processed is less than the fusing threshold. It can also be understood that the importance of the hierarchical relationship between the node to be processed and other nodes in the current stack is relatively low, and the relevance is relatively low. Therefore, a technical topic tree for this node to be processed can be constructed separately to avoid forced combination and obtain a technical topic tree with relatively low accuracy.

[0119] For example, for "Electronic Engineering / Medical Field / Heart Imaging", the relevance between the medical field and electronic engineering is small, and it has only appeared once in the output of the large language model. Moreover, the medical field already has a sufficiently rich hierarchy of sub-levels and sub-nodes as a topic. Therefore, a rule is needed to break it. Suppose the weight of the edge from electronic engineering to the medical field is 1, and the weight of the edge from the medical field to heart imaging is 5, then the fusing threshold is . Fuse the edge connecting electronic engineering to the medical field. With electronic engineering as the root node, a technical topic tree is formed. The medical field is determined as the root node of another technical topic tree, and the leaf node is heart imaging.

[0120] It should be noted that when determining the most important incoming edge of the node to be processed and identifying the other node connected by this incoming edge (excluding the node to be processed) as the parent node of the node to be processed, if the stack does not contain the other node connected by this incoming edge (excluding the node to be processed), then this other node is determined as the parent node of the node to be processed.

[0121] If the stack contains the other node connected by this incoming edge (excluding the node to be processed), then the other node connected by the target incoming edge (excluding the node to be processed) is determined as the parent node of the node to be processed. Here, the target incoming edge is the incoming edge with the highest importance among the incoming edges where the other node connected by it is not in the stack, and the other incoming edges are the incoming edges of the node to be processed other than the most important incoming edge.

[0122] If the stack contains the parent node of the node to be processed, it indicates that there is a cyclic structure in the target weighted directed graph, such as a pointing to b, b pointing to c, and c pointing to a. A cyclic structure will cause chaos in the technical topic tree. Therefore, the server can identify other parent nodes not in the stack to avoid the occurrence of a cyclic structure.

[0123] For example, currently the node to be processed is node b, and the determined parent node is node a. The outgoing edge connecting node a and node b is the most important in terms of importance among all the incoming edges of node b. However, node a is already in the stack, and node b also has another incoming edge, and the node of this other incoming edge is node d, which is not in the stack and is the most important in terms of importance among the other incoming edges. Then node d is determined as the parent node of node b, and the target operation is continued.

[0124] By calculating the importance of incoming edges, dynamically generating a stack, and multiple technical topic trees, it is possible to extract technical topic trees with clear hierarchies and no cycles from the target weighted directed graph. Through noise cleaning, such as removing self-loop edges and cyclic chain structures, the high quality and consistency of the generated technical topic trees are ensured. During the process of constructing the technical topic tree, the bottom-up construction logic is followed, progressing layer by layer from leaf nodes to root nodes, and finally organizing all nodes into multiple independent technical topic trees.

[0125] It is understandable that the embodiments of the present application adopt a large language model and a semantic extraction model, realizing the full - process automation from unstructured target text to hierarchical structure, reducing manual participation, and greatly improving efficiency and accuracy. Through the importance of incoming edges and noise cleaning, the clarity of the hierarchical structure of the technical theme tree and the rationality of semantic logic are ensured. The final obtained technical theme tree of the target text sequence more accurately reflects the semantic association of technical themes, supporting deeper knowledge mining. Through extracting semantic features and calculating semantic similarity, technical themes with similar semantics can be more precisely merged, and redundant or noisy nodes can be removed. When processing a large amount of text, high - quality technical themes can still be generated, avoiding problems such as chaotic technical themes or unreasonable hierarchies. The method for determining the retained technical themes based on the counting results enables the embodiments of the present application to exhibit excellent performance when processing large - scale text, quickly process a large amount of data, and ensure that the generated technical theme tree is complete and acyclic.

[0126] The finally obtained technical theme tree can also be displayed to the user, providing the user with a more intuitive view of the technical theme structure and supporting various application scenarios such as technical intelligence analysis and technical innovation navigation.

[0127] The embodiments of the present application provide a device for generating a technical theme tree, as Figure 10 shown. The device 100 for generating a technical theme tree may include: a target text sequence acquisition module 1001, a theme information extraction module 1002, an initial weighted directed graph determination module 1003, a to - be - merged information determination module 1004, a merging module 1005, and a theme tree construction module 1006, where: The target text sequence acquisition module 1001 is used to acquire a target text sequence, and the target text sequence includes a plurality of target texts; The theme information extraction module 1002 is used to extract the theme information of each of the target texts through a large language model. For each target text, the theme information of the target text includes at least one technical theme of the target text. When there are multiple technical themes of the target text, the theme information of the target text further includes the hierarchical relationship between the multiple technical themes; The initial weighted directed graph determination module 1003 is used to determine an initial weighted directed graph corresponding to the target text sequence according to the theme information of each of the target texts. Each node in the initial weighted directed graph represents a technical theme, the edge between two nodes represents the hierarchical relationship between the technical themes corresponding to the two nodes, and the weight of an edge represents the occurrence frequency of the hierarchical relationship between the two nodes connected by the edge in the hierarchical relationships extracted from the multiple target texts; The module 1004 for determining information to be merged is used to determine the semantic similarity between the technical themes represented by the nodes in the initial weighted directed graph, determine the semantic similarity between the technical themes that meet the preset conditions as the target similarity, and determine the nodes corresponding to each target similarity as the nodes to be merged; A merging module 1005 is used to merge the nodes to be merged, the edges associated with the nodes to be merged, and the weights of the edges in the initial weighted directed graph to obtain a target weighted directed graph; A topic tree construction module 1006 is used to construct a technical topic tree of the target text sequence according to the target weighted directed graph; The preset condition includes at least one of the following: The semantic similarity is not less than the preset similarity threshold; The similarities are sorted in descending order, with the semantic similarities of the preset number at the top.

[0128] Optionally, the topic information extraction module 1002 may be used to obtain preset first prompt information, where the first prompt information is used to prompt the large language model to output topic information of the target text; For a first target text in the target text sequence, extracting topic information of the target text through a large language model based on the first prompt information and the target text; For each target text except the first target text in the target text sequence, second prompt information is determined based on the extracted technical theme and the first prompt information, the second prompt information is used to prompt the large language model to output the theme information of the target text with reference to the extracted technical theme, and the theme information of the target text is extracted through the large language model based on the target text and the second prompt information, the extracted technical theme is the technical theme of each target text before the target text in the target text sequence.

[0129] Optionally, the initial weighted directed graph determination module 1003 may be used to construct a first weighted directed graph according to the topic information of the first target text; For each target text except the first target text, each time the topic information of a target text is extracted, the constructed weighted directed graph is updated according to the topic information of the target text, until a weighted directed graph updated based on the topic information of the last target text is obtained, and the weighted directed graph is determined as the initial weighted directed graph corresponding to the target text sequence.

[0130] Optionally, the initial weighted directed graph determination module 1003 may be used to count the first quantity of the extracted technical topics, and determine the number of occurrences of each technical topic, so as to obtain the counting result of each technical topic; If the first quantity is greater than a preset value, then according to the order of the counting results of the extracted technical topics from large to small, delete the topic information corresponding to the technical topics ranked after the preset value; Based on the topic information of the remaining target texts, determine the initial weighted directed graph corresponding to the target text sequence.

[0131] Optionally, the topic tree construction module 1006 may be used to delete the self-loop edges in the target weighted directed graph; For each node in the target weighted directed graph after deleting the self-loop edges, determine the importance of each incoming edge of the node according to the weights of the incoming edges of the node, so as to determine the parent-child relationship in the technical topic tree to be constructed according to the determined importance of each incoming edge; Construct the technical topic tree of the target text sequence according to the nodes in the target weighted directed graph and the importance of each incoming edge of each node.

[0132] Optionally, the topic tree construction module 1006 may be used to initialize the stack structure; By continuously performing the target operation until all the nodes in the target weighted directed graph are included in the constructed technical topic tree, determine the technical topic tree of the target text sequence as the technical topic tree constructed by each target operation; Wherein, the target operation includes: Determine the node with the smallest out-degree and / or the smallest sum of the weights of each outgoing edge in the current directed graph as the leaf node; wherein, the current directed graph of the first target operation is the target weighted directed graph; Use the leaf node as the first node to be processed in this target operation, and continuously perform the following operations on the node to be processed until the stop pushing condition is met: Push the node to be processed onto the stack; determine the incoming edge with the highest importance of the node to be processed, and determine the other node connected by the incoming edge except the node to be processed as the parent node of the node to be processed; If the stop pushing condition is not met, use the parent node of the node to be processed as the next node to be processed in this target operation; If the stop pushing condition is met, pop the nodes that have been pushed onto the stack, construct the technical topic tree corresponding to this target operation based on the popped nodes, delete the leaf node of this target operation from the current directed graph, and decrement the out-degree of the parent node of the leaf node by 1 to obtain the current directed graph of the next target operation.

[0133] Optionally, the subject tree construction module 1006 may be configured to, if there is no other node connected by the incoming edge in the stack except the node to be processed, determine the other node as the parent node of the node to be processed; The subject tree construction module 1006 may be configured to, if there is another node connected by the incoming edge in the stack except the node to be processed, determine the other node connected by the target incoming edge except the node to be processed as the parent node of the node to be processed, where the target incoming edge is the incoming edge with the highest importance among the incoming edges where the other node connected by the incoming edge is not in the stack, and the other incoming edges are the incoming edges of the node to be processed except the incoming edge with the highest importance.

[0134] Optionally, the stack stop condition includes at least one of the following: The in-degree of the node to be processed is 0; The weight of the incoming edge of the node to be processed is less than the fusing threshold; When the weight of the incoming edge of the node to be processed is less than the fusing threshold, the subject tree construction module 1006 may be configured to use the node adjacent to the last node pushed onto the stack in the stack as the root, and construct a technical subject tree based on the nodes pushed onto the stack before the last node pushed onto the stack; Use the last node pushed onto the stack as the root to construct a technical subject tree, and obtain the technical subject tree corresponding to the target operation this time.

[0135] Optionally, the information to be merged determination module 1004 may be configured to input the technical subject represented by each node in the initial weighted directed graph into the semantic extraction model, and obtain the semantic features of the technical subject represented by the node output by the semantic extraction model; Determine the semantic similarity between the technical subject represented by the node and the technical subjects represented by other nodes in the initial weighted directed graph according to the semantic features of the technical subject represented by the node.

[0136] Optionally, the merging module 1005 may be configured to, for each target similarity, if the node pair corresponding to the target similarity does not have the same node as the node pairs corresponding to other target similarities, use the node pair corresponding to the target similarity as the set to be merged, and if the node pair corresponding to the target similarity has the same node as the node pairs corresponding to other target similarities, use the nodes in the node pairs with the same node as the set to be merged; For each set to be merged, input the technical subjects represented by the nodes in the set to be merged and the preset task prompt data into the subject optimization model, and obtain the target node of the set to be merged, where the task prompt data is used to prompt the subject optimization model to screen out the target node from the input nodes; For each non-target node in the set to be merged, update the edges and their weights between the other nodes and this non-target node (excluding the nodes in this set to be merged) to the edges and their weights between the other nodes and the target node, and delete the edges, their weights, and the non-target nodes between the target node and the non-target nodes in this set to be merged, so as to obtain a target weighted directed graph.

[0137] The device according to the embodiment of the present application can execute the method provided by the embodiment of the present application, and its implementation principle is similar and has corresponding technical effects. The actions performed by each module in the device according to the embodiments of the present application correspond to the steps in the methods according to the embodiments of the present application. For the detailed function descriptions of the modules of the device, reference can specifically be made to the descriptions in the corresponding methods shown above, and details will not be repeated here.

[0138] In the embodiment of the present application, an electronic device is provided, which includes a memory, a processor, and a computer program stored on the memory. The processor executes the above computer program to implement the steps of the method provided by any optional embodiment of the present application. Compared with the prior art, the following can be achieved: The present application obtains a target text sequence, which includes multiple target texts. Based on the large language model, the automatic extraction of the theme information of the target text is realized, and the efficiency of obtaining the theme information of the target text is improved. Based on the theme information, an initial weighted directed graph is constructed. The initial weighted directed graph includes the hierarchical relationship and weights between technical themes, which is conducive to determining the technical theme tree subsequently and improving the efficiency of constructing the technical theme tree. The semantic similarity of the technical themes represented by the nodes is determined, and node merging can be performed according to the semantic similarity to obtain a target weighted directed graph, reducing redundant nodes, making the extracted technical themes more accurate, and making the hierarchy between the technical themes obtained from the target weighted directed graph clearer and the semantic logic stronger. In an optional embodiment, an electronic device is provided, as Figure 11 shown, as Figure 11 shown in Figure 11 which, the electronic device 2000 mainly includes at least one processor 2001 ( Figure 11 one is shown in

[0139] Among them, the memory 2002 can be used to store the operating system, application programs, etc. The application programs can include computer programs that implement the methods shown in the embodiments of the present application when called by the processor 2001, and can also include programs for implementing other functions or services. The memory 2002 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and computer programs. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0140] The processor 2001 is connected to the memory 2002 through the bus 2005 and realizes corresponding functions by calling the application programs stored in the memory 2002. Among them, the processor 2001 can be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof, which can implement or execute various exemplary logic blocks, modules, and circuits described in connection with the disclosure of the present application. The processor 2001 can also be a combination that realizes computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0141] The electronic device 2000 can be connected to a network through a communication module 2003 (which can include but is not limited to components such as a network interface) to communicate with other devices (such as user terminals or servers, etc.) through the network and achieve data interaction, such as sending data to other devices or receiving data from other devices. Among them, the communication module 2003 can include a wired network interface and / or a wireless network interface, etc., that is, the communication module can include at least one of a wired communication module or a wireless communication module.

[0142] The electronic device 2000 can be connected to the required input / output devices through an input / output interface 2004, such as a keyboard, a display device, etc. The electronic device 2000 itself can have a display device and can also externally connect other display devices through the interface 2004. Optionally, a storage device, such as a hard disk, etc., can also be connected through this interface 2004 to store the data in the electronic device 2000 into the storage device, or read the data in the storage device, and the data in the storage device can also be stored in the memory 2002. It can be understood that the input / output interface 2004 can be a wired interface or a wireless interface. According to different actual application scenarios, the devices connected to the input / output interface 2004 can be components of the electronic device 2000 or external devices connected to the electronic device 2000 when needed.

[0143] The bus 2005 for connecting each component can include a path to transmit information between the above components. The bus 2005 can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. According to different functions, the bus 2005 can be divided into an address bus, a data bus, a control bus, etc.

[0144] Optionally, for the solution provided in the embodiments of the present application, the memory 2002 can be used to store a computer program for executing the solution of the present application, and the processor 2001 runs it. When the processor 2001 runs this computer program, it implements the actions of the method or device provided in the embodiments of the present application.

[0145] The embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can implement the steps and corresponding contents of the foregoing method embodiments.

[0146] The embodiments of the present application also provide a computer program product, including a computer program. When the computer program is executed by a processor, it can implement the steps and corresponding contents of the foregoing method embodiments.

[0147] It should be understood that although the flowchart of the embodiments of the present application indicates each operation step by an arrow, the execution order of these steps is not limited to the order indicated by the arrow. Unless there is a clear description in this article, in some implementation scenarios of the embodiments of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage among these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present application do not limit this.

[0148] The above are only optional implementation manners of some implementation scenarios of the present application. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the present application, adopting other similar implementation means based on the technical idea of the present application also belongs to the protection scope of the embodiments of the present application.

Claims

1. A method for generating a technical theme tree, characterized in that: include: Acquire a target text sequence, wherein the target text sequence includes a plurality of target texts; By using a large language model, extracting the subject information of each target text respectively, wherein, for each target text, the subject information of the target text includes at least one technical subject of the target text, and when the target text has multiple technical subjects, the subject information of the target text also includes the hierarchical relationship between the multiple technical subjects; Determine an initial weighted directed graph corresponding to the target text sequence according to the subject information of each of the target texts, wherein each node in the initial weighted directed graph represents a technical subject, an edge between two nodes represents a hierarchical relationship between the technical subjects corresponding to the two nodes, and the weight of an edge represents the occurrence frequency of the hierarchical relationship corresponding to the two nodes connected by the edge in the hierarchical relationship extracted from the multiple target texts; Determine the semantic similarity between the technical themes represented by each node in the initial weighted directed graph, determine the semantic similarity between the technical themes that meet the preset conditions as the target similarity, and determine each node corresponding to each target similarity as the node to be merged; Merging the nodes to be merged, the edges associated with the nodes to be merged, and the weights of the edges in the initial weighted directed graph to obtain a target weighted directed graph; Constructing a technical theme tree of the target text sequence according to the target weighted directed graph; The preset condition includes at least one of the following: The semantic similarity is not less than the preset similarity threshold; The similarities are sorted in descending order, with the semantic similarities of the preset number at the top.

2. The method according to claim 1, characterized in that The extracting the subject information of each target text respectively through the large language model includes: Obtaining preset first prompt information, where the first prompt information is used to prompt the large language model to output subject information of the target text; For a first target text in the target text sequence, extracting topic information of the target text through a large language model based on the first prompt information and the target text; For each target text except the first target text in the target text sequence, second prompt information is determined based on the extracted technical theme and the first prompt information, the second prompt information is used to prompt the large language model to output the theme information of the target text with reference to the extracted technical theme, and the theme information of the target text is extracted through the large language model based on the target text and the second prompt information, the extracted technical theme is the technical theme of each target text before the target text in the target text sequence.

3. The method according to claim 2, characterized in that Determining an initial weighted directed graph corresponding to the target text sequence according to the subject information of each target text includes: Constructing a first weighted directed graph according to the subject information of the first target text; For each target text except the first target text, each time the topic information of a target text is extracted, the constructed weighted directed graph is updated according to the topic information of the target text, until a weighted directed graph updated based on the topic information of the last target text is obtained, and the weighted directed graph is determined as the initial weighted directed graph corresponding to the target text sequence.

4. The method according to any one of claims 1 to 3, characterized in that Determining the initial weighted directed graph corresponding to the target text sequence according to the subject information of each target text includes: Counting a first number of extracted technical topics, and determining the number of occurrences of each technical topic, to obtain a counting result for each technical topic; If the first number is greater than a preset value, deleting the subject information corresponding to the technical subjects ranked after the preset value according to the counting results of the extracted technical subjects in descending order; Based on the retained topic information of each target text, an initial weighted directed graph corresponding to the target text sequence is determined.

5. The method according to claim 4, characterized in that The step of constructing a technical theme tree of the target text sequence according to the target weighted directed graph includes: Deleting the self-loop edges in the target weighted directed graph; For each node in the target weighted directed graph from which the self-loop edge is deleted, the importance of each incoming edge of the node is determined according to the weight of each incoming edge of the node, so as to determine the parent-child relationship in the technical subject tree to be constructed according to the determined importance of each incoming edge; According to the importance of each node and each incoming edge of each node in the target weighted directed graph, a technical topic tree of the target text sequence is constructed.

6. The method according to claim 5, characterized in that The step of constructing a technical topic tree of the target text sequence according to the importance of each node and each incoming edge of each node in the target weighted directed graph comprises: Initialize the stack structure; By continuously executing the target operation until the constructed technical theme tree includes all nodes in the target weighted directed graph, the technical theme tree constructed by each target operation is determined as the technical theme tree of the target text sequence; The target operation includes: Determine the node with the smallest out-degree and / or the smallest sum of weights of each out-edge in the current directed graph as a leaf node; wherein the current directed graph of the first target operation is the target weighted directed graph; The leaf node is used as the first node to be processed for the target operation, and the following operations are continuously performed on the node to be processed until the stop stacking condition is met: Push the node to be processed into a stack; determine the most important incoming edge of the node to be processed, and determine another node other than the node to be processed connected by the incoming edge as the parent node of the node to be processed; If the stop stacking condition is not met, the parent node of the node to be processed is used as the next node to be processed for the target operation; If the stop pushing condition is met, pop the pushed nodes, build the technical theme tree corresponding to the target operation based on the popped nodes, delete the leaf node of the target operation from the current directed graph, and reduce the out-degree of the parent node of the leaf node by 1 to obtain the current directed graph of the next target operation.

7. The method according to claim 6, characterized in that The step of determining the most important incoming edge of the node to be processed, and determining another node other than the node to be processed connected to the incoming edge as the parent node of the node to be processed, includes: If there is no other node in the stack connected to the incoming edge other than the node to be processed, determining the other node as the parent node of the node to be processed; The method further comprises: If there exists another node in the stack connected to the incoming edge other than the node to be processed, then the other node connected to the target incoming edge other than the node to be processed is determined as the parent node of the node to be processed, wherein the target incoming edge is another node connected to other incoming edges that is not the most important incoming edge among the incoming edges of the stack, and the other incoming edges are the incoming edges of the node to be processed other than the most important incoming edge.

8. The method according to claim 6, characterized in that The stop push condition includes at least one of the following: The in-degree of the node to be processed is 0; The weight of the incoming edge of the node to be processed is less than the fuse threshold; When the weight of the incoming edge of the node to be processed is less than the fuse threshold, each node that has been stacked is popped out of the stack, and a technical theme tree corresponding to the target operation is constructed based on each popped node, including: Taking the node adjacent to the last node pushed into the stack as the root, a technical theme tree is constructed based on the nodes pushed into the stack before the last node pushed into the stack; Take the last node pushed into the stack as the root, build a technical theme tree, and obtain the technical theme tree corresponding to this target operation.

9. The method according to claim 1, characterized in that: Determining the semantic similarity between the technical topics represented by each node in the initial weighted directed graph includes: For each node in the initial weighted directed graph, the technical subject represented by the node is input into a semantic extraction model to obtain a semantic feature of the technical subject represented by the node output by the semantic extraction model; According to the semantic features of the technical subject represented by the node, the semantic similarity between the technical subject represented by the node and the technical subject represented by other nodes in the initial weighted directed graph is determined.

10. The method according to claim 1, characterized in that The step of merging the nodes to be merged, the edges associated with the nodes to be merged, and the weights of the edges in the initial weighted directed graph to obtain a target weighted directed graph includes: For each target similarity, if the node pair corresponding to the target similarity does not have the same nodes as the node pairs corresponding to other target similarities, the node pair corresponding to the target similarity is used as the set to be merged; if the node pair corresponding to the target similarity has the same nodes as the node pairs corresponding to other target similarities, the nodes in each node pair with the same nodes are used as the set to be merged; For each set to be merged, the technical subject represented by each node in the set to be merged and the preset task prompt data are input into the subject optimization model to obtain the target node of the set to be merged, and the task prompt data is used to prompt the subject optimization model to select the target node from each input node; For each non-target node in the set to be merged, the edges and edge weights between other nodes except the node in the set to be merged and the non-target node are updated to the edges and edge weights between other nodes and the target node, and the edges, edge weights and non-target nodes between the target node and the non-target node in the set to be merged are deleted to obtain the target weighted directed graph.

11. A device for generating a technical theme tree, characterized in that: include: A target text sequence acquisition module, used to acquire a target text sequence, wherein the target text sequence includes a plurality of target texts; A topic information extraction module, used to extract topic information of each target text respectively through a large language model, wherein, for each target text, the topic information of the target text includes at least one technical topic of the target text, and when the target text has multiple technical topics, the topic information of the target text also includes a hierarchical relationship between the multiple technical topics; An initial weighted directed graph determination module is used to determine an initial weighted directed graph corresponding to the target text sequence according to the subject information of each of the target texts, wherein each node in the initial weighted directed graph represents a technical subject, an edge between two nodes represents a hierarchical relationship between the technical subjects corresponding to the two nodes, and the weight of an edge represents the occurrence frequency of the hierarchical relationship corresponding to the two nodes connected by the edge in the hierarchical relationship extracted from the multiple target texts; The module for determining information to be merged is used to determine the semantic similarity between the technical themes represented by each node in the initial weighted directed graph, determine the semantic similarity between the technical themes that meet the preset conditions as the target similarity, and determine each node corresponding to each target similarity as the node to be merged; A merging module, used for merging the nodes to be merged, the edges associated with the nodes to be merged, and the weights of the edges in the initial weighted directed graph to obtain a target weighted directed graph; A topic tree construction module, used to construct a technical topic tree of the target text sequence according to the target weighted directed graph; The preset condition includes at least one of the following: The semantic similarity is not less than the preset similarity threshold; The similarities are sorted in descending order, with the semantic similarities of the preset number at the top.

12. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 10.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.

14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.

Citation Information

Patent Citations

  • Patent literature similarity measurement method based on ontology

    CN107247780A

  • Thematic term determination method and device and storage medium

    CN114186557A

  • Information processing device, information processing method and program

    JP2018026039A

  • Mining sequential patterns in weighted directed graphs

    US20100251210A1

  • Generating descriptive topic labels

    US20170103074A1