Text association management method, device and computer program product
By constructing knowledge graphs and entity similarity calculations, and automatically identifying and recording association relationships, the problem of large workload of test text association sorting in agile testing environment is solved, and the understanding and understanding of test text associations is improved.
Patent Information
- Application Number
- CN202510351600.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-18
Smart Images

Figure CN120336541A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text association, and in particular, to a text association management method, device, computer-readable storage medium, and computer program product. Background Art
[0002] Before test execution, how to achieve in-depth and comprehensive coverage of requirement content has always been a challenge that testers have been constantly exploring. Although we have accumulated rich test experience, various test means, and theoretical methods, in today's reality where agile testing is becoming increasingly popular, it is still quite difficult to comprehensively obtain the associated information of iterative requirements. In the specific test execution process, the problems faced are not only the consumption of human and material resources and the dispersion of business points, but also involve deeper challenges, which can be specifically summarized as follows:
[0003] 1) Business points that change in real time: In an agile environment, it is normal for business points to change. It is necessary to cope with the business points that change in real time to ensure the timeliness and accuracy of test tasks.
[0004] 2) Version management of requirement documents: As requirements are iterated, there may be multiple versions of requirement documents. It is necessary to perform effective version management of requirement documents to ensure that test tasks are executed based on the latest business points.
[0005] 3) Complexity of business associations: In the agile iteration of a project, business associations may be very complex, and it is very common to span multiple modules and teams. If these complex business associations are not properly handled, it is difficult to ensure the identification and adaptation of global business change points.
[0006] 4) Knowledge sharing of business understanding: Due to the dispersion of business points and information isolation between teams, there may be differences in business understanding. It is necessary to achieve knowledge sharing of business understanding to avoid the problem of information silos. Summary of the Invention
[0007] The main purpose of the present application is to provide a text association management method, device, computer-readable storage medium, and computer program product to at least solve the problem of large workload in sorting out test text associations in the prior art.
[0008] To achieve the above object, according to one aspect of the present application, there is provided a text association management method, including: analyzing a target text to obtain a plurality of first entities and a first relationship between the plurality of first entities, where the first entities are named words with specific meanings; matching the first entities to obtain a comparison text including all the first entities; calculating the similarity between the first entities and the second entities according to the first relationship between the first entities and the second relationship between the second entities in the comparison text; when the similarity between a first target entity and a second target entity is greater than a predetermined threshold, determining that the first target entity and the second target entity are associated entities, and recording the first relationship corresponding to the first target entity and the second relationship corresponding to the second target entity, where the first target entity is any one of the first entities, and the second target entity is any one of the second entities; when receiving an entity query instruction, outputting the text where the entity corresponding to the entity query instruction is located, the first relationship between the entities corresponding to the entity query instruction, the text where the associated entity of the entity corresponding to the entity query instruction is located, and the second relationship between the associated entities of the entity corresponding to the entity query instruction.
[0009] Optionally, analyzing the target text to obtain a plurality of first entities and a first relationship between the plurality of first entities includes: preprocessing the target text to obtain a processed target text, where the preprocessing includes removing unrecognizable special characters, deleting URLs, and commenting out specific nouns; performing word segmentation on the processed target text to obtain a plurality of target words; identifying all the target words to obtain a plurality of first entities; matching the plurality of first entities to obtain the first relationship between the plurality of first entities, where the first relationship is one of predefined relationship categories.
[0010] Optionally, calculating the similarity between the first entities and the second entities according to the first relationship between the first entities and the second relationship between the second entities in the comparison text includes: constructing a knowledge graph of the target text according to the plurality of first entities and the first relationship between the plurality of first entities to obtain a first knowledge graph, where the first knowledge graph includes nodes corresponding to the first entities and edges corresponding to the first relationship, and the nodes corresponding to the first entities include first key nodes and first function nodes; obtaining a knowledge graph corresponding to the comparison text to obtain a second knowledge graph, where the second knowledge graph includes nodes corresponding to the second entities and edges corresponding to the second relationship, and the nodes corresponding to the second entities include second key nodes and second function nodes; calculating the similarity between the first key nodes of the first knowledge graph and the second key nodes of the second knowledge graph.
[0011] Optionally, calculating the similarity between the first key node and the second key node includes: establishing an adjacency matrix for each of the first key nodes according to the connection relationships between each of the first key nodes and other nodes, where the elements of the adjacency matrix are used to represent whether there is a connection relationship between the first key node and other nodes; extracting the feature vectors of all entities in the entity library to obtain a feature matrix; inputting the feature matrix and the adjacency matrix into a heterogeneous graph convolutional neural network to obtain the vector representation of the first key node; calculating the similarity between the first key node and the second key node based on the vector representation of the first key node and the vector representation of the second key node.
[0012] Optionally, extracting the feature vectors of all entities in the entity library to obtain a feature matrix includes: segmenting all the entities in the entity library to obtain a plurality of words; encoding the words using One-hot to obtain a feature matrix, where one word corresponds to a row vector of the feature matrix.
[0013] Optionally, inputting the feature matrix and the adjacency matrix into a heterogeneous graph convolutional neural network to obtain the vector representation of the first key node includes: calculating, according to the feature matrix, all the adjacency matrices, and the weight matrix W of the heterogeneous graph convolutional neural network pr to obtain the node embeddings of all nodes of the first knowledge graph According to all the node embeddings calculate the average importance of each edge of the first knowledge graph, and normalize the average importance to obtain an attention coefficient; calculate the node embedding of the first key node and the attention coefficient to calculate the vector representation of the first key node.
[0014] Optionally, calculating the similarity between the first key node and the second key node based on the vector representation of the first key node and the vector representation of the second key node includes: using to calculate the similarity score(αi,αj) between the first key node αi and the second key node αj, where is the vector representation of the first key node, is the vector representation of the second key node.
[0015] According to another aspect of the present application, there is provided a text association management device, including: an analysis unit configured to analyze a target text to obtain a plurality of first entities and a first relationship between the plurality of first entities, where the first entities are name words with specific meanings; a matching unit configured to match the first entities to obtain a comparison text including all the first entities; a calculation unit configured to calculate a similarity between the first entities and the second entities according to the first relationship between the first entities and a second relationship between the second entities of the comparison text; a determination unit configured to, when the similarity between a first target entity and a second target entity is greater than a predetermined threshold, determine that the first target entity and the second target entity are associated entities, and record the first relationship corresponding to the first target entity and the second relationship corresponding to the second target entity, where the first target entity is any one of the first entities and the second target entity is any one of the second entities; an output unit configured to, when receiving an entity query instruction, output the text where the entity corresponding to the entity query instruction is located, the first relationship between the entities corresponding to the entity query instruction, the text where the associated entity of the entity corresponding to the entity query instruction is located, and the second relationship between the associated entities of the entity corresponding to the entity query instruction.
[0016] According to still another aspect of the present application, there is provided a computer-readable storage medium, where the computer-readable storage medium includes a stored program, and when the program runs, it controls the device where the computer-readable storage medium is located to execute any one of the above methods.
[0017] According to yet another aspect of the present application, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, it implements any one of the above methods.
[0018] Applying the technical solution of the present application in the above text association management method, by parsing the first entities and the first relationship between the first entities of the target text, the similarity between the first entities and the second entities is determined according to the first relationship between the first entities and the second relationship between the second entities of the comparison text. If the similarity meets the requirements, the first entity and the second entity are associated entities, and the corresponding first relationship and second relationship are recorded. When querying the text, the text where the entity corresponding to the entity query instruction is located, the first relationship between the entities corresponding to the entity query instruction, the text where the associated entity of the entity corresponding to the entity query instruction is located, and the second relationship between the associated entities of the entity corresponding to the entity query instruction are output, improving the user's understanding and awareness of the relevance of the test text, eliminating the need for testers to search for associated texts according to test requirements, greatly reducing the workload of test text collation, and solving the problem of large workload of test text association collation in the prior art. Brief Description of the Drawings
[0019] Figure 1 It shows a hardware structure block diagram of a mobile terminal for implementing a text association management method provided in an embodiment of the present application;
[0020] Figure 2 It shows a schematic flowchart of a text association management method provided in an embodiment of the present application;
[0021] Figure 3 It shows a schematic flowchart of a knowledge graph construction provided in an embodiment of the present application;
[0022] Figure 4 It shows a detailed schematic flowchart of a knowledge graph construction provided in an embodiment of the present application;
[0023] Figure 5 It shows a schematic flowchart of implementing graph neural network similarity calculation provided in an embodiment of the present application;
[0024] Figure 6 It shows a structure block diagram of a text association management device provided in an embodiment of the present application.
[0025] Among them, the above-mentioned drawings include the following reference numerals:
[0026] 102, processor; 104, memory; 106, transmission device; 108, input / output device. Detailed Description of the Embodiment
[0027] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0028] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0029] It should be noted that the terms "first", "second", etc. in the description, claims, and above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so as to implement the embodiments of this application described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0030] For ease of description, some nouns or terms related to the embodiments of this application are described below:
[0031] Natural Language Processing (NLP): A branch in the field of artificial intelligence that involves the interaction between computers and human natural languages. Its goal is to enable computers to understand, interpret, generate, and communicate effectively with human languages. NLP is used to analyze documents and divide text into key nodes for subsequent construction of graph structures.
[0032] Attention Mechanism: In deep learning, a technology that simulates the human attention mechanism and is used to assign different weights to information at different positions when processing sequential data, enabling the model to pay more attention to important parts when processing sequences or sets.
[0033] Graph: A structure that defines nodes (points) and connection methods (edges). Nodes and edges each have their own properties. In addition to nodes being able to express information, edges can also express information. The edges of a graph can be weighted or unweighted. For example, chemical molecules (atoms / bonds), urban subways (platforms / railways), social networks (people / relationships), etc.
[0034] Knowledge Representation: Refers to converting the knowledge in the real world into a form that can be understood and processed by a computer. Formalizing diverse knowledge into a form that can be operated by a computer system facilitates reasoning, learning, and problem-solving by the computer system. In a knowledge graph, knowledge representation is carried out in the form of a graph. The graph representation represents entities, relationships, and attributes as nodes and edges in the graph, thus forming a graphical knowledge representation structure. Knowledge representation can help us more intuitively understand the knowledge in the knowledge graph and perform related query and analysis operations.
[0035] Graph Neural Networks (GNN): A class of deep learning models for processing graph data. Different from traditional neural network models that focus on processing structured data (such as vectors or matrices), GNNs are designed to process unstructured data, such as graphs or networks. The core idea is to update the representation of nodes by iteratively propagating and aggregating node information. By learning the relationships and feature representations between nodes, tasks in graph data can be solved. It can be used for training data to determine the association dependencies between graphs, so as to quickly identify the relevance between documents.
[0036] As introduced in the background art, in the prior art, the workload of testing text association collation is large. To solve this technical problem, embodiments of the present application provide a text association management method, device, computer-readable storage medium, and computer program product.
[0037] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0038] The method embodiments provided in the embodiments of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 is a hardware structure block diagram of a mobile terminal of a text association management method according to an embodiment of the present invention. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1 only one is shown in Figure 1 processors 102 (the processors 102 may include, but are not limited to, processing devices such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown in Figure 1 is only schematic and does not limit the structure of the above mobile terminal. For example, the mobile terminal may further include more or fewer components than
[0039] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the text association management method in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above-mentioned method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the mobile terminal through a network. Examples of the above-mentioned network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof. The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.
[0040] In this embodiment, a text association management method running on a mobile terminal, a computer terminal, or a similar computing device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0041] Figure 2 is a flowchart of the text association management method according to the embodiments of the present application. As Figure 2 shown, the method includes the following steps:
[0042] Step S201: Analyze the target text to obtain a plurality of first entities and a first relationship between the plurality of first entities, where the first entities are named words with specific meanings;
[0043] Step S202: Match the first entities to obtain a comparison text that includes all the first entities;
[0044] Step S203: Calculate the similarity between the first entities and the second entities according to the first relationship between the first entities and the second relationship between the second entities in the comparison text;
[0045] Step S204, in the case where the similarity between the first target entity and the second target entity is greater than a predetermined threshold, determine that the first target entity and the second target entity are associated entities, and record the first relationship corresponding to the first target entity and the second relationship corresponding to the second target entity. The first target entity is any one of the first entities, and the second target entity is any one of the second entities;
[0046] Step S205, in the case of receiving an entity query instruction, output the text where the entity corresponding to the entity query instruction is located, the first relationship between the entities corresponding to the entity query instruction, the text where the associated entity of the entity corresponding to the entity query instruction is located, and the second relationship between the associated entities of the entity corresponding to the entity query instruction.
[0047] In the above text association management method, by parsing the first entities and the first relationships between the first entities in the target text, the similarity between the first entity and the second entity is determined according to the first relationships between the first entities and the second relationships between the second entities in the comparison text. If the similarity meets the requirements, the first entity and the second entity are associated entities, and the corresponding first relationship and second relationship are recorded. When querying the text, output the text where the entity corresponding to the entity query instruction is located, the first relationship between the entities corresponding to the entity query instruction, the text where the associated entity of the entity corresponding to the entity query instruction is located, and the second relationship between the associated entities of the entity corresponding to the entity query instruction, improving the user's understanding and recognition of the relevance of the test text, eliminating the need for testers to search for associated texts according to test requirements, greatly reducing the workload of test text collation, and solving the problem of large workload in test text association collation in the prior art.
[0048] It should be noted that text association requires the use of a knowledge graph. As Figure 3 shown, it shows the specific flowchart of knowledge graph construction. First, apply NLP technology to preprocess the text document, then further perform entity recognition and relationship extraction, and finally perform knowledge representation on the output of the previous step and output structured knowledge graph data. The specific implementation steps are as follows:
[0049] Text preprocessing: This step requires unifying the text format, removing some unrecognizable special characters, etc., and deleting URLs and annotating specific nouns such as email addresses and phone numbers through data cleaning. After performing word segmentation, further annotate the part of speech of each word, such as nouns, verbs, adjectives, etc., which is beneficial to the accurate and efficient use of entity recognition and relationship extraction in the next step.
[0050] Entity Recognition and Relationship Extraction: Entity recognition aims to identify named entities from text, such as instance names like person names, place names, and organization names. Specific implementation can be achieved using rule-based, statistical learning-based, or pre-trained language model-based methods. Based on the results of entity recognition, relationship extraction matches and combines entities in the text, identifies the relationships between entities, and maps them to predefined relationship categories. After completing the correspondence between relationships and entities, the identified entities are further stored in the entity library, and the management graph corresponding to the file containing all entities of the target file is extracted as the comparison graph input for the graph neural network.
[0051] Knowledge Representation and Output of Knowledge Graph: Use entity recognition technology to identify entities from text and disambiguate entities with the same name but different meanings. Also, identify references such as pronouns and noun phrases in the text and determine the specific entities they refer to to resolve co-reference. According to the semantic relevance between entities and relationships, represent the knowledge in a graph-based structure, where entities are nodes in the graph and relationships are edges in the graph. Organize the constructed knowledge representation form into the structure of a knowledge graph and save it in a suitable data format. At this time, the output of the knowledge graph can be stored in the database in structured data form or viewed using a graphical visualization tool. Add the knowledge graph of the target file to the graph library and use it as the target graph input to the graph neural network.
[0052] In order to implement entity recognition and relationship extraction of the target text, in an alternative implementation manner, the above step S201 includes:
[0053] Step S2011, preprocess the above target text to obtain the processed target text. The above preprocessing includes removing unrecognizable special characters, deleting URLs, and commenting out specific nouns;
[0054] Step S2012, perform word segmentation on the above processed target text to obtain multiple target words;
[0055] Step S2013, identify all the above target words to obtain multiple above first entities;
[0056] Step S2014, match multiple above first entities to obtain the above first relationships between multiple above first entities. The above first relationships are one of the predefined relationship categories.
[0057] In the above implementation manner, as Figure 4As shown, in the text preprocessing stage, the steps of unifying text format, data cleaning, word segmentation, and part-of-speech tagging are sequentially performed. In the entity recognition and relationship extraction stage, entity recognition, relationship extraction, and attribute extraction are respectively performed. After entity recognition, the entities are added to the entity library. In the knowledge representation and knowledge graph output stage, entity disambiguation, coreference resolution, and graph database management of the knowledge graph (constructing the knowledge graph and adding it to the graph library) are respectively performed. By preprocessing the target text, performing entity recognition and relationship extraction, the above-mentioned first entity and the above-mentioned first relationship between the above-mentioned first entities can be obtained.
[0058] In order to analyze the relevance of the text, in an alternative implementation, the above step S203 includes:
[0059] Step S2031, constructing a knowledge graph of the target text based on multiple above-mentioned first entities and the above-mentioned first relationships between multiple above-mentioned first entities, to obtain a first knowledge graph. The first knowledge graph includes nodes corresponding to the first entities and edges corresponding to the first relationships. The nodes corresponding to the first entities include first key nodes and first functional nodes;
[0060] Step S2032, obtaining the knowledge graph corresponding to the comparison text, to obtain a second knowledge graph. The second knowledge graph includes nodes corresponding to the second entities and edges corresponding to the second relationships. The nodes corresponding to the second entities include second key nodes and second functional nodes;
[0061] Step S2033, calculating the similarity between the first key nodes of the first knowledge graph and the second key nodes of the second knowledge graph.
[0062] In the above implementation, entity disambiguation and coreference resolution are performed on the above-mentioned first entities. Then, based on multiple above-mentioned first entities and the above-mentioned first relationships between multiple above-mentioned first entities, a knowledge graph of the target text can be constructed to obtain a first knowledge graph, which is added to the graph library. And the knowledge graph corresponding to the comparison text is obtained from the graph library to obtain a second knowledge graph. The first knowledge graph and the second knowledge graph are subjected to associated graph analysis. The higher the similarity, the higher the node relevance.
[0063] In order to calculate the similarity between key nodes, in an alternative implementation, the above step S2033 includes:
[0064] Step S20331, establishing an adjacency matrix for each of the above-mentioned first key nodes according to the connection relationship between each of the above-mentioned first key nodes and other nodes. The elements of the adjacency matrix are used to represent whether there is a connection relationship between the first key node and other nodes;
[0065] Step S20332, extracting the feature vectors of all entities in the entity library to obtain a feature matrix;
[0066] Step S20333: Input the above feature matrix and the above adjacency matrix into a heterogeneous graph convolutional neural network to obtain the vector representation of the above first key node;
[0067] Step S20334: Calculate the similarity between the above first key node and the above second key node based on the vector representation of the above first key node and the vector representation of the above second key node.
[0068] In the above embodiment, the entities in the knowledge graph are classified into nodes of different importance: key nodes P and functional nodes O; a graph is represented as G = (V, E), where V = P ∪ O, P = {p1, p2... pn}, O = {o1, o2... om}, and n and m respectively represent the number of key nodes and functional nodes included; E represents the relationships between the graph nodes. Nodes and edges are respectively associated with a mapping function Ψ: V → A and Φ: E → R, where A and R respectively represent the sets of nodes and edges. Two nodes can be connected through different paths. For example, the path r from node V to Vl+1 is defined as: where represents a synthesis process that ignores the specific transformation of the relationship. For example, Figure 5 shows the specific operation steps of the function of this module. First, analyze the adjacency matrix and feature matrix of the nodes in each graph, and then further use the graph convolutional neural network to aggregate the information of each key node and then aggregate different connections through attention to obtain the final representation of each graph; finally, calculate the similarity between the key nodes, and conduct associated graph analysis based on the similarity and output the associated graph. Construct the adjacency matrix of the key node pi in each document knowledge graph as Ai. For example, if the key node pi is connected to the functional node oj, then the matrix value of pi and oj is 1; if pi and pk are connected to the same node oj, then the matrix value of pi and pk is 1. If a graph has n key nodes, n adjacency matrices will be generated. Specifically, the adjacency matrix A ij has the following calculation formula:
[0069] In order to extract the feature matrix, in an optional embodiment, the above step S20332 includes:
[0070] Step S203321: Segment all the above entities in the above entity library to obtain multiple words;
[0071] Step S203322: Encode the above words using One-hot to obtain a feature matrix, and one of the above words corresponds to a row vector of the above feature matrix.
[0072] In the above embodiments, each key node has a uniquely corresponding feature matrix, which is encoded using the One-hot method. First, the entities in the entire entity library are tokenized and counted, and then the feature vectors are extracted using One-hot to obtain the feature matrix X.
[0073] In order to implement the vector representation of nodes, in an alternative embodiment, the above step S20333 includes:
[0074] Step S203331, according to the above feature matrix, all the above adjacency matrices, and the weight matrix W of the above heterogeneous graph convolutional neural network pr Calculate the node embeddings of all nodes of the above first knowledge graph
[0075] Step S203332, according to all the above node embeddings Calculate the average importance of each edge of the above first knowledge graph, and normalize the above average importance to obtain the attention coefficient;
[0076] Step S203333, calculate the above node embedding of the above first key node And calculate the vector representation of the above first key node using the above attention coefficient.
[0077] In the above embodiments, GCN (Graph Convolutional Neural Network) can learn graph-structured data and continuously update parameters through convolution, improving the accuracy of the model and reducing the operation time. Because traditional GCN is used for homogeneous networks and only considers a single type of node, while the graph nodes we process are more complex and cannot be directly applied to traditional GCN. A heterogeneous graph convolutional neural network is adopted, considering the diversity of different types of nodes. The heterogeneous graph convolutional neural network fully considers different types of nodes in the graph and divides the original large graph into small graphs centered on key nodes. P = {P1, P2,..., PS} corresponds to the adjacency matrix A = {A(1), A(2),..., A(s)}. Among them, s represents the number of key nodes and the number of their corresponding adjacency matrices. To ensure the retention of node information, the adjacency matrix is processed as A' = A + I. Where I is the identity matrix. Taking the learning of vector representation by double-layer graph convolution as an example, taking the feature matrix X and all adjacency matrices A as the original input, the generated embedding representation is as follows: Among them, Is a sub-matrix of the adjacency matrix A', And Are the weight matrices of the first layer and the second layer of the double-layer graph convolution respectively. By averaging the importance of all nodes, the importance of each edge is obtained Among them, V is the sum of the number of key nodes and functional nodes, W is the weight of the node, and b is a constant. Finally, the softmax function is used for IP r Perform normalization to obtain the final attention coefficient β P r . At this time, the obtained final vector representation is
[0078] In order to calculate the similarity, in an alternative implementation, the above step S20334 includes:
[0079] Step S203341, using Calculate the similarity score score(αi,αj) between the above first key node αi and the above second key node αj, where is the vector representation of the above first key node, is the vector representation of the above second key node.
[0080] In the above implementation, given the vector representations of the key nodes in a target file, the cosine similarity is used to calculate the similarity score between nodes. When the score is closer to 1, it means the similarity between the two is higher. By sorting according to the score size, the nodes with high correlation can be analyzed. This step requires comparing all the key nodes in the target graph, and finally finding the nodes that meet the set threshold; if there are no nodes that meet the conditions, it means the correlation between the two graphs is low and does not meet the screening conditions. When there are key nodes with similar scores between two graphs, mark the similar nodes and their connection information in both graphs and use this as the output.
[0081] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0082] The embodiment of the present application also provides a text association management device. It should be noted that the text association management device of the embodiment of the present application can be used to execute the text association management method provided by the embodiment of the present application. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that realizes a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0083] The following introduces the text association management device provided by the embodiment of the present application.
[0084] Figure 6 is the structural block diagram of the text association management device according to the embodiment of the present application. AsFigure 6 As shown in the figure, the device includes:
[0085] An analysis unit 10, configured to analyze a target text to obtain a plurality of first entities and a first relationship between the plurality of first entities, where the first entities are named words with specific meanings;
[0086] A matching unit 20, configured to match the first entities to obtain a comparison text including all the first entities;
[0087] A calculation unit 30, configured to calculate a similarity between the first entities and second entities according to the first relationship between the first entities and a second relationship between second entities in the comparison text;
[0088] A determination unit 40, configured to determine that the first target entity and the second target entity are associated entities and record the first relationship corresponding to the first target entity and the second relationship corresponding to the second target entity when the similarity between the first target entity and the second target entity is greater than a predetermined threshold, where the first target entity is any one of the first entities and the second target entity is any one of the second entities;
[0089] An output unit 50, configured to output the text where the entity corresponding to the entity query instruction is located, the first relationship between the entities corresponding to the entity query instruction, the text where the associated entity of the entity corresponding to the entity query instruction is located, and the second relationship between the associated entities of the entity corresponding to the entity query instruction when receiving an entity query instruction.
[0090] In the above text association management device, by parsing the first entities and the first relationship between the first entities of the target text, the similarity between the first entities and the second entities is determined according to the first relationship between the first entities and the second relationship between the second entities in the comparison text. If the similarity meets the requirements, the first entities and the second entities are associated entities, and the corresponding first relationship and second relationship are recorded. When querying the text, the text where the entity corresponding to the entity query instruction is located, the first relationship between the entities corresponding to the entity query instruction, the text where the associated entity of the entity corresponding to the entity query instruction is located, and the second relationship between the associated entities of the entity corresponding to the entity query instruction are output, improving the user's understanding and recognition of the relevance of the test text, eliminating the need for testers to search for associated texts according to test requirements, greatly reducing the workload of test text collation, and solving the problem of large workload of test text association collation in the prior art.
[0091] It should be noted that text association requires the use of a knowledge graph, such as Figure 3As shown, it presents the specific flowchart of knowledge graph construction. First, NLP technology is applied to preprocess the text document, then entity recognition and relation extraction are further carried out, and finally, the output of the previous step is represented as knowledge and the structured knowledge graph data is output. The specific implementation steps are as follows:
[0092] Text preprocessing: In this step, it is necessary to unify the text format, remove some unrecognizable special characters, etc., and delete URLs and annotate specific nouns through data cleaning, such as email addresses and phone numbers. After performing word segmentation, further part-of-speech tagging is carried out for each word, such as nouns, verbs, adjectives, etc., which is conducive to the accurate and efficient use of entity recognition and relation extraction in the next step.
[0093] Entity recognition and relation extraction: Entity recognition aims to identify named entities from the text, such as instance names of person names, place names, organization names, etc. The specific implementation can be achieved using rule-based, statistical learning-based, or pre-trained language model-based methods. Based on the results of entity recognition, relation extraction matches and combines the entities in the text, identifies the relationships between entities, and maps them to predefined relation categories. After completing the correspondence between relations and entities, the recognized entities are further stored in the entity library, and the management graph corresponding to the file containing all entities of the target file is extracted as the comparison graph input for the graph neural network.
[0094] Knowledge representation and output of the knowledge graph: Use entity recognition technology to identify entities from the text, disambiguate entities with the same name but different meanings. And identify references such as pronouns and noun phrases in the text, and determine the specific entities they refer to to resolve co-reference. According to the semantic relevance between entities and relations, the knowledge is represented as a graph structure, where entities are nodes in the graph and relations are edges in the graph. Organize the constructed knowledge representation form into the structure of the knowledge graph and save it in a suitable data format. At this time, the output of the knowledge graph can be stored in the database in the form of structured data, or can be viewed using a graphical visualization tool. Add the knowledge graph of the target file to the graph library and input it into the graph neural network as the target graph.
[0095] To implement entity recognition and relation extraction of the target text, in an alternative implementation, the above analysis unit includes:
[0096] A preprocessing module for preprocessing the above target text to obtain the processed target text, where the preprocessing includes removing unrecognizable special characters, deleting URLs, and annotating specific nouns;
[0097] A word segmentation module for performing word segmentation on the above processed target text to obtain multiple target words;
[0098] An identification module, configured to identify all of the above target words to obtain a plurality of the above first entities;
[0099] A matching module, configured to match a plurality of the above first entities to obtain the above first relationships between the plurality of the above first entities, where the above first relationships are one of predefined relationship categories.
[0100] In the above embodiment, as Figure 4 shown, in the text preprocessing stage, uniform text format, data cleaning, word segmentation processing, and part-of-speech tagging are sequentially performed. In the entity recognition and relationship extraction stage, entity recognition, relationship extraction, and attribute extraction are respectively performed. After entity recognition, they are added to the entity library. In the knowledge representation and output knowledge graph stage, entity disambiguation, co-reference resolution, and graph database management of the graph (constructing the knowledge graph and adding it to the graph library) are respectively performed. By preprocessing the target text, entity recognition, and relationship extraction, the above first entities and the above first relationships between the above first entities can be obtained.
[0101] To analyze the relevance of the text, in an alternative embodiment, the above computing unit includes:
[0102] A construction subunit, configured to construct a knowledge graph of the above target text according to a plurality of the above first entities and the above first relationships between the plurality of the above first entities, to obtain a first knowledge graph, where the first knowledge graph includes nodes corresponding to the first entities and edges corresponding to the above first relationships, and the nodes corresponding to the first entities include first key nodes and first function nodes;
[0103] An acquisition subunit, configured to acquire the knowledge graph corresponding to the above comparison text to obtain a second knowledge graph, where the second knowledge graph includes nodes corresponding to the second entities and edges corresponding to the above second relationships, and the nodes corresponding to the second entities include second key nodes and second function nodes;
[0104] A calculation subunit, configured to calculate the similarity between the above first key nodes of the above first knowledge graph and the above second key nodes of the above second knowledge graph.
[0105] In the above embodiment, entity disambiguation and co-reference resolution are performed on the above first entities, and then a knowledge graph of the above target text can be constructed according to a plurality of the above first entities and the above first relationships between the plurality of the above first entities, to obtain a first knowledge graph, which is added to the graph library, and the knowledge graph corresponding to the above comparison text is acquired from the graph library to obtain a second knowledge graph. An association graph analysis is performed on the first knowledge graph and the second knowledge graph. The higher the similarity, the higher the node relevance.
[0106] To calculate the similarity between key nodes, in an alternative embodiment, the above calculation subunit includes:
[0107] A building module, configured to establish an adjacency matrix for each of the above-mentioned first key nodes according to the connection relationships between each of the above-mentioned first key nodes and other nodes, where the elements of the adjacency matrix are used to represent whether there is a connection relationship between the first key node and other nodes;
[0108] An extraction module, configured to extract the feature vectors of all entities in the entity library to obtain a feature matrix;
[0109] An input module, configured to input the feature matrix and the adjacency matrix into a heterogeneous graph convolutional neural network to obtain a vector representation of the first key node;
[0110] A calculation module, configured to calculate the similarity between the first key node and the second key node according to the vector representation of the first key node and the vector representation of the second key node.
[0111] In the above embodiment, the entities in the knowledge graph are classified into nodes of different importance: key nodes P and functional nodes O; a graph is represented as G=(V, E), where V = P ∪ O, P = {p1, p2... pn}, O = {o1, o2... om}, and n and m respectively represent the number of key nodes and functional nodes included; E represents the relationships between the graph nodes. Nodes and edges are respectively associated with a mapping function Ψ: V → A and Φ: E → R, where A and R respectively represent the sets of nodes and edges. Two nodes can be connected through different paths. For example, the path r from node V to Vl+1 is defined as: where represents a synthesis process that ignores the specific transformation of the relationship. For example, Figure 5 As shown, it shows the specific operation steps of the function of this module. First, analyze the adjacency matrix and the feature matrix of the nodes in each graph, and then further use the graph convolutional neural network to aggregate the information of each key node and then aggregate different connections through attention to obtain the final representation of each graph; finally, calculate the similarity between the key nodes, and perform associated graph analysis according to the similarity and output the associated graph. The adjacency matrix representing the key node pi in each document knowledge graph is denoted as Ai. For example, if the key node pi is connected to the functional node oj, then the matrix value of pi and oj is 1; if pi and pk are connected to the same node oj, then the matrix value of pi and pk is 1. If a graph has n key nodes, n adjacency matrices will be generated. Specifically, the adjacency matrix A ij The calculation formula is as follows:
[0112] In order to extract the feature matrix, in an optional embodiment, the above-mentioned extraction module includes:
[0113] A word segmentation sub-module for segmenting all the above entities in the above entity library to obtain a plurality of words;
[0114] An encoding sub-module for encoding the above words using One-hot to obtain a feature matrix, where one of the above words corresponds to a row vector of the above feature matrix.
[0115] In the above implementation manner, each key node has a uniquely corresponding feature matrix, which is encoded using the One-hot method. First, the entities in the entire entity library are segmented and counted, and then the feature vectors are extracted using One-hot to obtain the feature matrix X.
[0116] In order to implement the vector representation of nodes, in an optional implementation manner, the above input module includes:
[0117] A first calculation sub-module for calculating the node embeddings of all nodes of the above first knowledge graph according to the above feature matrix, all of the above adjacency matrices, and the weight matrix W of the above heterogeneous graph convolutional neural network pr The node embeddings of all nodes of the above first knowledge graph are calculated and obtained
[0118] A second calculation sub-module for calculating the average importance of each edge of the above first knowledge graph according to all of the above node embeddings and normalizing the above average importance to obtain an attention coefficient;
[0119] A third calculation sub-module for calculating the above node embedding of the above first key node and calculating the vector representation of the above first key node according to the above attention coefficient.
[0120] In the above implementation manner, GCN (Graph Convolutional Neural Network) can learn graph-structured data and continuously update parameters through convolution, improving the accuracy of the model and reducing the operation time. Because traditional GCN is used for homogeneous networks and only considers single-type nodes, while the graph nodes we process are more complex and cannot be directly applied to traditional GCN. The heterogeneous graph convolutional neural network is adopted, considering the diversity of different types of nodes. The heterogeneous graph convolutional neural network fully considers different types of nodes in the graph and divides the original large graph into small graphs centered on key nodes. P = {P1, P2,..., PS} corresponds to the adjacency matrix A = {A(1), A(2),..., A(s)}. Among them, s represents the number of key nodes and the number of their corresponding adjacency matrices. To ensure the retention of node information, the adjacency matrix is processed as A' = A + I. Where I is the identity matrix. Taking the learning of vector representation by double-layer graph convolution as an example, the feature matrix X and all adjacency matrices A are used as the original inputs, and the generated embedding representation is as follows: Among them, is a sub-matrix of the adjacency matrix A', and are the weight matrices of the first and second layers of the double-layer graph convolution respectively. By averaging the importance of all nodes, the importance of each edge is obtained where V is the sum of the number of key nodes and functional nodes, W is the weight of the nodes, and b is a constant. Finally, the softmax function is used to normalize I P r to obtain the final attention coefficient β P r . At this time, the obtained final vector representation is
[0121] In order to calculate the similarity, in an alternative implementation, the above calculation module includes:
[0122] A fourth calculation sub-module for using to calculate the similarity between the above first key node and the above second key node, where is the vector representation of the above first key node, is the vector representation of the above second key node.
[0123] In the above implementation, given the vector representation of the key nodes in a target file, the cosine similarity is used to calculate the similarity score between nodes. When the score is closer to 1, it means that the similarity between the two is higher. By sorting according to the score size, the nodes with high relevance can be analyzed. This step requires comparing all the key nodes in the target graph, and finally finding the nodes that meet the set threshold; if there are no nodes that meet the conditions, it means that the relevance between the two graphs is low and does not meet the screening conditions. When there are key nodes with similar scores between two graphs, the similar-related nodes and their connection information are marked in the two graphs and used as the output.
[0124] The above text association management device includes a processor and a memory. The above analysis unit, matching unit, calculation unit, determination unit, output unit, etc. are all stored in the memory as program units, and the processor executes the above program units stored in the memory to implement the corresponding functions. The above modules are all located in the same processor; or, the above modules are respectively located in different processors in any combination form.
[0125] The processor contains a kernel, and the kernel retrieves the corresponding program units from the memory. One or more kernels can be set, and by adjusting the kernel parameters, the problem of large workload in testing text association collation in the prior art can be solved.
[0126] The memory may include non - permanent memory in the form of computer - readable media, such as random access memory (RAM) and / or non - volatile memory, such as read - only memory (ROM) or flash RAM, and the memory includes at least one storage chip.
[0127] An embodiment of the present invention provides a computer - readable storage medium. The above - mentioned computer - readable storage medium includes a stored program. Wherein, when the above - mentioned program runs, it controls the device where the above - mentioned computer - readable storage medium is located to execute the above - mentioned text association management method.
[0128] Specifically, the text association management method includes:
[0129] Step S201: Analyze the target text to obtain a plurality of first entities and a first relationship between the plurality of first entities. The above - mentioned first entities are named words with specific meanings.
[0130] Step S202: Match the above - mentioned first entities to obtain a comparison text that includes all of the above - mentioned first entities.
[0131] Step S203: Calculate the similarity between the above - mentioned first entities and the second entities according to the above - mentioned first relationship between the first entities and the second relationship between the second entities in the above - mentioned comparison text.
[0132] Step S204: When the similarity between the first target entity and the second target entity is greater than a predetermined threshold, determine that the above - mentioned first target entity and the above - mentioned second target entity are associated entities, and record the above - mentioned first relationship corresponding to the above - mentioned first target entity and the above - mentioned second relationship corresponding to the above - mentioned second target entity. The above - mentioned first target entity is any one of the above - mentioned first entities, and the above - mentioned second target entity is any one of the above - mentioned second entities.
[0133] Step S205: When receiving an entity query instruction, output the text where the entity corresponding to the above - mentioned entity query instruction is located, the above - mentioned first relationship between the entities corresponding to the above - mentioned entity query instruction, the text where the associated entity of the entity corresponding to the above - mentioned entity query instruction is located, and the above - mentioned second relationship between the associated entities of the entity corresponding to the above - mentioned entity query instruction.
[0134] An embodiment of the present invention provides a processor. The above - mentioned processor is used to run a program. Wherein, when the above - mentioned program runs, it executes the above - mentioned text association management method.
[0135] Specifically, the text association management method includes:
[0136] Step S201: Analyze the target text to obtain a plurality of first entities and a first relationship between the plurality of first entities. The above - mentioned first entities are named words with specific meanings.
[0137] Step S202: Match the above first entities to obtain a comparison text containing all of the above first entities;
[0138] Step S203: Calculate the similarity between the above first entities and the above second entities according to the above first relationships between the above first entities and the second relationships between the second entities in the above comparison text;
[0139] Step S204: When the similarity between the first target entity and the second target entity is greater than a predetermined threshold, determine that the above first target entity and the above second target entity are associated entities, and record the above first relationship corresponding to the above first target entity and the above second relationship corresponding to the above second target entity, where the above first target entity is any one of the above first entities, and the above second target entity is any one of the above second entities;
[0140] Step S205: When receiving an entity query instruction, output the text where the entity corresponding to the above entity query instruction is located, the above first relationship between the entities corresponding to the above entity query instruction, the text where the above associated entity of the entity corresponding to the above entity query instruction is located, and the above second relationship between the above associated entities of the entity corresponding to the above entity query instruction.
[0141] An embodiment of the present invention provides a device, which includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, it implements at least the following steps:
[0142] Step S201: Analyze the target text to obtain a plurality of first entities and a first relationship between the plurality of above first entities, where the above first entities are name words with specific meanings;
[0143] Step S202: Match the above first entities to obtain a comparison text containing all of the above first entities;
[0144] Step S203: Calculate the similarity between the above first entities and the above second entities according to the above first relationships between the above first entities and the second relationships between the second entities in the above comparison text;
[0145] Step S204: When the similarity between the first target entity and the second target entity is greater than a predetermined threshold, determine that the above first target entity and the above second target entity are associated entities, and record the above first relationship corresponding to the above first target entity and the above second relationship corresponding to the above second target entity, where the above first target entity is any one of the above first entities, and the above second target entity is any one of the above second entities;
[0146] Step S205, when receiving an entity query instruction, output the text where the entity corresponding to the entity query instruction is located, the first relationship between the entities corresponding to the entity query instruction, the text where the associated entities of the entity corresponding to the entity query instruction are located, and the second relationship between the associated entities of the entity corresponding to the entity query instruction.
[0147] The present application also provides a computer program product, which, when executed on a data processing device, is adapted to execute a program initialized with at least the following method steps:
[0148] Step S201, analyze the target text to obtain a plurality of first entities and a first relationship between the plurality of first entities, where the first entities are name words with specific meanings;
[0149] Step S202, match the first entities to obtain a comparison text containing all the first entities;
[0150] Step S203, calculate the similarity between the first entities and the second entities according to the first relationship between the first entities and the second relationship between the second entities in the comparison text;
[0151] Step S204, when the similarity between the first target entity and the second target entity is greater than a predetermined threshold, determine that the first target entity and the second target entity are associated entities, and record the first relationship corresponding to the first target entity and the second relationship corresponding to the second target entity, where the first target entity is any one of the first entities and the second target entity is any one of the second entities;
[0152] Step S205, when receiving an entity query instruction, output the text where the entity corresponding to the entity query instruction is located, the first relationship between the entities corresponding to the entity query instruction, the text where the associated entities of the entity corresponding to the entity query instruction are located, and the second relationship between the associated entities of the entity corresponding to the entity query instruction.
[0153] Obviously, those skilled in the art should understand that the various modules or steps of the present invention described above can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed over a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present invention is not limited to any specific combination of hardware and software.
[0154] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0155] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks
[0156] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks
[0157] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing in the processFigure 1 one or more processes and / or blocks Figure 1 steps of functions specified in one block or more blocks.
[0158] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0159] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0160] Computer-readable media includes permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0161] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.
[0162] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:
[0163] 1) In the text association management method of the present application, by parsing the first entity in the target text and the first relationship between the first entities, the similarity between the first entity and the second entity is determined according to the first relationship between the first entities and the second relationship between the second entities in the comparison text. If the similarity meets the requirements, the first entity and the second entity are associated entities, and the corresponding first relationship and second relationship are recorded. When querying the text, the text where the entity corresponding to the entity query instruction is located, the first relationship between the entities corresponding to the entity query instruction, the text where the associated entity of the entity corresponding to the entity query instruction is located, and the second relationship between the associated entities of the entity corresponding to the entity query instruction are output, improving the understanding and awareness of the relevance of the test text by the user. There is no need for the tester to search for associated texts according to the test requirements, greatly reducing the workload of test text collation, and solving the problem of large workload of test text association collation in the prior art.
[0164] 2) In the text association management device of the present application, by parsing the first entity in the target text and the first relationship between the first entities, the similarity between the first entity and the second entity is determined according to the first relationship between the first entities and the second relationship between the second entities in the comparison text. If the similarity meets the requirements, the first entity and the second entity are associated entities, and the corresponding first relationship and second relationship are recorded. When querying the text, the text where the entity corresponding to the entity query instruction is located, the first relationship between the entities corresponding to the entity query instruction, the text where the associated entity of the entity corresponding to the entity query instruction is located, and the second relationship between the associated entities of the entity corresponding to the entity query instruction are output, improving the understanding and awareness of the relevance of the test text by the user. There is no need for the tester to search for associated texts according to the test requirements, greatly reducing the workload of test text collation, and solving the problem of large workload of test text association collation in the prior art.
[0165] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A text association management method, characterized in that, including: Analyze the target text to obtain multiple first entities and a first relationship between the multiple first entities, where the first entities are named words with specific meanings; Match the first entities to obtain a comparison text that includes all of the first entities; Calculate the similarity between the first entities and the second entities according to the first relationship between the first entities and the second relationship between the second entities in the comparison text; When the similarity between a first target entity and a second target entity is greater than a predetermined threshold, determine that the first target entity and the second target entity are associated entities, and record the first relationship corresponding to the first target entity and the second relationship corresponding to the second target entity, where the first target entity is any one of the first entities and the second target entity is any one of the second entities; When receiving an entity query instruction, output the text where the entity corresponding to the entity query instruction is located, the first relationship between the entities corresponding to the entity query instruction, the text where the associated entity of the entity corresponding to the entity query instruction is located, and the second relationship between the associated entities of the entity corresponding to the entity query instruction.
2. The method according to claim 1, characterized in that, Analyze the target text to obtain multiple first entities and a first relationship between the multiple first entities, including: Preprocess the target text to obtain a processed target text, where the preprocessing includes removing unrecognizable special characters, deleting URLs, and commenting out specific nouns; Perform word segmentation on the processed target text to obtain multiple target words; Identify all of the target words to obtain multiple first entities; Match the multiple first entities to obtain the first relationship between the multiple first entities, where the first relationship is one of the predefined relationship categories.
3. The method according to claim 1, characterized in that, Calculate the similarity between the first entities and the second entities according to the first relationship between the first entities and the second relationship between the second entities in the comparison text, including: Construct a knowledge graph of the target text according to the multiple first entities and the first relationship between the multiple first entities to obtain a first knowledge graph, where the first knowledge graph includes nodes corresponding to the first entities and edges corresponding to the first relationship, and the nodes corresponding to the first entities include first key nodes and first function nodes; Obtain the knowledge graph corresponding to the comparison text to obtain a second knowledge graph, where the second knowledge graph includes nodes corresponding to the second entities and edges corresponding to the second relationship, and the nodes corresponding to the second entities include second key nodes and second function nodes; Calculate the similarity between the first key nodes of the first knowledge graph and the second key nodes of the second knowledge graph.
4. The method according to claim 3, wherein Calculate the similarity between the first key nodes and the second key nodes, including: Establish an adjacency matrix for each of the first key nodes according to the connection relationship between each of the first key nodes and other nodes, where the elements of the adjacency matrix are used to represent whether there is a connection relationship between the first key node and other nodes; Extract the feature vectors of all entities in the entity library to obtain a feature matrix; Input the feature matrix and the adjacency matrix into a heterogeneous graph convolutional neural network to obtain the vector representation of the first key node; Calculate the similarity between the first key node and the second key node based on the vector representations of the first key node and the second key node.
5. The method according to claim 4, wherein Extract the feature vectors of all entities in the entity library to obtain a feature matrix, including: Segment all the entities in the entity library to obtain a plurality of words; Encode the words using One-hot to obtain a feature matrix, where one word corresponds to one row vector of the feature matrix.
6. The method according to claim 4, wherein Input the feature matrix and the adjacency matrix into a heterogeneous graph convolutional neural network to obtain the vector representation of the first key node, including: According to the feature matrix, all of the adjacency matrices, and the weight matrix W of the heterogeneous graph convolutional neural network pr the node embeddings of all nodes of the first knowledge graph are calculated According to all the node embeddings Calculate the average importance of each edge of the first knowledge graph, and normalize the average importance to obtain an attention coefficient; Calculate the node embedding of the first key node and the attention coefficient to calculate the vector representation of the first key node.
7. The method according to claim 4, characterized in that, Calculate the similarity between the first key node and the second key node based on the vector representations of the first key node and the second key node, including: Adopt Calculate the similarity score(αi,αj) between the first key node αi and the second key node αj, where is the vector representation of the first key node is the vector representation of the second key node 8. A text association management device, characterized in that, Including: An analysis unit for analyzing the target text to obtain a plurality of first entities and the first relationships between the plurality of first entities, where the first entities are name words with specific meanings; A matching unit for matching the first entities to obtain a comparison text containing all the first entities; A calculation unit for calculating the similarity between the first entities and the second entities according to the first relationships between the first entities and the second relationships between the second entities in the comparison text; A determination unit for determining that the first target entity and the second target entity are associated entities and recording the first relationship corresponding to the first target entity and the second relationship corresponding to the second target entity when the similarity between the first target entity and the second target entity is greater than a predetermined threshold, where the first target entity is any one of the first entities and the second target entity is any one of the second entities; An output unit for outputting the text where the entity corresponding to the entity query instruction is located, the first relationship between the entities corresponding to the entity query instruction, the text where the associated entity of the entity corresponding to the entity query instruction is located, and the second relationship between the associated entities of the entity corresponding to the entity query instruction when receiving an entity query instruction.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, where, when the program runs, it controls the device where the computer-readable storage medium is located to execute the method according to any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the method according to any one of claims 1 to 7.