A retrieval method and device based on a task relation graph (MRG)
Patent Information
- Application Number
- CN202610891433.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-18
- Publication Date
- 2026-09-29
AI Technical Summary
[0006]本发明旨在解决现有技术中存在的MRG检索方案仅依赖单一维度、缺乏针对性结构编码、且效率与精度难以兼顾的技术问题,从而提供一种能够同时兼顾语义相关性和结构可复用性,并能平衡检索效率与匹配精度的MRG检索方法及装置
1.通过生成语义向量和结构向量的双路编码方式,并对二者进行融合后检索,能够同时兼顾MRG的文本语义相关性和图结构可复用性,使得检索结果不仅在语义上相关,在结构上也更具兼容性,从而提高了检索结果的准确性和可用性。
Smart Images

Figure CN122838536A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a retrieval method and apparatus based on task relationship graph (MRG). Background Technology
[0002] In existing technologies, graph structure retrieval schemes mainly include semantic indexing based on graph coding, graph matching based on graph neural networks, and exact matching based on graph isomorphism or subgraph isomorphism. Currently, graph structure retrieval schemes mainly fall into the following categories: The first type is semantic indexing schemes based on graph coding. The core idea is to encode the semantic information of the graph into fixed-length feature vectors and then retrieve the results based on vector similarity. However, this type of scheme focuses on the semantic features of the graph, ignoring or only roughly processing the topological structure information of the graph, resulting in low reusability of the retrieval results at the structural level.
[0003] The second category is graph matching schemes based on graph neural networks, which embed graphs into vector spaces for matching. However, these schemes are usually designed for general graph structures and do not specifically encode the structural elements unique to MRGs (such as node type systems like task element nodes, task nodes, and operable nodes, as well as AND, OR, and sequential logic type systems), making it difficult to accurately reflect the functionality and logical compatibility of MRGs.
[0004] The third category is exact matching schemes based on graph isomorphism or subgraph isomorphism, which determine whether two graphs match through rigorous graph structure comparison. These schemes have high matching accuracy, but their computational complexity increases dramatically with the size of the graph (the VF2 algorithm has factorial complexity). When directly applied to large-scale MRG databases, their retrieval efficiency is extremely low, making them impractical for real-world engineering applications.
[0005] The fourth category is graph clustering-based solutions, which use methods such as closure trees or random walks to cluster graphs to accelerate retrieval. However, these solutions are mainly designed for general graph data and lack support for cross-type retrieval scenarios in MRGs. Summary of the Invention
[0006] The present invention aims to solve the technical problems of existing MRG retrieval schemes that rely on only a single dimension, lack targeted structural encoding, and are difficult to balance efficiency and accuracy. Therefore, it provides an MRG retrieval method and device that can simultaneously take into account semantic relevance and structural reusability, and balance retrieval efficiency and matching accuracy.
[0007] To achieve the above objectives, this invention provides a retrieval method based on Task Relationship Graph (MRG), comprising: obtaining a query MRG; generating a serialized text sequence according to a predetermined traversal order of the query MRG and the text elements contained therein, and encoding the serialized text sequence using a text embedding model to generate a semantic vector; generating a structure vector to represent the structural information of the query MRG, wherein the structural information includes: node type sequence features, edge type sequence features, node degree sequence features, and graph statistical features; performing a weighted summation of the semantic vector and the structure vector to generate a unified retrieval vector; performing a vector similarity retrieval in a database containing multiple target MRG representations based on the unified retrieval vector to obtain a candidate MRG set; determining the semantic similarity of each candidate MRG based on the semantic vector of the query MRG and the semantic vectors of each candidate MRG; and calculating the structural similarity between the candidate MRG and the query MRG for each candidate MRG in the candidate MRG set by performing graph isomorphism detection separately, or by combining subgraph isomorphism determination and node type alignment scoring. Based on the obtained semantic similarity and the calculated structural similarity, a final ranking score is obtained by weighted summation, and the candidate MRG set is finally ranked according to the final ranking score to obtain the final retrieval results.
[0008] Optionally, the step of obtaining the MRG to be queried includes: receiving a query request in natural language form, and parsing the query request into the MRG to be queried using a preset large language model.
[0009] Optionally, the calculation of the structural similarity between the candidate MRG and the query MRG includes: if graph isomorphism is true, the structural similarity score is directly set to the maximum score of 1.0; otherwise, the structural similarity is calculated based on the following formula: structural similarity = w1 × graph isomorphism detection score + w2 × subgraph isomorphism score + w3 × node type alignment score; where w1, w2, and w3 are configurable weights.
[0010] Optionally, the step of generating structural information to characterize the MRG to be queried includes: converting the type of each node into a numerical code according to a predetermined traversal order of nodes in the MRG to generate the node type sequence feature; converting the logical type of each edge into a numerical code according to a predetermined order of edges in the MRG to generate the edge type sequence feature; extracting the in-degree and out-degree of each node according to a predetermined traversal order of nodes in the MRG to generate the node degree sequence feature; and calculating the depth, average branch factor, total number of nodes, total number of edges, and proportion of each type of node and edge in the MRG to generate the graph statistical features.
[0011] Furthermore, the step of generating the structure vector includes: normalizing the node type sequence features, the edge type sequence features, the node degree sequence features, and the graph statistical features to the [0,1] interval respectively, and then concatenating them into a unified structure vector, and aligning the dimensions to the same dimension as the semantic vector through a linear projection layer, so that the two vectors can be weighted and fused in a unified vector space.
[0012] Optionally, the step of weighted summation of the semantic vector and the structural vector is implemented by the following formula: V_retrieval = α V_sem + β V_struc, where V_retrieval is the unified retrieval vector, V_sem is the semantic vector, V_struc is the structural vector, and α and β are preset weight coefficients.
[0013] Furthermore, the weighting coefficients α and β are dynamically adjusted according to a preset retrieval scenario.
[0014] Optionally, the vector similarity retrieval is an approximate nearest neighbor (ANN) retrieval.
[0015] Optionally, the step of obtaining the final ranking score through weighted summation is implemented by the following formula: Score_final = w_sem Score_sem + w_struc Score_struc, where Score_final is the final ranking score, Score_sem is the semantic similarity, Score_struc is the structural similarity, and w_sem and w_struc are preset weight coefficients.
[0016] This invention also provides a retrieval device based on a Task Relationship Graph (MRG), comprising: a modeling unit for modeling business entities as a Task Relationship Graph (MRG) structure to obtain a query MRG; a text processing unit for generating a serialized text sequence according to a predetermined traversal order of the query MRG and the text elements contained therein; a semantic encoding unit for encoding the serialized text sequence using a text embedding model to generate a semantic vector; a structural feature extraction unit for generating node type sequence features, edge type sequence features, node degree sequence features, and graph statistical features according to the query MRG; and a vector fusion unit for generating a structural vector based on the node type sequence features, edge type sequence features, node degree sequence features, and graph statistical features, and processing the semantic vector and the structural vector. A weighted summation is performed to generate a unified retrieval vector; a coarse screening unit is used to perform vector similarity retrieval in a database containing multiple target MRG representations based on the unified retrieval vector to obtain a candidate MRG set; and, based on the semantic vector of the MRG to be queried and the semantic vectors of each candidate MRG, to determine the semantic similarity corresponding to each candidate MRG; a fine ranking unit is used to calculate the structural similarity between each candidate MRG in the candidate MRG set and the MRG to be queried by performing graph isomorphism detection individually, or by combining subgraph isomorphism judgment and node type alignment scoring; and based on the obtained semantic similarity and the calculated structural similarity, a final ranking score is obtained by weighted summation, and the candidate MRG set is finally ranked according to the final ranking score to obtain the final retrieval result.
[0017] Compared with the prior art, the present invention has the following beneficial effects: 1. By generating semantic vectors and structural vectors through a dual-path encoding method and then fusing them for retrieval, the textual semantic relevance and graph structure reusability of MRG can be taken into account simultaneously. This makes the retrieval results not only semantically relevant but also structurally more compatible, thereby improving the accuracy and usability of the retrieval results.
[0018] 2. By extracting the unique node type sequence, edge type sequence, node degree sequence, and graph statistical features of MRG to generate structural vectors, targeted encoding of MRG structural features is achieved. Compared with general graph encoding methods, it can more accurately reflect the structural semantics of MRG and improve the accuracy of structural similarity assessment.
[0019] 3. A two-stage retrieval architecture of "vector coarse screening + structural fine ranking" is adopted. The first stage quickly narrows down the candidate range through efficient vector retrieval, while the second stage performs high-computational-complexity precise structural matching on only a small number of candidates. This greatly improves the retrieval efficiency in large-scale MRG databases while ensuring matching accuracy.
[0020] 4. The unified retrieval framework can adapt to different query construction methods and target libraries, supporting various cross-type retrieval scenarios such as "finding workstations based on needs" and "finding similar workstations based on workstations", which reduces the complexity of system implementation and maintenance and enhances the versatility and scalability of the solution. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart illustrating a retrieval method based on a Task Relationship Graph (MRG) according to an embodiment of the present invention.
[0023] Figure 2 This is a schematic block diagram of a retrieval device based on a task relationship graph (MRG) according to an embodiment of the present invention.
[0024] Figure 3 This is a schematic diagram of a structure vector generation process according to an embodiment of the present invention.
[0025] Figure 4 This is a schematic diagram of an example of a Task Relationship Graph (MRG) according to an embodiment of the present invention.
[0026] The reference numerals in the attached figures are explained as follows: 10-Modeling Unit, 20-Text Processing Unit, 30-Semantic Encoding Unit, 40-Structural Feature Extraction Unit, 50-Vector Fusion Unit, 60-Coarse Screening Unit, 70-Fine Ranking Unit, 80-Database; S101-Obtaining the MRG to be Queryed Step, S102-Generating Semantic Vector Step, S103-Generating Structural Vector Step, S104-Generating a Unified Retrieval Vector Step, S105-Vector Similarity Retrieval Step, S106-Calculating Semantic Similarity Step, S107-Calculating Structural Similarity Step, S108-Final Ranking Step; 301-Node Type Sequence Features, 302-Edge Type Sequence Features, 303-Node Degree Sequence Features, 304-Graph Statistical Features, 305-Concatenation Operation, 306-Linear Projection Layer, 307-Structural Vector; 401-Task Meta Node (MNS), 402-Task Node (MN), 403-Task Node (MN), 404-Task Derivation Condition Node (ME), 405-Operable Node (OMN), 406-Constraint Node (CONST), 407-Constraint Node (CONST), 408-Task Node (MN). Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be considered as limitations on this application. All other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this application. Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by those skilled in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.
[0028] Before providing a further detailed description of the embodiments of this application, some of the nouns and terms involved in the embodiments of this application will be explained. The nouns and terms involved in the embodiments of this application are subject to the following interpretations.
[0029] (1) Task Relationship Graph (MRG): This refers to a formal representation of tasks, which can use a directed acyclic graph (DAG) structure to record the logical relationships and operational dependencies of each task, while defining the task objectives of the corresponding agent. For example... Figure 4 As shown, an MRG instance can include multiple types of nodes, such as task meta node (MNS) 401, task node (MN) 402, 403, 408, task derivation condition node (ME) 404, operable node (OMN) 405, and constraint node (CONST) 406, 407. These nodes are connected by edges of different logical types and together describe a complete task.
[0030] (2) Semantic Vector: This refers to the vector representation used to capture the textual semantic information of the MRG. In a specific implementation, it can be obtained by encoding a semantic document composed of the text descriptions of all nodes and the label text of the edges in the MRG through a text embedding model. This vector mainly focuses on "what the MRG says", that is, the content and intent of the task.
[0031] (3) Structure Vector: This refers to the vector used to represent the structural information of the MRG to be queried. This vector mainly focuses on "what the MRG looks like," that is, its topological shape and organization. For example... Figure 3 As shown, the structure vector can be formed by fusing structural information from multiple dimensions, such as node type sequence features 301, edge type sequence features 302, node degree sequence features 303, and graph statistical features 304.
[0032] (4) Graph isomorphism detection: This refers to an operation that determines whether two graphs have completely identical graph structures without ignoring the specific text labels of the nodes. For example, algorithms well-known to those skilled in the art, such as the VF2 algorithm, can be used to implement this. If two MRG graphs are isomorphic, it means that they are completely consistent in terms of task flow, dependency relationships, and other structural aspects, and have high reusability.
[0033] (5) Approximate Nearest Neighbor (ANN) retrieval: This refers to an efficient vector retrieval technique used to quickly find several vectors most similar to the query vector in a large vector database, rather than performing a precise but time-consuming and computationally expensive full comparison. This technique is particularly suitable for scenarios that require rapid retrieval from massive amounts of data.
[0034] (6) Large Language Model (LLM): refers to a pre-trained deep learning model with natural language understanding and generation capabilities. In the embodiments of this application, it can be used to automatically parse and convert unstructured natural language query requests input by users into structured query MRGs.
[0035] Please see Figure 1 This application provides a retrieval method based on a Task Relationship Graph (MRG), aiming to solve the problem in existing technologies that cannot simultaneously consider semantic relevance and structural reusability when relying solely on semantics or structure for retrieval. It improves matching accuracy while maintaining retrieval efficiency through a dual-encoding, two-stage retrieval architecture. The method first obtains the MRG to be queried in step S101. This MRG can be an existing MRG directly provided by the user, or it can be generated through other means.
[0036] After obtaining the MRG to be queried, this method employs a dual-path parallel encoding strategy. On one hand, in step S102, the method focuses on the semantic content of the MRG. Specifically, based on the predetermined traversal order (e.g., topological sorting) of the MRG to be queried and the text elements it contains (such as node names, descriptions, etc.), a serialized text sequence is generated. Subsequently, a text embedding model is used to encode this text sequence to generate a semantic vector. The purpose of this step is to capture the core semantics such as the task intent and business logic described by the MRG, thus overcoming the deficiency of traditional structure matching methods that ignore text content. By converting the text information of the graph into a vector, semantic-based similarity calculation becomes possible.
[0037] On the other hand, in step S103, the method focuses on the structural morphology of the MRG. Specifically, a structure vector is generated to represent the structural information of the MRG to be queried. This structural information is multi-dimensional and may include node type sequence features, edge type sequence features, node degree sequence features, and graph statistical features, etc. The purpose of this design is that semantic similarity alone is not enough; the reusability of the task largely depends on the compatibility of its process structure. By extracting these MRG-specific structural features and vectorizing them, this method can quantitatively evaluate the structural similarity of different MRGs, making up for the shortcomings of traditional text retrieval that completely ignores structural information.
[0038] In step S104, to utilize both types of information simultaneously in subsequent retrievals, the semantic vector and the structural vector need to be fused. This method performs a weighted summation of the semantic vector and the structural vector to generate a unified retrieval vector. This fusion strategy ensures that the final representation vector contains information from both semantic and structural dimensions, laying the foundation for subsequent joint retrieval. By adjusting the weights, the importance of semantics and structure in retrieval can be flexibly controlled, thus adapting to different application scenarios.
[0039] After generating a unified retrieval vector, the method enters a two-stage retrieval process. The first stage is coarse screening, as shown in step S105. Based on the unified retrieval vector, vector similarity is searched in a database containing multiple target MRG representations to quickly obtain a set of candidate MRGs. This stage aims to efficiently filter out a small subset of candidate sets that are most likely to match from massive amounts of data, ensuring the system's response speed. Simultaneously, to provide a basis for the fine-tuning in the second stage, the semantic similarity of each candidate MRG is determined based on the semantic vector of the MRG to be queried and the semantic vectors of each candidate MRG.
[0040] The second stage is fine-tuning. For each candidate MRG in the coarsely screened candidate MRG set, in step S107, its structural similarity to the query MRG is evaluated through more precise calculations. This can be achieved in various ways, such as performing graph isomorphism detection alone, or combining subgraph isomorphism judgment and node type alignment scoring. The purpose of this step is to perform precise structural verification of the coarsely screened results, avoiding misjudgments caused by the approximation of vector representations, thereby significantly improving the accuracy of the final result.
[0041] Finally, in step S108, a final ranking score is obtained by weighted summation based on the semantic similarity obtained in the coarse screening stage (the result of step S106) and the structural similarity calculated in the fine ranking stage (the result of step S107). The candidate MRG set is then ranked according to this final ranking score to obtain the final retrieval results. In this way, the top-ranked results will be those MRGs that are both relevant to the query in terms of task content and compatible with the query structure in terms of execution flow, thus achieving an effective balance between semantic relevance and structural reusability.
[0042] In a preferred embodiment, to lower the barrier to entry for users, the process of obtaining the MRG to be queried in step S101 can be made more intelligent. Specifically, it can receive query requests entered by users in natural language, such as "help me find a workstation where I can do sales data analysis," and then use a preset Large Language Model (LLM) to automatically parse this ambiguous natural language request into a structured MRG to be queried that can be processed in subsequent steps. In this way, even if users are not familiar with the professional format of MRGs, they can easily initiate a search, improving the usability and ease of use of the system.
[0043] In another preferred embodiment, to make the structural similarity calculated in step S107 more refined and accurate, a case-by-case calculation strategy can be adopted. For example, graph isomorphism detection is performed first. If the graph structure of the query MRG is completely isomorphic to that of the candidate MRG, this represents a high degree of structural similarity, and the structural similarity score can be directly set to the maximum score of 1.0 without further complex calculations. If graph isomorphism is not established, a weighted formula can be used to comprehensively evaluate the degree of structural similarity. This formula can be: Structural similarity = w1 × Graph isomorphism detection score + w2 × Subgraph isomorphism score + w3 × Node type alignment score Here, w1, w2, and w3 are configurable weights (e.g., default values of 0.4, 0.3, and 0.3 respectively). The graph isomorphism detection score is determined as follows: if graph isomorphism is valid, the score is a maximum of 1.0. The subgraph isomorphism score reflects whether the candidate MRG can be reused to complete some or all of the functions of the query MRG. The node type alignment score reflects the structural similarity of two MRGs at the abstract level. Through this combined scoring mechanism, even if two MRGs are not completely isomorphic, a reasonable and quantifiable structural similarity score can be obtained, making the ranking more reasonable.
[0044] Furthermore, in order to ensure that the structure vector generated in step S103 can fully reflect the structural characteristics of the MRG, its generation process may include more specific steps. For example... Figure 3As shown, features from multiple dimensions can be extracted first. For example, based on the predetermined traversal order of nodes in the MRG to be queried, the types of each node (such as MNS, MN, OMN, etc.) are converted into numerical codes to generate node type sequence features 301; based on the predetermined order of edges in the MRG to be queried, the logical types of each edge (such as AND, OR, SEQ, etc.) are converted into numerical codes to generate edge type sequence features 302; based on the predetermined traversal order of nodes, the in-degree and out-degree of each node are extracted to generate node degree sequence features 303; and the depth, average branch factor, total number of nodes, total number of edges, and proportion of each type of node and edge in the MRG to be queried are calculated to generate graph statistical features 304. These features characterize the structure of the MRG from different perspectives, and are more targeted and interpretable than general graph coding methods.
[0045] In an optional implementation, after generating the aforementioned multiple structural feature components, to form a unified structural vector 307, the node type sequence feature 301, the edge type sequence feature 302, the node degree sequence feature 303, and the graph statistical feature 304 can be normalized (e.g., normalized to the [0,1] interval), and then merged into a higher-dimensional vector through a concatenation operation 305. To enable effective fusion with the semantic vector, a linear projection layer 306 can be used for dimension alignment, making its dimension the same as that of the semantic vector. This ensures that the structural information is fully and reasonably represented before entering the fusion step, providing a foundation for generating a unified retrieval vector.
[0046] Furthermore, the weighted summation of the semantic vector and the structural vector in step S104 can be implemented using an explicit formula, for example: V_retrieval = α V_sem + β V_struc Where V_retrieval is the generated unified retrieval vector, V_sem is the semantic vector, V_struc is the structural vector, and α and β are configurable preset weight coefficients (for example, the default can be set to α=β=0.5). This formula clearly defines the method of dual-path information fusion.
[0047] In one optional implementation, the weighting coefficients α and β are not fixed but can be dynamically adjusted according to a preset retrieval scenario. For example, when the user's intent is "to find similar workstations for best practice migration," structural similarity is more important than semantic similarity. In this case, the value of β can be dynamically increased (e.g., set to 0.7), while the value of α can be decreased (e.g., set to 0.3). Conversely, when the user's intent is "to find possible functions based on a vague requirement description," semantic matching is more critical. In this case, the value of α can be increased, while the value of β can be decreased. This dynamic weighting mechanism allows the retrieval behavior to better adapt to the user's specific intent, thereby helping to provide retrieval results that better match the user's intent.
[0048] In a preferred embodiment, the vector similarity retrieval in step S105 can specifically employ the Approximate Nearest Neighbor (ANN) retrieval algorithm. ANN retrieval can significantly reduce the time complexity of retrieval (typically logarithmic) while maintaining recall, thus ensuring that the system can still efficiently complete the retrieval even when the MRG database reaches hundreds of thousands or even millions in size, improving the retrieval speed in large-scale databases. This is one of the key technologies for achieving an efficient coarse screening stage.
[0049] Furthermore, the process of obtaining the final ranking score through weighted summation in step S108 can also be implemented using an explicit formula, for example: Score_final = w_sem Score_sem + w_struc Score_struc Wherein, Score_final is the final ranking score, Score_sem is the determined semantic similarity, Score_struc is the calculated structural similarity, and w_sem and w_struc are preset weight coefficients. This formula defines how to combine the similarity scores of both semantic and structural dimensions to perform the final ranking of the candidate set during the fine-ranking stage, thereby helping to select a more comprehensive result that better meets the user's needs.
[0050] Please see Figure 2 This application also provides a retrieval device based on a Task Relationship Graph (MRG), which is a specific physical or logical implementation of the above-described method. The device includes: a modeling unit 10, a text processing unit 20, a semantic encoding unit 30, a structural feature extraction unit 40, a vector fusion unit 50, a coarse screening unit 60, and a fine ranking unit 70. These units work together to complete the entire process from receiving a query to outputting the final retrieval result.
[0051] Specifically, modeling unit 10 is used to model business entities as a task relationship graph (MRG) structure to obtain the MRG to be queried. For example, it can receive natural language input from the user and call LLM to convert it into an MRG, corresponding to step S101 in the method. Text processing unit 20 is used to generate a serialized text sequence according to the predetermined traversal order of the MRG to be queried and the text elements contained therein. Semantic encoding unit 30 is used to encode the text sequence using a text embedding model to generate a semantic vector. Text processing unit 20 and semantic encoding unit 30 jointly implement step S102 in the method.
[0052] Simultaneously, the structural feature extraction unit 40 generates node type sequence features, edge type sequence features, node degree sequence features, and graph statistical features based on the MRG to be queried. This corresponds to the part of the method that generates structural information. The vector fusion unit 50 is responsible for generating structural vectors based on these structural features and performing a weighted summation of the semantic vector output by the semantic encoding unit 30 and its own generated structural vector to generate a unified retrieval vector. The work of the vector fusion unit 50 corresponds to steps S103 and S104 in the method.
[0053] The coarse screening unit 60 performs vector similarity retrieval in the database 80 based on the unified retrieval vector generated by the vector fusion unit 50 to obtain a set of candidate MRGs and determine the semantic similarity of each candidate MRG. This corresponds to the coarse screening stage of the method, namely steps S105 and S106. The fine ranking unit 70 is responsible for refining the candidate MRG set output by the coarse screening unit 60. It calculates the structural similarity between the candidate MRGs and the query MRGs through graph isomorphism detection and other methods, and calculates the final ranking score based on semantic similarity and structural similarity to finally rank the candidate set and obtain the final retrieval results. The work of the fine ranking unit 70 corresponds to the fine ranking stage of the method, namely steps S107 and S108.
[0054] The technical solution of the present invention will be described in detail below in conjunction with specific application scenarios. In a "requirement-finding workstation" search scenario, suppose a product manager needs to design a "member points calculation workstation" for the membership system and enters a natural language requirement: "I need a workstation that can calculate the points that a user should receive based on the user's spending amount, spending frequency, and membership level." First, modeling unit 10 receives the requirement and calls the large language model to parse it into a temporary requirement MRG. The structure of this MRG can be a task meta node (MNS) "Member Points Calculation", which is connected to three parallel task nodes (MN) "Get Consumption Amount", "Get Consumption Frequency", and "Get Membership Level". These three task nodes then converge into an operable node (OMN) "Calculate Points". This process completes step S101.
[0055] Next, the system uses the temporary MRG as the query and performs dual-path encoding. The text processing unit 20 and the semantic encoding unit 30 extract text such as "consumption amount", "membership level", and "points calculation" from the MRG to generate a semantic vector (step S102). At the same time, the structural feature extraction unit 40 analyzes its graph structure, encodes its node type distribution of "1 MNS + 3 MN + 1 OMN" and the corresponding edge type distribution, and generates a structural vector (step S103). The vector fusion unit 50 merges these two vectors into a unified retrieval vector (step S104).
[0056] Subsequently, the coarse screening unit 60 uses the unified retrieval vector to perform an approximate nearest neighbor (ANN) search in a database 80 containing, for example, 2000 existing workstation MRGs, quickly recalling 50 candidate workstations (step S105). This stage takes a short time, for example, about 15 milliseconds.
[0057] In the fine-ranking stage, fine-ranking unit 70 performs precise structural similarity calculations on each of the 50 candidate workstations (step S107). For example, the system finds a candidate MRG named "VIP Integral Calculation Workstation". The graph isomorphism detection using the VF2 algorithm fails because this VIP workstation has an additional node, "Get VIP Level", compared to the query MRG. However, subgraph isomorphism detection using the Ullmann algorithm reveals that the query MRG is a subgraph of this VIP workstation MRG, resulting in a subgraph isomorphism score of 0.92. Simultaneously, the node type alignment score is calculated to be 0.88. The overall structural similarity is then calculated according to a preset formula. Structural similarity = w1 × graph isomorphism detection score + w2 × subgraph isomorphism score + w3 × node type alignment score = 0.4 × 0 + 0.3 × 0.92 + 0.3 × 0.88 = 0.54.
[0058] Meanwhile, the semantic similarity calculated based on vector cosine similarity is 0.88 (step S106). Finally, the final ranking score is calculated according to the formula (step S108), taking an example where the weights of w_sem and w_struc are both 0.5: Final ranking score = w_sem × semantic similarity + w_struc × structural similarity = 0.5 × 0.88 + 0.5 × 0.54 = 0.71.
[0059] The system sorts all candidates based on this final ranking score and recommends this "VIP Points Calculation Workstation" to the product manager. The product manager finds it to be a good match for their needs and can reuse it with minor adjustments, thus helping to reduce the workload of creating workstations from scratch.
[0060] In another application scenario, the method of this invention can be used for the migration of best practices. For example, the system detects that the "East China Sales Data Analysis Workstation" has a high value score. The system can automatically trigger a similar workstation search, directly using the MRG of the high-scoring workstation as the MRG to be queried (step S101, no LLM conversion is required in this scenario), and search for other workstations besides itself in the workstation MRG database.
[0061] In this scenario, since the goal is to find structurally similar workstations to replicate their successful patterns, weights can be dynamically adjusted during vector fusion (step S104) and final sorting (step S108) to increase the weight coefficients of structural similarity (e.g., β and w_struc). The search results might reveal that the "South China Sales Data Analysis Workstation" has a high overall score, for example, a structural similarity of 0.87 and a semantic similarity of 0.85. The system can then suggest to managers that the successful configurations of workstations in East China (such as prompt word templates, model parameters, etc.) be migrated to South China, thereby facilitating the replication and promotion of best practices within the organization and improving overall operational efficiency.
[0062] Furthermore, the retrieval method of the present invention can also be applied to cross-type entity matching, such as finding suitable agents for workstations. When an agent bound to a "customer consultation workstation" goes offline due to a fault, the system can automatically trigger a replacement retrieval. At this time, the system uses the MRG of the "customer consultation workstation" as the MRG to be queried (step S101) and performs a retrieval in an "Agent Capability Declaration MRG Library" that stores descriptions of all available agent capabilities (step S105).
[0063] Since the Agent's capabilities are also modeled in MRG format, the entire retrieval process is completely consistent with the aforementioned embodiment. The system calculates the matching score between the MRG of each candidate Agent's capabilities and the MRG of the workstation's requirements through dual-path encoding and two-stage retrieval. Ultimately, the system recommends a highly matched Agent (e.g., whose capabilities cover the workstation's requirements) as a substitute and automatically completes the binding switch, thereby ensuring business continuity.
[0064] Furthermore, this invention can also be used to directly identify capability gaps in an organization from high-level strategic objectives. For example, when the CEO sets a strategic objective: "Increase customer satisfaction from 85% to 92% in the fourth quarter," the system first calls the LLM through modeling unit 10 to convert this strategic objective described in natural language into a structured requirement MRG (step S101). This MRG may contain sub-task branches such as "customer satisfaction analysis" and "recommendation of improvement measures."
[0065] Subsequently, the system uses this requirement MRG as the query and performs a search in the Agent Capability Declaration MRG database. If, during the fine-tuning stage (steps S107 and S108), it is found that the final matching score of all candidate Agents is lower than a preset threshold (e.g., 0.5), this clearly indicates that the organization currently lacks an Agent capable of fulfilling this strategic objective. Based on this finding, the system can automatically generate a report and make recommendations to managers: "To achieve the goal of improving customer satisfaction, it is recommended to create a 'Customer Satisfaction Analysis Agent' and a 'Recommendation Agent for Improvement Measures'." This enables managers to plan organizational capability building based on data-driven insights, demonstrating the strategic application value of this invention.
Claims
1. A retrieval method based on Task Relationship Graph (MRG), characterized in that, include: Obtain the MRG to be queried; generate a serialized text sequence based on the predetermined traversal order of the MRG and the text elements contained therein, and encode the serialized text sequence using a text embedding model to generate a semantic vector; generate a structure vector to represent the structural information of the MRG to be queried, wherein the structural information includes: node type sequence features, edge type sequence features, node degree sequence features, and graph statistical features; perform a weighted summation of the semantic vector and the structure vector to generate a unified retrieval vector; based on the unified retrieval vector, perform vector similarity analysis in a database containing multiple target MRG representations. The system performs a degree search to obtain a set of candidate MRGs; and determines the semantic similarity of each candidate MRG based on the semantic vector of the MRG to be queried and the semantic vectors of each candidate MRG; for each candidate MRG in the set of candidate MRGs, it calculates the structural similarity between the candidate MRG and the MRG to be queried by performing graph isomorphism detection separately, or by combining subgraph isomorphism determination and node type alignment scoring; based on the obtained semantic similarity and the calculated structural similarity, it obtains a final ranking score by weighted summation, and performs a final ranking of the set of candidate MRGs according to the final ranking score to obtain the final search results.
2. The method according to claim 1, characterized in that, The step of obtaining the MRG to be queried includes: receiving a query request in natural language form, and using a preset large language model to parse the query request into the MRG to be queried.
3. The method according to claim 1, characterized in that, The calculation of the structural similarity between the candidate MRG and the query MRG includes: If the graphs are identical, the structural similarity score is directly set to the maximum score of 1.0; Otherwise, calculate the structural similarity based on the following formula: Structural similarity = w1 × graph isomorphism detection score + w2 × subgraph isomorphism score + w3 × node type alignment score; Among them, w1, w2, and w3 are configurable weights.
4. The method according to claim 1, characterized in that, The step of generating structural information to characterize the MRG to be queried includes: converting the type of each node into a numerical code according to a predetermined traversal order of the nodes in the MRG to generate the node type sequence feature; converting the logical type of each edge into a numerical code according to a predetermined order of the edges in the MRG to generate the edge type sequence feature; extracting the in-degree and out-degree of each node according to a predetermined traversal order of the nodes in the MRG to generate the node degree sequence feature; and calculating the depth, average branch factor, total number of nodes, total number of edges, and proportion of each type of node and edge in the MRG to generate the graph statistical features.
5. The method according to claim 4, characterized in that, The steps for generating the structure vector include: normalizing the node type sequence features, edge type sequence features, node degree sequence features, and graph statistical features to the [0,1] interval, concatenating them into a unified structure vector, and aligning the dimensions to the same dimension as the semantic vector through a linear projection layer, so that the two vectors can be weighted and fused in a unified vector space.
6. The method according to claim 1, characterized in that, The step of performing a weighted summation of the semantic vector and the structural vector is achieved through the following formula: V_retrieval = α All_here + β In_struc; Wherein, V_retrieval is the unified retrieval vector, V_sem is the semantic vector, V_struc is the structural vector, and α and β are preset weight coefficients.
7. The method according to claim 6, characterized in that, The weighting coefficients α and β are dynamically adjusted according to the preset retrieval scenario.
8. The method according to claim 1, characterized in that, The vector similarity retrieval is an approximate nearest neighbor (ANN) retrieval.
9. The method according to claim 1, characterized in that, The step of obtaining the final ranking score through weighted summation is implemented using the following formula: Score_final = w_sem Score_sem + w_struc Score_struc, where Score_final is the final ranking score, Score_sem is the semantic similarity, Score_struc is the structural similarity, and w_sem and w_struc are preset weight coefficients.
10. A retrieval device based on a task relationship graph (MRG), characterized in that, include: The modeling unit is used to model business entities as a task relationship graph (MRG) structure to obtain the MRG to be queried. A text processing unit is configured to generate a serialized text sequence based on a predetermined traversal order of the MRG to be queried and the text elements contained therein. A semantic encoding unit is used to encode the serialized text sequence using a text embedding model to generate a semantic vector; The structural feature extraction unit is used to generate node type sequence features, edge type sequence features, node degree sequence features, and graph statistical features based on the MRG to be queried. The vector fusion unit is used to generate a structure vector based on the node type sequence features, the edge type sequence features, the node degree sequence features, and the graph statistical features, and to perform a weighted summation of the semantic vector and the structure vector to generate a unified retrieval vector. The coarse screening unit is used to perform vector similarity retrieval in a database containing multiple target MRG representations based on the unified retrieval vector to obtain a set of candidate MRGs; and to determine the semantic similarity of each candidate MRG based on the semantic vector of the MRG to be queried and the semantic vector of each candidate MRG. The fine ranking unit is used to calculate the structural similarity between each candidate MRG in the candidate MRG set and the query MRG by performing graph isomorphism detection individually, or by combining subgraph isomorphism determination and node type alignment scoring; and based on the obtained semantic similarity and the calculated structural similarity, obtain the final ranking score by weighted summation, and perform the final ranking of the candidate MRG set according to the final ranking score to obtain the final retrieval result.