Graph contrastive learning method for solving graph combinatorial optimization problems

By constructing an attributed graph database and dynamically generating a taboo prototype library, the problems of pattern explosion and causal ambiguity in existing technologies are solved, efficient and accurate solution of graph combination optimization problems is achieved, and the efficiency and success rate of VLSI design are improved.

CN120523870BActive Publication Date: 2025-09-23NANJING ARTIFICIAL INTELLIGENCE CHIPS RES INST OF AUTOMATION CHINESE ACAD OF SCI +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511030422.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-09-23
Estimated Expiration
2045-07-25

Smart Images

  • Figure CN120523870B_ABST
    Figure CN120523870B_ABST
Patent Text Reader

Abstract

This invention discloses a graph contrast learning method for solving graph combinatorial optimization problems. The method includes: constructing VLSI design data with success or failure performance labels into an attributed graph database; discovering high-contrast subgraph instances strongly correlated with failure cases through contrast subgraph mining; dynamically abstracting the massive subgraph instances into a controllable number of taboo prototypes through online structural clustering; extracting the minimum causal core from the taboo prototypes through differential perturbation analysis; and converting this minimum causal core knowledge base into high-penalty terms, which are integrated into a standard combinatorial optimization solver to guide it in actively avoiding known design flaws when solving new problems, and outputting optimized graph problem solutions. This invention improves solution efficiency and success rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence, and in particular to a graph comparison learning method for solving graph combination optimization problems. Background Art

[0002] Graph combinatorial optimization (GCO) is a core optimization problem in graph theory. Its goal is to find the optimal combination of vertices, edges, or subgraphs in a given graph that satisfies specific conditions. Research on GCO has far-reaching significance: theoretically, it continues to advance the development of graph theory; and in applications, it plays a key role in numerous fields, including network flow analysis, logistics planning, and bioinformatics. However, most GCO problems are NP-hard, meaning that the time cost of solving them with traditional exact algorithms increases exponentially as the problem scale increases, making them unacceptable for practical applications. In VLSI (very large-scale integrated circuit) physical design, layout and routing is one of the most typical and challenging applications of GCO. A chip designed at an advanced process node may contain billions of transistors and tens of millions of interconnects. Any small local wiring structure flaw can lead to the failure of the entire chip or a significant reduction in yield. Therefore, how to efficiently solve such large-scale GCO problems, especially how to learn and avoid failure modes from massive amounts of design data, has significant theoretical and commercial value for improving the level of chip design automation and shortening the R&D cycle.

[0003] Currently, the industry is exploring various approaches to uncover design patterns from historical data. One mainstream approach is to apply data mining techniques known as Frequent Subgraph Mining (FSM), such as the gSpan algorithm, to discover frequently occurring subgraph structures by setting a support threshold. Machine learning methods are also being widely studied. Supervised learning algorithms face training bottlenecks due to the difficulty in obtaining high-quality labeled datasets. While reinforcement learning algorithms do not require labels, they lack global differentiability and face difficulties in gradient estimation, severely impacting training efficiency and solution quality. Existing general-purpose unsupervised algorithms struggle to directly obtain good optimization strategies. Graph Contrastive Learning (GCL), a key unsupervised learning paradigm, has achieved remarkable results in tasks such as graph feature extraction. However, existing GCL frameworks have not been specifically designed for the characteristics of the GCO problem, hindering its further application in this area.

[0004] However, existing technologies face some technical problems when extracting valuable and generalizable taboo patterns from massive design data, including pattern explosion and topological rigidity problems as well as causal ambiguity and core structure drowning problems. Summary of the Invention

[0005] The purpose of the invention is to provide a graph comparative learning method for solving graph combinatorial optimization problems, in order to solve at least one technical problem existing in the prior art.

[0006] The technical solution is a graph contrast learning method for solving graph combinatorial optimization problems, including:

[0007] Construct an attributed graph database based on the original VLSI layout database and the associated performance label set;

[0008] Scan the attributed graph database and perform contrast analysis based on performance labels to mine high-contrast subgraph instance flows;

[0009] Receive a stream of high-contrast subgraph instances and dynamically abstract and generate a taboo prototype library through online structural clustering;

[0010] Combined query of the taboo prototype library and the attributed graph database, by performing differential perturbation analysis on the prototypes in the library, to extract the minimum causal core knowledge base;

[0011] Load the minimal causal core knowledge base, build a guided combinatorial optimization solver, and use it to solve the new graph problem to be optimized, and output the optimized graph problem solution.

[0012] Preferably, further:

[0013] Obtain the original VLSI layout database containing GDSII, LEF / DEF format files and the associated design performance label set, and construct an attributed graph database by mapping the geometric entities (metal trace segments, vias) in the VLSI layout into graph nodes and mapping the connections or adjacency relationships between geometric entities into graph edges. In this way, physical attributes (net ID, metal layer, etc.) are attached to the graph nodes, and geometric attributes (parallel length, spacing, etc.) are attached to the graph edges.

[0014] Scan the VLSI layout graph data with success and failure labels in the attributed graph database, perform contrast analysis by calculating the support count ratio of subgraph structures in the failure graph and the success graph, and mine high-contrast subgraph instance flows that are strongly correlated with VLSI design failures;

[0015] For each subgraph instance in the high-contrast subgraph instance stream, a structural feature vector containing multi-scale local degree distribution, cyclic basis features, and attribute distribution is calculated. By using an online structural clustering method that compares the structural feature vector with existing prototypes in the taboo prototype library, a taboo prototype library reflecting the VLSI design defect patterns is dynamically abstracted and generated.

[0016] A differential perturbation analysis is performed on each prototype in the taboo prototype library with a single element removed. The causal contribution score of each element is calculated by comparing the contrast scores before and after removal, and the minimum causal core knowledge base containing elements with high causal contribution is extracted.

[0017] Each core in the minimum causal core knowledge base is converted into a set of logical predicates and a corresponding penalty function is constructed. The function is then integrated into the VLSI detailed wiring solver to construct a guided combinatorial optimization solver. This solver can detect and avoid known design defect patterns in real time when solving new VLSI wiring problems, and output a VLSI layout solution that includes optimized wiring path coordinates, metal layer allocation, and via locations.

[0018] Beneficial effect: The present invention solves the problems of pattern explosion, topological rigidity and causal ambiguity that exist in the existing technology when mining the causal core of failure modes from massive design data, and improves the efficiency and success rate of solving large-scale graph combinatorial optimization problems. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 A flowchart of the steps of a graph comparative learning method for solving graph combinatorial optimization problems provided in an embodiment of the present application.

[0020] Figure 2 A flowchart of the steps for mining high-contrast subgraph instance flow provided in an embodiment of the present application.

[0021] Figure 3 A flowchart of the steps for generating an extended candidate set provided in an embodiment of the present application.

[0022] Figure 4 A flowchart of the steps for generating a taboo prototype library provided in an embodiment of the present application. DETAILED DESCRIPTION

[0023] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0024] It should be noted that the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or are inherent to these processes, methods, products or apparatus.

[0025] Research has found that traditional FSM-based methods generate massive, highly redundant pattern outputs. For example, a three-node core structure that actually causes a failure, along with any four-node or five-node structures containing it, will be reported as separate patterns, resulting in significant redundancy and making subsequent analysis difficult. Furthermore, FSMs require strict graph isomorphism matching. If a taboo pattern appears in another graph with only an insignificant additional branch or slightly different attributes, the FSM will treat it as a completely different pattern, thus losing its ability to generalize universal laws, known as topological rigidity. Furthermore, there are deeper issues of causal ambiguity and core structure drowning. The so-called failure patterns found by existing methods are essentially merely a reflection of correlation, not proof of causality. A highly correlated pattern that is discovered is often a relatively bloated structure containing the true core of the failure, mixed with a large amount of structural noise that has no direct causal connection to the failure. Existing technologies lack an effective mechanism to separate the true from the false and accurately locate and isolate the minimum causal core substructure. As a result, the extracted rules are either too large to be applied, or have poor generalization capabilities due to the inclusion of non-causal elements, and are unable to fundamentally guide subsequent design optimization.

[0026] like Figure 1 As shown in the figure, a graph contrastive learning method for solving graph combinatorial optimization problems is proposed, which includes the following steps:

[0027] An attributed graph database is constructed based on the original VLSI layout database and the associated performance label set.

[0028] Specifically, the original, multi-format VLSI layout design data and its associated success or failure performance labels are converted into a unified attributed graph database with rich physical properties and performance labels, laying the data foundation for subsequent analysis.

[0029] Scan the attributed graph database and perform contrast analysis based on performance labels to mine high-contrast subgraph instance flows in VLSI layout designs.

[0030] Specifically, based on contrast scores rather than traditional support, high-risk local structural patterns that are significantly enriched in failure cases are mined from the attributed graph database to form a high-contrast subgraph instance flow in VLSI layout design.

[0031] It receives a stream of high-contrast subgraph instances and dynamically abstracts and generates a library of taboo prototypes through online structural clustering.

[0032] Specifically, in order to solve the problem of excessive instances, an online clustering method is used to dynamically and in real time summarize the massive subgraph instances into a controllable number of representative taboo prototypes, and store them in the taboo prototype library.

[0033] The taboo prototype library and the attributed graph database are jointly queried, and the minimum causal core knowledge base is extracted by performing differential perturbation analysis on the prototypes in the library.

[0034] Specifically, the differential perturbation analysis method is used to refine each taboo prototype, separate and identify the most critical minimum causal core structure that actually causes failure, and form a minimum causal core knowledge base.

[0035] The minimal causal core knowledge base is loaded, and a guided combinatorial optimization solver is constructed. This solver is then used to solve the new graph problem to be optimized, outputting an optimized graph solution. This solution includes detailed routing paths for each signal net across multiple metal layers; safe spacing between critical signal pairs; optimized via placement to reduce signal integrity issues; and a hierarchical routing strategy for clock and data networks. The routing solution avoids known failure modes identified in the minimal causal core knowledge base.

[0036] Specifically, the extracted minimum causal core knowledge base is converted into specific mathematical penalty terms and integrated into a standard combinatorial optimization solver, thereby constructing a guided solver that can actively avoid known design risks to solve new optimization problems.

[0037] In one embodiment of the present application, the process of constructing an attributed graph database is as follows: reading and parsing the original VLSI layout database and the associated design performance label set, the database contains files in formats such as GDSII, LEF / DEF, and the label set marks each layout as successful or failed, thereby obtaining an original layout object set. Each layout object in the original layout object set is traversed, and geometric entities such as metal routing segments and vias therein are mapped as nodes of the graph, and the connections or adjacency relationships between geometric entities are mapped as edges of the graph, thereby generating a basic topology graph set. Based on the physical information in the original layout object set, each graph in the basic topology graph set is attributed, and physical attributes are attached to each node and edge (for example, the node is attached with the corresponding wire net ID and metal layer; the edge is attached with the parallel length, spacing, etc.), thereby obtaining an intermediate attribute graph set. The success or failure label in the design performance label set is associated one-to-one with each graph in the intermediate attribute graph set, and finally an attributed graph database is constructed.

[0038] In some embodiments, the solver performs the following technical operations: planning a signal transmission path on a specified routing layer;

[0039] Optimize routing congestion to improve routing yield; meet geometric constraints for design rule checking; and minimize total routing length to reduce latency.

[0040] In summary, this embodiment improves solution efficiency and success rate, including reducing the total wiring length; reducing the risk of signal crosstalk; improving the manufacturability of the layout; and shortening the design verification cycle.

[0041] like Figure 2 As shown, according to one aspect of the present application, mining high-contrast subgraph instance flows, that is, mining high-contrast wiring structure instance flows that are strongly correlated with circuit performance failures based on success and failure labels of VLSI layout design, includes:

[0042] Scan the attributed graph database, extract unilateral subgraphs, and calculate the support counts of the unilateral subgraphs in the success graph and the failure graph to form the initial mining queue;

[0043] Extract candidate subgraphs from the initial mining queue and perform attribute-topology composite expansion on the candidate subgraphs to generate an expanded candidate set;

[0044] For each extended candidate subgraph in the extended candidate set, the contrast score is calculated by calculating the support count of the extended candidate subgraph in the success graph and the failure graph;

[0045] The expansion candidate subgraphs with contrast scores higher than a preset threshold are output to the high-contrast subgraph instance stream, and the remaining expansion candidate subgraphs are added to the initial mining queue.

[0046] For example, wiring structure examples include: topological connection relationships of signal line networks; via connection information between metal layers; and geometric parameters of key signal paths.

[0047] like Figure 3 As shown, according to one aspect of the present application, generating an extended candidate set includes:

[0048] Traverse the boundary nodes of the candidate subgraph and generate a pure topological extension set by adding new edges and associated nodes;

[0049] Identify the generalized attributes in the candidate subgraph, replace the generalized attributes with specific attribute instances retrieved from the attributed graph database, and generate attribute specialization sets;

[0050] The pure topological extension set and the attribute specialization set are merged, and duplicate subgraphs are removed through graph isomorphism detection to form an extension candidate set.

[0051] In one embodiment of the present application, an attributed graph database is scanned, in which each graph is marked as successful or failed. All unique unilateral subgraphs with attributes are extracted, and the support count of each unilateral subgraph in the two types of graphs, success and failure, is calculated, that is, the number of times it appears in the two types of graphs. These unilateral subgraphs and their support counts together constitute the initial mining queue, which serves as the starting point for subsequent extended mining. Attribute-topology composite expansion is performed on the candidate subgraphs to generate an extended candidate set, aiming to solve the problems of topological rigidity and weak attribute perception ability of traditional frequent subgraph mining (FSM) algorithms. The candidate subgraph g is taken out from the initial mining queue and attribute-topology composite expansion is performed on it. This expansion involves two parallel operations: pure topological expansion, which traverses the boundary nodes of the candidate subgraph g and generates a set of new subgraphs that are topologically larger than g by adding new edges and associated nodes; and attribute specialization, which identifies generalized attributes in the candidate subgraph g. For example, if the net type of a node is marked as any, the attributed graph database is queried and the generalized attributes are replaced with specific attribute instances, such as clock nets or data nets. This generates a set of new subgraphs that are topologically isomorphic to g but with more specific attributes. The subgraph sets generated by the two methods are merged, and duplicate subgraphs are removed through graph isomorphism detection, ultimately forming an expanded candidate set. For each expanded candidate subgraph g′ in the expanded candidate set, the attributed graph database is rescanned to calculate its support count Freq(Fail′) in the failure graph and its support count Freq(Success′) in the success graph. Based on the support count, the contrast score Score(g′) is calculated for each extended candidate subgraph g′. The calculation formula is: Score(g′) = Freq(Fail′) / (Freq(Success′) + ε); where Score(g′) represents the contrast score of the candidate subgraph g′. The higher the score, the stronger its correlation with the failure case. Freq(Fail′) is the frequency (support count) of the candidate subgraph g′ in the failure graph database; Freq(Success′) is the frequency (support count) of the candidate subgraph g′ in the success graph database; ε is a very small positive smoothing factor, for example, 10 -6 , which prevents the denominator from being zero and ensures computational stability. Subgraph candidates with contrast scores Score(g′) exceeding the preset threshold σ are output to the high-contrast subgraph instance stream for subsequent processing. The remaining subgraph candidates are re-added to the mining queue, awaiting the next round of iterative expansion, until the entire mining queue is empty.

[0052] This example utilizes a composite attribute-topology expansion mechanism and uses contrast scores as the core driver of iterative mining, enabling early focus on highly discriminative failure modes that are tightly coupled to physical properties. The mining process in this example possesses dual exploration capabilities: it can explore new topological structures by adding edges and nodes (topological expansion), and it can also explore new attribute combinations (attribute specialization) within existing structures by refining the physical properties of their nodes (for example, specializing any net into a clock net). Furthermore, guided by contrast scores, each algorithm expansion step is designed to maximize the distinction between failure and success cases, rather than searching for useless patterns that are common in both types of cases. In VLSI wiring scenarios, this means the algorithm can automatically discover complex patterns that pose a high risk only when specific net types (such as clock lines) are combined with specific structures (such as U-shaped windings), improving mining efficiency and the relevance of the results.

[0053] Successful cases refer to layout designs that pass timing analysis and signal integrity checks; failed cases refer to layout designs with problems such as setup time violations and excessive crosstalk; the contrast score reflects the strength of the causal relationship between a specific wiring structure and design failure.

[0054] like Figure 4 As shown, according to one aspect of the present application, generating a taboo prototype library includes:

[0055] For each high-contrast sub-image instance in the high-contrast sub-image instance stream, a structural feature vector of each high-contrast sub-image instance is calculated by using a graph theory statistical method;

[0056] Compare the structural feature vector with the feature vector of each prototype in the pre-built taboo prototype library to determine the minimum structural distance;

[0057] Based on the relationship between the minimum structural distance and a preset threshold, it is determined whether the high-contrast subgraph instance belongs to an existing prototype or represents a newly discovered pattern, and the taboo prototype library is updated or expanded accordingly.

[0058] According to one aspect of the present application, updating or expanding the taboo prototype library based on the relationship between the minimum structural distance and a preset threshold value includes:

[0059] Determine the best matching prototype in the taboo prototype library based on the minimum structural distance;

[0060] If the minimum structural distance is less than the preset threshold, the high-contrast sub-image instance is determined to be a new instance of the best matching prototype, and the statistical information of the best matching prototype is updated without adding a new prototype;

[0061] Otherwise, it is determined that the high-contrast subgraph instance represents a newly discovered pattern, and a new prototype is created in the taboo prototype library based on the topological structure and structural feature vector of the high-contrast subgraph instance.

[0062] According to one aspect of the present application, the structural feature vector is calculated by a graph theory statistical method, including:

[0063] Read high contrast sub-image instances, in order:

[0064] Calculate the multi-scale local degree distribution and obtain the multi-scale distribution histogram;

[0065] Calculate the cyclic basis characteristics and obtain the cyclic basis eigenvalues;

[0066] Count the attribute distribution of nodes and edges to obtain the attribute distribution histogram;

[0067] The multi-scale distribution histogram, cyclic basis eigenvalue and attribute distribution histogram are concatenated and normalized to form the structural feature vector of the VLSI wiring diagram.

[0068] According to one aspect of the present application, obtaining a cyclic basis eigenvalue includes:

[0069] Determine the size of the cyclic basis of the high-contrast subgraph instance, which is the number of edges minus the number of nodes plus the number of connected components of the high-contrast subgraph instance;

[0070] Find the lengths of the k shortest independent loops in the high-contrast sub-image instance; where k is a preset parameter;

[0071] The size of the cyclic basis and the lengths of the k independent loops are combined to form the eigenvalue of the cyclic basis.

[0072] In one embodiment of the present application, an empty taboo prototype library of VLSI design defects is initialized. The library is used to store abstracted taboo prototypes, where each prototype contains at least: a representative topological structure, a structural feature vector, and associated statistical information (such as instance count, average contrast score, etc.). Each high-contrast subgraph instance p in the high-contrast subgraph instance stream is received one by one. new , for each received instance p new , calculate its structural feature vector v by using a non-neural network method based on graph theory statistics new The vector is intended to capture the morphological and attribute characteristics of the subgraph from multiple dimensions. In a preferred implementation, the feature vector is composed of the following parts: new For each node v in instance p, calculate its new The internal degree is called the first-order degree deg1(v); the sum of the first-order degrees of all its neighboring nodes is calculated as its second-order degree deg2(v) =Σu∈Neighbors(v) deg1(u), where u is a neighbor node and Neighbors(v) is the set of neighbor nodes of node v. Generate a normalized histogram of the first-order and second-order degrees of all nodes, and concatenate the two histograms to form a multi-scale distribution histogram, which reflects the local connection density of the nodes. Calculation example p new The size of the cyclic basis of is given by the formula |E|-|V|+|C|, where |E| is the number of edges, |V| is the number of nodes, and |C| is the number of connected components; find the instance p new The lengths of the k shortest independent cycles in the graph (k is a preset parameter, for example, k=3); the size of the cycle basis and the lengths of the k shortest cycles are combined to form the cycle basis eigenvalue, which is intended to describe the cycle characteristics of the subgraph. new The number of nodes with different attributes (such as wire mesh type and metal layer) in the graph is counted to form a node attribute histogram; the number of edges with different attributes is counted, and continuous attributes (such as parallel length) are binned and counted to form an edge attribute histogram; the attribute histograms of nodes and edges are spliced ​​to obtain the final attribute distribution histogram. The multi-scale distribution histogram, cyclic basis eigenvalue and attribute distribution histogram are flattened and spliced, and then L2 norm normalized to form a high contrast subgraph instance p new The structural characteristic vector v new The calculated structural feature vector v new Compare the feature vectors of each prototype in the taboo prototype library, for example, by calculating the cosine similarity to measure the similarity between vectors to determine the minimum structural distance d min , for any prototype A in the library i , whose eigenvector is v i , then p new With A i The structural distance d i It can be calculated as: d i =1- v new ·v i / (∣∣v new ∣∣∣∣v i ∣∣); By traversing all prototypes in the library, you can find the prototype that matches p new The most similar best matching prototype A best , and record its minimum structural distance d min .

[0073] According to the minimum structural distance d min The relationship between the similarity threshold δ and the preset similarity threshold is used to perform the judgment and update operation: if d min <δ, indicating that the new instance p new The best matching prototype Abest If the structure is highly similar, then p new Belongs to the existing prototype Abest The complexity of the model should not be increased at this time, that is, no new prototype is created, only the existing prototype A is best The statistics of d are updated online. For example, its instance count is increased by one, and its average contrast score is updated by rolling average, which effectively suppresses the redundant growth of the pattern. On the contrary, if d min ≥δ, indicating that the new instance p new If there are significant differences from all known prototypes in the current library, it may represent a new taboo pattern that has not been discovered before, then p is determined to be new Represents a newly discovered pattern, this time with a new instance p new The topological structure and eigenvector v new Based on the new instance p, a new prototype is created in the taboo prototype library. The initial statistics of the new entry (such as instance count and average contrast score) are based on the new instance p. new This process continues until all high-contrast subgraph instances have been processed, ultimately resulting in a controllable number of taboo prototypes with complete content. This process aims to address the pattern explosion and result redundancy issues caused by instance mining. Instead of simply storing all mined high-contrast subgraph instances in a list, by repeatedly performing the above characterization, comparison, and conditional update processes on all instances in the high-contrast subgraph instance stream, dozens or hundreds of representative taboo prototypes can be summarized and abstracted from tens of thousands of specific instances, forming a complete taboo prototype library.

[0074] This embodiment calculates a structural feature vector for each high-contrast subgraph instance, composed of multi-scale distributions, cyclic basis features, and attribute distributions. During the mining process, this vector is compared in real time with a dynamically maintained library of taboo prototypes. Online clustering operations, either absorption or creation, are performed based on the relationship between distance and thresholds, achieving efficient summarization and abstraction of massive, redundant failure mode instances. This embodiment does not employ rigid graph isomorphism matching, but instead maps each subgraph into a feature space that characterizes its shape and complexity. In this space, even if two subgraphs differ slightly in topological details (e.g., having an extra or missing insignificant edge), as long as their core structures are similar (e.g., two long parallel lines), their feature vectors will be spatially close to each other. This enables the system to automatically summarize thousands of specific, subtle pattern variations into dozens of representative taboo prototypes, improving the usability and interpretability of mining results.

[0075] According to one aspect of the present application, a minimum causal core knowledge base of VLSI wiring failure modes is extracted, including:

[0076] Perform single-element perturbation on the topological structure of each prototype in the taboo prototype library to generate a set of perturbation patterns;

[0077] For each perturbation pattern in the perturbation pattern set, recalculate the contrast score of each perturbation pattern in the attributed graph database;

[0078] A causal contribution score is calculated for each removed element in the prototype by comparing the contrast score of each perturbed pattern with the original contrast score of that prototype;

[0079] Identify all elements whose causal contribution scores are above a preset threshold, constituting a single minimal causal core of the archetype;

[0080] The single minimal causal core of all prototypes is aggregated to form a minimal causal core knowledge base of VLSI wiring failure modes.

[0081] According to one aspect of the present application, calculating a causal contribution score for each removed element in the prototype includes:

[0082] The causal contribution score is obtained by subtracting the contrast score of the perturbation pattern corresponding to the removed element from the original contrast score of the prototype.

[0083] According to one aspect of the present application, the contrast score of each disturbance pattern in the attributed graph database is recalculated as follows:

[0084] Find the graph ID list corresponding to each edge that constitutes the perturbation pattern in the preconfigured edge-graph ID index; where the edge-graph ID index maps each unique attributed edge to the graph ID list containing the edge;

[0085] Perform an intersection operation on the graph ID list to obtain the final graph ID list containing the perturbation pattern;

[0086] The differential support counts of the perturbation pattern in the success graph and the failure graph are determined based on the final graph ID list, and the contrast scores of the success and failure cases of the VLSI design are calculated.

[0087] In one embodiment of the present application, prototypes are selected one by one from the taboo prototype library for analysis. For the prototype A to be analyzed, a single element perturbation is performed on its topological structure, that is, a node or an edge in A is systematically and one by one removed, thereby generating a set of corresponding perturbation pattern sets {A′}. For each perturbation pattern A′ in the perturbation pattern set, it is necessary to recalculate its contrast score Score(A′) in the entire attributed graph database. In order to achieve efficient calculation, an edge-graph ID index can be pre-constructed, which maps each unique, attributed edge to a list of IDs of all graphs containing the edge. The specific calculation process includes: finding each edge that constitutes the perturbation pattern A′ and each edge e that constitutes A′ in the index. j The corresponding graph ID list L j , forming a graph ID list; performing efficient intersection operations (such as bit operations) on the multiple graph ID lists found to obtain the final graph ID list L containing the perturbation pattern A′ intersection = L1∩L2∩...; Based on the design performance labels, count the number of successful and failed graphs in the final graph ID list to obtain the differential support count (Freq(Fail′), Freq(Success′)); Based on the differential support count, calculate the contrast score (Score(A′)) of the perturbation pattern A′ using the formula Score(A') = Freq(Fail') / (Freq(Success') + ε). By comparing the contrast score of each perturbation pattern with the original contrast score of the prototype, calculate the causal importance score (CIS) for each removed element x (node ​​or edge) in prototype A using the following formula: CIS(x) = Score(A) - Score(Ax); where CIS(x) represents the causal contribution score of element x; Score(A) is the original contrast score of prototype A; and Score(Ax) is the contrast score of the perturbation pattern after element x is removed. The CIS(x) score intuitively measures the decrease in the failure or taboo level of the pattern when element x is missing. A high CIS score indicates that element x plays a key causal role in the formation of the failure pattern. All elements with CIS scores above a preset causal threshold τ are identified. The subgraph consisting of these highly contributing elements is defined as the single minimal causal core of the archetype. The minimal causal cores of all archetypes are aggregated to form a minimal causal core knowledge base. This aims to address the problems of causal ambiguity and core structure drowning, specifically pinpointing the minimal structure that plays a decisive role within complex failure patterns.

[0088] This embodiment systematically perturbs taboo prototypes with single elements, efficiently recalculates the contrast score of each perturbation pattern using edge-graph ID indexing, and then quantifies the causal contribution of each element according to a formula, effectively extracting the minimum causal core from strongly correlated patterns. This transforms the abstract causal problem into a computable and quantifiable differential analysis problem. In VLSI design, a failed prototype containing 10 nodes may actually be caused by the specific adjacency relationship between two of these nodes (for example, two specific signal lines). By removing these 10 nodes one by one and observing the decrease in their criticality (i.e., contrast score), this embodiment uses a data-driven approach to precisely identify the two critical nodes whose removal will cause a sharp drop in the score. The final output is no longer a vague, bloated high-risk area, but rather a verified, irreducible pathogenic gene, making it possible to generate precise, non-redundant taboo rules that can be directly used to guide design.

[0089] According to one aspect of the present application, loading a minimal causal core knowledge base to construct a guided combinatorial optimization solver and using it to solve the problem includes:

[0090] Convert each minimum causal core in the minimum causal core knowledge base into a high penalty term or constraint paradigm;

[0091] Integrate all high-penalty terms or constraint paradigms into the cost function or constraint engine of the standard combinatorial optimization solver to build a guided combinatorial optimization solver;

[0092] When using a guided combinatorial optimization solver to solve a new graph problem to be optimized, the solver checks in real time during the search process whether the local solution it generates is isomorphic to any core in the minimum causal core knowledge base, and avoids the solution path through high penalty terms or constraint paradigms when isomorphic.

[0093] According to one aspect of the present application, each minimal causal core is converted into a high-penalty term, comprising:

[0094] Decompose the graph structure and attribute information of each minimum causal core into a corresponding set of formalized logical predicates;

[0095] Based on the set of logical predicates of each core, a single pattern penalty function is constructed. The function returns the preset penalty value if and only if the local solution satisfies all the logical predicates of the core;

[0096] The individual pattern penalty functions corresponding to all minimum causal cores are summed to form an aggregate penalty term, which is the high penalty term integrated into the cost function.

[0097] In one embodiment of the present application, a minimum causal core knowledge base is loaded. For each minimum causal core in the knowledge base, it is converted into a high penalty term or a hard constraint paradigm. In a preferred implementation, the conversion to a high penalty term is adopted. Specifically, the penalty terms corresponding to all cores in the knowledge base are integrated into the cost function of a standard combinatorial optimization solver (e.g., a VLSI detailed router) to construct a guided combinatorial optimization solver. The process of converting a single minimum causal core into a high penalty term includes: decomposing into a set of logical predicates: decomposing the graph structure and attribute information of the minimum causal core C into a formalized set of logical predicates. For example, a clock line and a data line with a parallel length of more than 50 microns on the M3 layer will be decomposed into a set of formalized logical predicates. Or a core representing two different types of wire nets running in parallel for a long distance on a specific metal layer can be decomposed into: {IsType(n1, TypeA), IsType(n2, TypeB), IsLayer(e1, M4), IsParallel(e1, e2), Length(e1)>L th}, where n represents a node (line network), e represents an edge (line segment), and L th IsType(n1, TypeA) indicates that node n1 (i.e., a certain line network) belongs to type TypeA; IsLayer(e1, M4) indicates that edge e1 (a line segment) is located on metal layer M4; IsParallel(e1, e2) indicates that edges e1 and e2 are parallel; Length(e1) indicates the length of edge e1. Construct a single pattern penalty function: Based on the above set of logical predicates, construct a single pattern penalty function P for this core C. C (S), which takes the local solution S generated by the solver during the search process as input. If and only if the local solution S satisfies all logical predicates of the core, the function returns a preset, maximum penalty value W C , otherwise it returns 0. This function can be expressed as the product of the indicator functions corresponding to all predicates: P C (S) = W C * ∏ i I(predicate i (S)), where ∏ is the product symbol, I( ) is the indicator function, and predicate i Represents the i-th logical predicate. Forming the aggregate penalty term: Aggregate (for example, sum) the individual pattern penalty functions corresponding to all minimum causal cores to form the aggregate penalty term P total (S)=∑ C P C (S). The aggregate penalty term P total(S) is integrated into the cost function of a standard combinatorial optimization solver (e.g., a VLSI detailed router) to form a new objective function: new (S)=Obj original (S)+P total (S), where Obj original (S) is the original cost function of the combinatorial optimization solver when the knowledge base guidance term (i.e., penalty term) is not introduced. The modified solver is a guided combinatorial optimization solver. A new graph problem to be optimized (e.g., a new VLSI wiring design) is input into the guided combinatorial optimization solver. In the process of searching the solution space, when the solver tries to generate a local solution (e.g., planning a specific wiring path), it will check in real time and efficiently whether the local solution satisfies all the logical predicates corresponding to any core in the minimum causal core knowledge base, that is, whether it is isomorphic with any taboo pattern. If it is isomorphic, it means that once the local solution is formed, it will trigger P total The high penalty value in (S) makes the cost of the overall solution including this local solution Obj new (S) increases dramatically. Based on its inherent optimization mechanism, the solver naturally abandons or significantly reduces the priority of selecting this solution path, effectively mitigating the risk of known design flaws. The solver finds the optimal solution within the effectively guided search space and outputs the optimized graph problem solution. This process is the final step in applying the mined causal knowledge to actual optimization problems. For example, in a VLSI detailed routing scenario, preliminary learning has stored a rule in the minimum causal core knowledge base: two adjacent signal lines belonging to high-speed buses A and B, running parallel on the M4 metal layer with a spacing less than a safe value, must not exceed a maximum of 20 microns. When routing a new chip, if the guided router attempts to generate a path that places two lines from buses A and B in close proximity and parallel on the M4 layer for 25 microns, its real-time checking module immediately detects that this path is isomorphic to the core in the knowledge base. In this case, the cost function returns a very large penalty, forcing the router to abandon this path and explore other routing solutions that bypass or increase spacing. Tests show that the solver guided by this embodiment reduces the occurrence rate of such critical design defects.

[0098] This embodiment introduces graph node degree information, which is of great significance to multiple graph combinatorial optimization problems, as heuristic information during the comparison sample generation stage. This method performs data augmentation by deleting edges based on the node degree. Each time, the edge connected to the node with the smallest degree in the current graph is deleted, enabling hierarchical learning of graph structural information. Edges connected to specific nodes are gradually and heuristically deleted from the graph, allowing the model to gradually remove edge information from the dataset and more effectively focus on, capture, and utilize the structural information inherent in the graph. This overcomes the problem that traditional graph comparative learning frameworks only randomly perturb the graph structure, making it difficult to generate comparison samples that meet the characteristics of graph combinatorial optimization problems.

[0099] In a specific embodiment of the present application, in order to improve the model's learning ability for comparison samples, a bootstrapping comparison framework is used as the basic framework of the model, and gradient pruning technology is used to prevent the model from overfitting phenomena such as feature collapse during training. In terms of comparison target setting, the fused Gromov-Wasserstein distance in optimal transportation theory is introduced to measure the differences in the representation results of different comparison samples. It fully considers the local structural information and global structural information in the graph, and measures the differences between different distributions and different spaces. It enhances its ability to interpret structural data in different domains and enhances the model's ability to solve graph combination optimization problems. It solves the problem that traditional graph comparison learning methods average nodes in the setting of training targets and ignore important topological structure information in the graph.

[0100] To facilitate the subsequent description, we briefly define the problem background. Assume that the graph to be solved is G=(V, E), where V={v1, v2, ..., v n} is the set of vertices in the graph, E (contained in V×V) is the set of edges in the graph, and if two vertices v i and v j There is an edge between them, then the edge e ij =(v i , v j )∈E (a node is considered to be connected to itself). The adjacency matrix of graph G is A={0, 1} N×N , when there is an edge connecting node i and node j, the element a in the adjacency matrix ij =1, otherwise a ij =0; the degree matrix of graph G is D={d1, d2, …, d n}, where d i Indicates that in the graph G, i The number of connected edges, N is the total number of nodes in the graph. Since the graph G will be changed and iterated during the training process, the input graph in the i-th round of training is defined as Gi =(V i , E i ), where V i is the node set of the i-th round graph, E i is the edge set of the graph in round i. In round i training, the graph G of the current input is calculated i =(V i , E i )’s degree matrix D i , and sort the nodes in ascending order according to degree to get V Di According to the manually set edge deletion ratio r and the total number of current edges, calculate the number of edges m that need to be deleted in the current round of training i =r×sizeof(E i ), where sizeof(E i ) represents the graph G in the current i-th round of training i Middle edge set E i The total number of edges, that is, the number of elements in the set. Based on this order, take out the node with the current non-zero and minimum degree in turn, and delete all the edges adjacent to it until the number of deleted edges is not less than m i . Node sequence V after sorting by degree di ={V Di [1], V Di [2],…,V Di [n i ]},n i ≤n,∑ j=1 ni deg(V di [j])≥m i , where n i is the number of nodes selected for edge deletion in round i, and n is the total number of nodes in the graph. i A new graph G is generated in i+1 :G i+1 =(V i+1 , E i+1 );E i+1 =E i \ {(v i , v j )∈E i |v i ∈V di ∪v j ∈V di};G i and G i+1 As a set of comparison samples, they are input into different encoders (the parameters of the online encoder ε are represented by θ). θ and the target encoder ε with parameters denoted by Φ Φ ). Among them, Gi will be input into the encoder ε θ Get the encoding result H i =ε θ (V i , E i ), and G i+1 will be input into the encoder ε Φ Get the encoding result H i+1 =ε Φ (V i+1 , E i+1 ). The encoding result H i Input to Multi-Layer Perceptron (MLP) p, output online representation Z i According to the online representation Z i And the encoding result H i+1 The comparison loss function of this round is calculated as: L(A i , A i+1 , Z i , H i+1 )=min T (1-α) ∑ j,k T jk C(A i [j],A i+1 [k])+α∑ j,k,j',k' T jk T j' k' | |A i [j, k]-A i+1 [j', k']| | 2 ;stT1=Z i , T T 1=H i+1 , T≥0; where A i is the adjacency matrix of the graph used in the i-th round of training. The matrix T represents all the transportation plans between the two given distributions. C(.,.) represents the distance between different nodes. α is an interpolation parameter with a value of [0, 1], which is used to balance the model's measurement of local and global information of the graph. j, k, j′, k′ represent indices in the loss function, which are used to traverse the relationship between node pairs. (j, k) comes from the graph G i , (j′, k′) from graph G i+1 ; T1 is the result of multiplying the matrix T and the all-1 vector; T T 1 is the result of multiplying the transpose of matrix T by the all-ones vector.

[0101] The model parameters are updated according to the loss function. The encoder ε is updated using the gradient back propagation method. θUpdate, where η represents the learning rate: θ←optimize(θ, η, Ψ θ l(Z i , H i+1 )), where optimize() represents the operation of performing gradient update on the online encoder parameters θ; θ l(Z i , H i+1 ) is Z i and H i+1 The loss function expression is input to perform gradient updates on the parameters θ; the exponential moving average method is used to gradually update the encoder ε Φ , where τ∈[0,1] represents the encoder ε Φ Attenuation parameter for the degree of retention of the original parameters: encoder ε Φ The parameter set Φ←τΦ+(1-τ)θ is as follows; the i-th round of training is completed. In the i+1-th round of training, the graph G i+1 Repeat the above steps to obtain graph G i+2 , and take graph G i+1 and Figure G i+2 is a new set of comparison samples; and so on until the training is completed. After the training process is completed, the trained model ε is used θ To predict the input data and output its node encoding representation. By decoding the encoding result to obtain the integer solution S, the final prediction result of the graph combination optimization problem is output. Assuming that the input graph is G, according to the size of each element in the encoding result H, each vertex is sorted in descending order; according to the constraints of the graph combination optimization problem, from the jth (1≤j≤k, where k is a hyperparameter used to limit the number of searches) vertex v j Start searching and get the solution S j , add it to the solution set Ω={S j} j=1 kThe optimal solution S* for a given graph combinatorial optimization problem is selected from the solution set Ω. Extensive experiments on multiple graph combinatorial optimization problems demonstrate that, compared to existing machine learning methods, the degree-inspired progressive graph contrastive learning (DP-GCL) method can effectively improve model performance and efficiency, achieving superior success on most datasets, particularly large-scale datasets such as RB1000. Compared to contemporary commercial solvers such as Gurobi, this embodiment demonstrates superior performance in solving large-scale graph combinatorial optimization problems, effectively addressing the shortcomings of current graph combinatorial optimization solutions. To further evaluate the performance and generalization of the model proposed in this embodiment, experiments were conducted on graph combinatorial optimization problems of varying scales and in few-shot learning scenarios. The experiments demonstrate that, compared to other machine learning methods, this embodiment can more effectively solve large-scale graph combinatorial optimization problems, with minimal performance degradation as the problem size increases. Even in few-shot learning scenarios, it demonstrates good stability and versatility, possessing strong robustness and the ability to effectively address a wide range of real-world problems.

[0102] In some embodiments, for example, to develop a radio frequency front-end chip for a 5G base station, a 28-nanometer CMOS process is used. The chip contains three main functional modules: a clock distribution network, a high-speed data transmission bus, and an analog radio frequency circuit. In the early design, engineers found that about 35% of the wiring schemes would cause the clock deviation to exceed the design specification, seriously affecting chip performance. In order to improve the wiring success rate, it was decided to use the graph comparison learning method of the present invention to analyze historical design data and guide new wiring optimization. The process is as follows:

[0103] Complete design data for 150 similar RF chips from the past three years was collected. This included: a GDSII-formatted layout file for each design, which recorded detailed routing information on eight metal layers (M1 to M8); process rule files in LEF / DEF format, which defined physical constraints such as minimum line width and minimum spacing; and a design verification report, which indicated the final status of each design. Ninety designs passed all verification tests (marked as successful), while 60 designs had timing violations or signal integrity issues (marked as failed). Each GDSII file was parsed to identify the basic geometric elements in the published drawings. For example, for one successful design, the system identified a 250-micron horizontal trace on the M4 metal layer for the main clock network CLK_MAIN and mapped it to a node in the diagram. The node's attributes included the net identifier CLK_MAIN, metal layer M4, a line width of 0.5 microns, and the clock assignment for the functional module to which it belongs. At the same time, a 180-micron-long parallel wiring of the data bus DATA_BUS

[15] on the M3 layer is identified and mapped to another node. Its attributes include the line identifier DATA_BUS

[15] , the metal level M3, the line width 0.4 microns, and the signal type digital data. Since the two wiring segments are parallel in space with a length of 150 microns and the vertical spacing is only 1.2 microns, the system maps the adjacency relationship between them into an edge in the graph. The attributes of the edge include the parallel length of 150 microns, the vertical spacing of 1.2 microns, and the relative position relationship CLK on the upper layer. After the complete conversion process, 150 designs were converted into 150 attributed graphs, each containing an average of about 8,000 nodes and 12,000 edges, forming a complete attributed graph database. The attributed graph database is scanned to find local structural patterns that are strongly correlated with design failures. First, all single-edge subgraphs are extracted as initial candidates. For example, the system found a specific edge: a parallel relationship between the clock line and the data line with a spacing of less than 2 microns on the M3 layer. Statistical analysis revealed that this type of edge appeared 45 times in 60 failed designs and only 8 times in 90 successful designs. Applying the contrast calculation formula, the contrast score is 45 divided by (8 plus a smoothing factor of 0.000001), which is approximately 5.625, well above the preset threshold of 2.0. A composite attribute-topology expansion was performed on the high-contrast single-edge structure. Regarding topological expansion, the system attempted to add adjacent wiring structures to the clock line-data line pair and found that when a third power line was present near the pair, a three-node subgraph structure was formed. Regarding attribute specialization, the system specialized the originally generalized data line attributes into a specific 16-bit address bus and found that this particular combination was more harmful.After multiple rounds of iterative expansion, the system ultimately identified a high-contrast subgraph instance consisting of four nodes and five edges. This instance describes a critical failure mode: when the 16-bit address bus, main clock line, and power line form a tightly parallel layout on the M3 layer, with a total parallel length exceeding 200 microns, severe crosstalk issues can occur. This instance has a contrast score of 8.7, indicating a strong correlation with design failure.

[0104] The structural eigenvectors for the high-contrast subgraph example above are calculated. First, the multi-scale local degree distribution is calculated: the degrees of the four nodes are 2, 3, 2, and 1, respectively, resulting in a degree distribution histogram of [2, 1, 1, 0] (representing the number of nodes with degrees 1, 2, 3, and 4, respectively). Second-order degrees are also calculated, summing the degrees of each node's neighbors to form a second-order degree distribution histogram. The cyclic basis features are calculated: this subgraph has five edges, four nodes, and one connected component, with a cyclic basis size of 5-4+1=2, indicating the presence of two independent loops. The system finds the three shortest loops, whose lengths are 3, 4, and 5, respectively, which together form the cyclic basis eigenvalues ​​[2, 3, 4, 5]. Statistical attribute distribution: Node attributes include one clock line node, one address bus node, one power line node, and one via node, forming a node attribute histogram of [1, 1, 1, 1]. Edge attributes include two parallel edges (with a spacing of 1.2 microns), two connecting edges, and one crossing edge, with a parallel length distribution of [150 microns, 200 microns]. All features are concatenated and normalized to obtain a 48-dimensional structural feature vector for this instance. This feature vector is compared with existing prototypes in the current taboo prototype library. A similar prototype in the current library describes a parallel structure of clock and data lines. The cosine similarity between its feature vector and the new instance is 0.85, corresponding to a structural distance of 0.15. Because this distance is less than the preset threshold of 0.3, the system determines that the new instance is a variant of the existing prototype. Therefore, instead of creating a new prototype, the existing prototype's statistics are updated: the instance count increases from 15 to 16, and the average contrast score is updated from 6.2 to 6.3. A differential perturbation analysis was performed on the prototype with the highest contrast score in the taboo prototype library. This prototype describes a three-wire parallel structure: clock line, address bus, and power line. Each element in the prototype was systematically removed one by one: after removing the clock line node, the contrast score was recalculated to 2.1, a decrease of 6.6 from the original score of 8.7; after removing the address bus node, the contrast score dropped to 1.8, a decrease of 6.9; after removing the power line node, the contrast score dropped to 7.2, a decrease of only 1.5; and after removing the parallel edge between the clock line and address bus, the contrast score dropped to 0.9, a decrease of 7.8.

[0105] According to the calculated causal contribution scores, the clock line node had a contribution of 6.6, the address bus node had a contribution of 6.9, the power line node had a contribution of 1.5, and the critical parallel edge had a contribution of 7.8. Setting a causal threshold of 5.0, the system identified the clock line node, the address bus node, and the critical parallel edge as the prototype's minimum causal core, while the power line node was excluded due to its low contribution. The resulting minimum causal core consists of two nodes (the clock line and the 16-bit address bus) and one edge (a parallel relationship with a spacing less than 2 microns), accurately capturing the root cause of the design failure. A guided solver was constructed and applied, converting the minimum causal core into a set of logical predicates: IsType(n1, CLK_MAIN), IsType(n2, ADDR_BUS_16), IsLayer(e1, M3), IsParallel(e1, e2), and Distance(e1, e2) < 2.0μm. A penalty function is constructed based on these predicates: when the local routing solution satisfies all conditions simultaneously, a penalty value of 10000 is returned, otherwise 0 is returned.

[0106] During the detailed routing of a new 5G radio chip, a guided solver detected a candidate routing scheme: arranging the main clock network CLK_MAIN and the 16-bit address bus ADDR_BUS[15:0] in parallel on the M3 layer, with a spacing of 1.8 microns and a parallel length of 220 microns. A real-time check revealed that this scheme triggered a taboo mode, with the penalty function returning a high penalty value of 10,000, which dramatically increased the total cost of the scheme. The solver then abandoned this scheme and explored alternatives: moving the address bus to the M4 layer or increasing the routing spacing to 3 microns, ultimately obtaining an optimized routing solution.

[0107] For example, in some cases, the optimized routing scheme is as follows: Routing path of the main clock network CLK_MAIN: M4 layer segment 1: starting point (1000, 2000), end point (3500, 2000), line width 0.5μm; via V3_1: position (3500, 2000), connecting M4 to M3; M3 layer segment 1: starting point (3500, 2000), end point (3500, 4500), line width 0.4μm; via V3_2: position (3500, 4500), connecting M3 to M4; M4 layer segment 2: starting point (3500, 4500), end point (6000, 4500), line width 0.5μm; Routing path of the 16-bit address bus ADDR_BUS[15:0]: ADDR_BUS[0]: M3 layer segment, starting point (1200, 1500), end point (5800, 1500), line width 0.3μm; ADDR_BUS[1]: M3 layer segment, start point (1200, 1200), end point (5800, 1200), line width 0.3μm; ... (specific coordinates of the remaining 14 address lines).

[0108] Layer allocation results: M1 layer: local connection lines, total length 8500μm, utilization rate 75%; M2 layer: power distribution network, total length 12000μm, utilization rate 68%; M3 layer: main data bus wiring layer, total length 15500μm, utilization rate 82%; M4 layer: dedicated clock network layer, total length 6800μm, utilization rate 45%; M5-M8 layers: global signals and long-distance connections, total length 20000μm, utilization rate 60%.

[0109] Safe spacing configuration: Minimum spacing between clock and data lines: 3.2μm (exceeding the 2.0μm safety threshold); minimum spacing between clock and power lines: 2.8μm; spacing between high-speed data lines: 1.5μm; isolation distance between analog and digital signals: 5.0μm. Routing quality improvements: Total routing length: Optimized from the original 68,500μm to 60,800μm, a reduction of 11.2%; Number of vias: Reduced from 2,800 to 3,280, a 17% increase (to avoid high-risk structures); Clock skew: Reduced from a maximum of 450ps to 180ps, a 60% improvement; Signal integrity violations: Reduced from an estimated 15 to 2, an 87% improvement; Avoided taboo structures include: Eliminating eight dangerous parallel sections between the clock and 16-bit address bus on the M3 layer; Adjusting five crossover areas between power and high-speed data lines; and Rerouting 12 critical signal lines to avoid crosstalk hotspots.

[0110] Generated design files: Updated DEF file: contains the precise coordinates and layer information of all routing; SPEF file: extracted parasitic parameters for timing analysis verification; design rule check report: confirms that all geometric constraints are met; routing congestion analysis report: routing density distribution in each area.

[0111] This invention addresses the rigidity and redundancy issues of traditional pattern extraction methods through contrast mining with a composite extension of attributes and topology. It also abstracts massive instances into prototypes through online structural clustering, resolving pattern confusion. Through differential perturbation analysis, it refines the minimal causal core with direct contributions from the prototypes, addressing the causal ambiguity inherent in existing failure mode analysis techniques. Integrating this high-value causal core knowledge base into the optimization solver guides the solution process to proactively avoid known design pitfalls, thereby improving the efficiency and success rate of solving large-scale graph combinatorial optimization problems. Using a dynamic taboo prototype abstraction method based on online structural clustering, the massive, diversely detailed subgraph instances mined are mapped to a feature space using non-neural network structural feature vectors and then clustered online based on distance. Thousands of instances with slightly different topologies but essentially belonging to the same design defect are automatically and efficiently summarized into a manageable number of representative taboo prototypes. This approach directly overcomes the fundamental drawbacks of traditional FSM methods, which generate massively redundant results due to pattern explosion and cannot identify pattern variants due to topological rigidity. Using a minimal causal core extraction method based on differential perturbation analysis, a mechanism was established to calculate the causal contribution of each element within each bloated taboo archetype. By systematically removing each element from the archetype and observing the quantitative changes in its harmfulness (contrast score), causality can be inferred from correlation, accurately isolating the irreducible minimal causal core that truly causes failure. This directly addresses the underlying issue of existing methods, which only discover correlations but fail to locate the causal core, resulting in the mined patterns containing a significant amount of structured noise.

[0112] The preferred embodiments of the present invention are described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the scope of protection of the present invention.

Claims

1. A graph contrast learning method for solving graph combinatorial optimization problems, characterized by: include: An attributed graph database is constructed based on the original VLSI layout database and the associated performance label set. This includes: mapping geometric entities in the VLSI layout into graph nodes, and mapping connections or adjacency relationships between geometric entities into graph edges. Geometric entities include metal trace segments and vias. Physical attributes are added to graph nodes, and geometric attributes are added to graph edges. Physical attributes include metal layer and net ID, and geometric attributes include parallel length and spacing. Scan the attributed graph database and perform contrast analysis based on performance labels to mine high-contrast subgraph instance flows; Receive a stream of high-contrast subgraph instances and dynamically abstract and generate a taboo prototype library through online structural clustering; Combined query of the taboo prototype library and the attributed graph database, by performing differential perturbation analysis on the prototypes in the library, to extract the minimum causal core knowledge base; Loading a minimal causal core knowledge base, constructing a guided combinatorial optimization solver, and using it to solve new graph problems to be optimized, outputting the optimized graph problem solution. This includes: converting each core in the minimal causal core knowledge base into a set of logical predicates and constructing a corresponding penalty function, integrating it into the VLSI detailed wiring solver to construct a guided combinatorial optimization solver. Utilizing this solver to solve new VLSI wiring problems, it detects and avoids known design defect patterns in real time, and outputs a VLSI layout solution that includes optimized wiring path coordinates, metal layer assignments, and via locations. The high contrast sub-graph instance flow is mined, including: Scan the attributed graph database, extract the unilateral subgraph and calculate its support count in the success graph and failure graph to form the initial mining queue; Extract candidate subgraphs from the initial mining queue and perform attribute-topology composite expansion on them to generate an expansion candidate set; For each expansion candidate subgraph in the expansion candidate set, calculate the contrast score by calculating its support count in the success graph and the failure graph; The expansion candidate subgraphs with contrast scores higher than a preset threshold are output to the high-contrast subgraph instance stream, and the remaining expansion candidate subgraphs are added to the initial mining queue.

2. The method according to claim 1, characterized in that Generate an expansion candidate set, including: Traverse the boundary nodes of the candidate subgraph and generate a pure topological extension set by adding new edges and associated nodes; Identify generalized attributes in the candidate subgraph and replace them with specific attribute instances retrieved from the attributed graph database to generate attribute specialization sets; The pure topological extension set and the attribute specialization set are merged, and duplicate subgraphs are removed through graph isomorphism detection to form an extension candidate set.

3. The method according to claim 1, characterized in that Generate a library of taboo prototypes, including: For each high-contrast sub-image instance in the high-contrast sub-image instance stream, its structural feature vector is calculated by using graph theory statistical methods; Compare the structural feature vector with the feature vector of each prototype in the pre-built taboo prototype library to determine the minimum structural distance; Based on the relationship between the minimum structural distance and a preset threshold, it is determined whether the high-contrast subgraph instance belongs to an existing prototype or represents a newly discovered pattern, and the taboo prototype library is updated or expanded accordingly.

4. The method according to claim 3, characterized in that Update or expand the taboo prototype library based on the relationship between the minimum structural distance and the preset threshold, including: Determine the best matching prototype in the taboo prototype library based on the minimum structural distance; If the minimum structural distance is less than a preset threshold, the high-contrast sub-image instance is determined to be a new instance of the best matching prototype, and the statistical information of the best matching prototype is updated; Otherwise, it is determined that the high-contrast subgraph instance represents a newly discovered pattern, and a new prototype is created in the taboo prototype library based on its topological structure and structural feature vector.

5. The method according to claim 3, characterized in that The structural eigenvectors are calculated using graph theory statistical methods, including: Read high contrast sub-image instances, in order: Calculate the multi-scale local degree distribution and obtain the multi-scale distribution histogram; Calculate the cyclic basis characteristics and obtain the cyclic basis eigenvalues; Count the attribute distribution of nodes and edges to obtain the attribute distribution histogram; The multi-scale distribution histogram, cyclic basis eigenvalue and attribute distribution histogram are concatenated and normalized to form a structural feature vector.

6. The method according to claim 5, characterized in that Obtain cyclic basis eigenvalues, including: Determine the size of the cyclic basis of the high-contrast subgraph instance, which is the number of edges minus the number of nodes plus the number of connected components of the high-contrast subgraph instance; Find the lengths of the k shortest independent loops in the high-contrast sub-image instance; where k is a preset parameter; The size of the cyclic basis and the lengths of the k independent loops are combined to form the eigenvalue of the cyclic basis.

7. The method according to claim 1, characterized in that Extract the minimum causal core knowledge base, including: Perform single-element perturbation on the topological structure of each prototype in the taboo prototype library to generate a set of perturbation patterns; For each perturbation pattern in the perturbation pattern set, recalculate its contrast score in the attributed graph database; A causal contribution score is calculated for each removed element in the prototype by comparing the contrast score of each perturbed pattern with the original contrast score of that prototype; Identify all elements whose causal contribution scores are above a preset threshold, constituting a single minimal causal core of the archetype; Aggregate the single minimal causal core of all prototypes to form a minimal causal core knowledge base.

8. The method according to claim 7, characterized in that Calculate a causal contribution score for each removed element in the prototype, including: The causal contribution score is obtained by subtracting the contrast score of the perturbation pattern corresponding to the removed element from the original contrast score of the prototype.

9. The method according to claim 7, characterized in that Recalculate the contrast score of each perturbation pattern in the attributed graph database as follows: Find the graph ID list corresponding to each edge that constitutes the perturbation pattern in the preconfigured edge-graph ID index; where the edge-graph ID index maps each unique attributed edge to the graph ID list containing the edge; Perform an intersection operation on the graph ID list to obtain the final graph ID list containing the perturbation pattern; The differential support count of the perturbation pattern in the success image and the failure image is determined based on the final image ID list, and the contrast score is calculated.

Citation Information

Patent Citations

  • Path planning method, application and device based on knowledge and data combination

    CN117808180A

  • Method and system to generate knowledge graph and sub-graph clusters to perform root cause analysis

    US20230050889A1