An AI model training and version iteration process graphing management method

CN121145922BActive Publication Date: 2026-10-09GUANGDONG INFORMATION NETWORK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511252520.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-03
Publication Date
2026-10-09
Estimated Expiration
2045-09-03

AI Technical Summary

Technical Problem

然而,现有的模型管理方式多依赖于人工命名、线性存档或基础版本控制工具,缺乏结构化的演化关系建模手段,难以准确还原各版本之间的继承链条和依赖路径

Benefits of technology

[0041] This invention constructs a model graph with dual semantics of temporal evolution and causal influence by simultaneously introducing a first type of edge (parameter inheritance path) and a second type of edge (cross-generational parameter dependency). This graph can comprehensively express the explicit inheritance relationship and implicit dependency transmission between different versions. This "dual topological dependency structure" breaks through the limitations of the traditional single-chain model recording method and realizes a structured, computable, and interpretable graph representation of the model's evolutionary history, providing a solid foundation for subsequent analysis and control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121145922B_ABST
    Figure CN121145922B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of model version control and iteration tracking, in particular to a graph-based management method for AI model training and version iteration process, which comprises the following steps: taking model version meta information of each model as a node, constructing a double-topology dependency relationship network based on the node, a first type of edge and a second type of edge; taking the double-topology dependency relationship network as input, propagating node state changes on the second type of edge through a graph neural network based on the topological connectivity of the first type of edge, and outputting a cross-generation abnormal conduction path of an evaluation index abnormality; loading full model files of a previous version of the starting version node from a model storage address, and locking the hyperparameter configuration group to realize automatic version rollback. The application overcomes the problems of difficult positioning of abnormal reasons and non-transparent dependency between versions, and improves the abnormal explainability and intelligent management level of an AI system in a multi-version evolution process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of model version control and iteration tracking technology, and in particular to a graph-based management method for the training and version iteration process of AI models. Background Technology

[0002] With the rapid expansion of the scale and application scenarios of artificial intelligence models, the model iteration speed has accelerated significantly, and model version management faces increasingly complex challenges. In actual development and deployment, different version models often differ in training data, hyperparameter configurations, training strategies, etc., and their performance often exhibits unpredictable fluctuations. However, existing model management methods mostly rely on manual naming, linear archiving, or basic version control tools, lacking structured evolutionary relationship modeling methods, making it difficult to accurately reconstruct the inheritance chains and dependency paths between versions.

[0003] Especially during multiple rounds of parameter tuning, minor adjustments to parameters in certain historical versions may have a non-obvious impact on subsequent, distant versions, leading to abnormal model evaluation metrics. However, due to the lack of effective cross-generational dependency modeling and state propagation mechanisms, the root cause of the anomaly often cannot be accurately traced, severely affecting the model's interpretability and maintenance efficiency. Furthermore, there is currently a lack of an automated version rollback mechanism. Once an abnormal version is deployed, the recovery process still relies on manual analysis and deployment, which is risky and slow to respond. Summary of the Invention

[0004] This invention provides a graph-based management method for AI model training and version iteration processes. This method has structured modeling capabilities, state propagation and traceability capabilities, and automatic rollback capabilities, thereby improving the transparency, stability, and intelligence level of the model lifecycle.

[0005] A graph-based management method for AI model training and version iteration processes includes the following steps:

[0006] S1, Construct a dual topological dependency network: Use the model version meta-information of each model as nodes; generate the first type of edges based on the hyperparameter inheritance rate between adjacent versions of the same model, generate the second type of edges based on the implicit dependency propagation of non-adjacent model versions, and construct a dual topological dependency network based on nodes, the first type of edges, and the second type of edges.

[0007] S2, Cross-generational anomaly path tracing: Taking the dual topological dependency network as input, utilizing the topological connectivity of the first type of edge constraints, the graph neural network propagates node state changes on the second type of edges, and outputs the cross-generational anomaly propagation path of the evaluation index anomaly.

[0008] S3, Topology-driven rollback decision: When an abnormal cross-generational propagation path is detected, backtrack along the first type of edge to the starting version node of the abnormal cross-generational propagation path, load the full model file of the previous version of the starting version node from the model storage address, and lock its hyperparameter configuration group in the no-code training environment to achieve automatic version rollback.

[0009] Optionally, the model version metadata includes the training dataset fingerprint, hyperparameter configuration group, evaluation metric values, and model storage address, wherein:

[0010] Training dataset fingerprint: An identifier obtained by hashing the training data content;

[0011] Hyperparameter configuration group: Stores all key-value pairs of adjustable parameters;

[0012] Evaluation metrics: The numerical values ​​of the model on the validation set;

[0013] Model storage address: A URI path pointing to the model binary file in immutable storage;

[0014] The training dataset fingerprint, hyperparameter configuration group, evaluation metric value, and model save address are created as node attributes for each model version.

[0015] Optionally, the first type of edge generation rules include:

[0016] For adjacent version V n-1 With V n Calculate the hyperparameter inheritance rate. When the hyperparameter inheritance rate is greater than the first threshold, in V... n-1 With V n Create first-class edges with edge weights.

[0017] The calculation of hyperparameter inheritance rate involves comparing the hyperparameter configuration sets of adjacent model versions, counting the number of elements in the intersection and union of the hyperparameter key name sets of the two versions, and obtaining the inheritance rate percentage based on the ratio of the intersection size to the union size.

[0018] Optionally, the second type of edge generation rules include:

[0019] For non-adjacent version V m With V k Define the implicit dependency propagation condition. If the implicit dependency propagation condition is satisfied, then create a second type of edge with edge weights.

[0020] The implicit dependency propagation condition is defined as follows: if there is a hyperparameter adjustment in an earlier version between non-adjacent model versions, and it affects the evaluation index of the subsequent version, there is an implicit dependency propagation path even if there is no direct inheritance relationship between the two. By analyzing the training graph structure through graph neural networks, the sensitivity of the target evaluation index to historical hyperparameters is calculated during backpropagation.

[0021] Optionally, S2 further includes graph neural network initialization:

[0022] The evaluation index value of each node in the dual topology-dependent network is normalized into an initial state vector. An anomaly detection threshold is preset. When the state value of a node exceeds the anomaly detection threshold, it is marked as an initial abnormal node.

[0023] Optionally, S2 includes the following stages:

[0024] Topology constraint propagation phase: State propagation analysis is carried out in the connected subgraph of the model evolution path. Within the scope of the connected subgraph, nodes transmit state information through the parameter influence path. Each node updates its state based on the state of its dependent nodes and the dependency strength.

[0025] Abnormal path backtracking phase: After the topological constraint propagation phase is completed, the state values ​​of all nodes are traversed, the target node that exceeds the propagation anomaly threshold is obtained, that is, the target node is determined to be affected by an anomaly, and the backtracking mechanism is initiated.

[0026] Optionally, the propagation process of the state propagation analysis is carried out in multiple rounds, with each round updating the node state to simulate how abnormal states are propagated structurally and in terms of parameters. The model evolution path is a first-type edge, and the parameter influence path is a second-type edge.

[0027] Optionally, the backtracking process of the backtracking mechanism includes:

[0028] Identify the historical predecessor node that has the greatest impact on the current abnormal node;

[0029] Starting from the predecessor node, trace back along the dependency path to find the source node with the most drastic state change, which is regarded as the starting point of the anomaly propagation;

[0030] Continue tracing back along the evolution path to the previous version of the source node, identify it as a safe and recoverable version, and use it for the rollback operation;

[0031] The backtracking path based on the aforementioned backtracking mechanism is the intergenerational anomaly propagation path.

[0032] Optionally, S3 specifically includes:

[0033] S31, Extract the starting version node of the anomaly propagation path;

[0034] S32, First type of edge reverse backtracking: Starting from the starting version node, traverse the version node sequence in reverse along the first type of edge until the first consecutive version pair that satisfies the first type of edge generation rule conditions, then take the later version node in the consecutive version pair as the rollback target version node.

[0035] S33 accesses the model storage address of the rollback target version node, loads the full model file from immutable storage into the no-code training environment, and completes automatic deployment in the no-code training environment.

[0036] Optionally, S3 further includes parameter locking deployment, specifically including:

[0037] Inject the hyperparameter configuration group of the target version node to be rolled back into the model runtime environment;

[0038] Freeze all hyperparameter modification interfaces;

[0039] Generate a version snapshot including model hash value and parameter lock, and complete the automatic rollback.

[0040] The beneficial effects of this invention are:

[0041] This invention constructs a model graph with dual semantics of temporal evolution and causal influence by simultaneously introducing a first type of edge (parameter inheritance path) and a second type of edge (cross-generational parameter dependency). This graph can comprehensively express the explicit inheritance relationship and implicit dependency transmission between different versions. This "dual topological dependency structure" breaks through the limitations of the traditional single-chain model recording method and realizes a structured, computable, and interpretable graph representation of the model's evolutionary history, providing a solid foundation for subsequent analysis and control.

[0042] This invention applies graph neural networks to version management graphs. By propagating evaluation index states along second-type edges within a connected subgraph defined by first-type edges, it achieves dynamic propagation modeling and precise backtracking of model performance anomalies. Combined with node state mutation detection and edge weight sensitivity judgment, it can automatically identify cross-generational propagation paths and their initial anomaly sources, overcoming the current difficulties of "difficulty in locating anomaly causes and opaque inter-version dependencies," and improving the interpretability and management intelligence of AI systems in the process of multi-version evolution.

[0043] This invention, based on graph-based results, backtracks along the model evolution path to the target version that satisfies a stable inheritance relationship. By reading the model save addresses stored in the nodes, it automatically loads the full model file and locks parameter configurations in a no-code training environment, forming a complete "identification-location-rollback-protection" closed-loop process. This mechanism features high automation, high security, and zero human intervention, enabling rapid, accurate, and traceable version recovery when abnormal fluctuations occur during model training, thus improving the stability and self-healing capabilities of AI systems in production environments. Attached Figure Description

[0044] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only for this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0045] Figure 1 This is a schematic diagram of the method flow according to an embodiment of the present invention;

[0046] Figure 2 This is a schematic diagram illustrating the cross-generational anomaly path tracing in an embodiment of the present invention. Detailed Implementation

[0047] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. For some well-known technologies, those skilled in the art may also use other alternative methods to implement the invention. Moreover, the accompanying drawings are only for more specific description of the embodiments and are not intended to specifically limit the present invention.

[0048] like Figures 1-2 As shown, a graph-based management method for the training and version iteration process of AI models includes the following steps:

[0049] S1, Construct a dual topological dependency network: Use the model version meta-information of each model as nodes; generate the first type of edges based on the hyperparameter inheritance rate between adjacent versions of the same model, generate the second type of edges based on the implicit dependency propagation of non-adjacent model versions, and construct a dual topological dependency network based on nodes, the first type of edges, and the second type of edges.

[0050] S11, Node Generation Rules: Create a unique node for each model version. Node attributes include:

[0051] Training dataset fingerprint: A 128-bit identifier obtained by hashing the training data content, generated using the SHA-3 algorithm to ensure data version traceability.

[0052] Hyperparameter configuration group: All adjustable hyperparameter key-value pairs are stored in JSON format, such as learning rate, batch size, and optimizer, supporting subsequent configuration comparison and inheritance analysis.

[0053] Evaluation metrics: A set of performance metrics for the model on the validation set, including accuracy, recall, and F1 score, used to evaluate the model's training effectiveness.

[0054] Model storage address: An immutable storage URI path pointing to the model binary file, used for historical version loading and automatic rollback operations.

[0055] S12, First type of edge generation rule: For adjacent version nodes V n-1 With V n Calculate the hyperparameter inheritance rate:

[0056]

[0057] When condition: R inherit ≥80%; then in V n-1 With V n Create weighted first-class edges between them, with edge weights of: W1 = R inherit ;

[0058] Among them, P n-1 P n Version V n-1 and V n The hyperparameter configuration set, |·| represents the number of elements in the set, R inherit The parameter inheritance rate is used to measure the hyperparameter similarity between two versions. W1 represents the weight of the first type of edge, which is used for backtracking priority path scoring. When the inheritance rate is below the 80% threshold, the instability of version evolution increases significantly.

[0059] The calculation of hyperparameter inheritance rate is based on the principle of measuring set similarity, using the Jaccard similarity coefficient as a quantification method. The Jaccard coefficient is used to evaluate the degree of overlap between two sets, defined as the ratio of the size of the intersection to the size of the union of the two sets. During model iteration, different versions of hyperparameter configurations may be reused, adjusted, or added. By calculating the intersection and union of the hyperparameter key name sets of two adjacent versions, the inheritance relationship in hyperparameter design between the two versions can be reflected. If the intersection ratio is high, it indicates that the current version reuses a large amount of parameter configuration from the previous version, belonging to "lightweight iteration"; if the intersection ratio is low, it indicates that the current version is more likely to be a redesigned structure or parameters, belonging to "reconstructive iteration". Therefore, using the Jaccard similarity coefficient (i.e., intersection divided by union) as the basis for hyperparameter inheritance rate can objectively reflect the degree of inheritance between versions, thus being used for the generation of the first type of edge and the setting of edge weights in the topological network.

[0060] Hyperparameter inheritance rate R inherit The calculation is explained in detail below:

[0061] Given two adjacent model versions:

[0062] V n-1 In the previous version, its hyperparameter configuration set was P. n-1 ;

[0063] V n The current version has a hyperparameter configuration set of P. n ;

[0064] Extract the set of hyperparameter key names from two versions (ignoring values, only comparing whether the key names are the same), for example:

[0065] P n-1 ={"Learning Rate":0.01,"Batch Size":64,"Optimizer":"Adam"};

[0066] P n ={"Learning Rate":0.01,"Optimizer":"Adam","Weight Decay":0.0001};

[0067] Calculate the number of elements in the intersection and union of sets:

[0068] Intersection P n-1 ∩P n {"learning rate", "optimizer"}, size 2;

[0069] Union of P n -1∪P n = {"Learning rate", "Batch size", "Optimizer", "Weight decay"}, size is 4;

[0070] Calculate inheritance rate:

[0071]

[0072] Intersection: Represents hyperparameters that are retained in both versions;

[0073] Union: Represents all hyperparameter keys that have appeared in both versions.

[0074] S13, Second type of edge generation rule: For non-adjacent version nodes V m With V k (where k>m+1), if the following implicit dependency transmission condition is satisfied:

[0075]

[0076] Then, a second type of edge is created, with the edge weight as follows:

[0077] Among them, evaluation indicators k Version V k The target performance metrics (such as accuracy), hyperparameters m Version V m A specific hyperparameter (such as the learning rate) is used, where δ represents the dependency sensitivity threshold, controlling whether to generate second-type edges. The gradient distribution of the influence of all historical versions on the current metric is extracted from the model training log, and its mean μj and standard deviation σj are calculated. δ is set to μj + kj·σj, with kj = 1 to 1.5, to filter out low-impact paths and retain only upper-layer strong dependency edges. This indicates the dependence strength of the hyperparameter change on the metric of subsequent versions. The dependence strength is automatically estimated through the backpropagation mechanism of a graph neural network (GNN). Essentially, it calculates the gradient impact of hyperparameter perturbations in historical version nodes on the evaluation metric of subsequent versions. The following is the specific process of estimating the dependence strength through GNN backpropagation:

[0078] I. GNN Graph Structure Modeling: Using the model version as a node, construct a directed graph that includes first-class edges (direct inheritance) and second-class edge candidate paths (from historical versions to subsequent versions);

[0079] Node attribute input: Each node input contains attributes such as the training data fingerprint, hyperparameter configuration group, and evaluation metric value for that version;

[0080] Initial feature encoding: Hyperparameters and indices are converted into low-dimensional feature vector representations through an embedding layer.

[0081] II. Forward Propagation: Propagation Modeling of Evaluation Indicators

[0082] Graph Neural Network Propagation: Information is propagated through several rounds of GNN to capture the dependency paths and feature combination methods between versions;

[0083] Target output node: At the target node, such as version Vk, the accuracy of the predicted evaluation metric is output through MLP;

[0084] Loss function definition: The mean squared error (MSE) is used to calculate the error between the predicted value and the actual evaluation index.

[0085] III. Backpropagation: Dependency Strength Estimation:

[0086] Perform backpropagation: Starting from the target evaluation index node, propagate the error back to all its connected predecessor nodes (i.e., candidate historical versions) in the network;

[0087] Extracting gradient information: During backpropagation, the system automatically calculates the partial derivative (gradient) of the hyperparameter input with respect to the target index output for each historical version;

[0088] Dependency strength calculation: Take its absolute value as the edge weight W2 of the second type of edge. If the value is greater than the preset sensitivity threshold δ, it is considered that the historical hyperparameter has an implicit transmission effect on the current index, thus generating the second type of edge.

[0089] Dual-topology dependency network: When constructing the evolutionary graph of the model training process, the "model version" is used as the core node, and two types of edges with different semantics are combined to form a dependency network with dual-topology features.

[0090] Node Definition: Each model version is represented as a node in the graph. Each node contains complete model metadata, including the training dataset fingerprint, hyperparameter configuration groups, evaluation metric values, and model storage location. This information collectively identifies the version's input conditions, training strategy, and performance results, serving as the basic unit for subsequent analysis.

[0091] Type I edges: Explicit topological paths: Type I edges connect adjacent versions that have a direct temporal relationship and whose hyperparameter inheritance rate exceeds a threshold, forming the "version evolution main path." These edges constitute the temporally continuous topological structure of the network, used to describe the backbone path of the model's natural evolution and provide a backtracking link for rollback operations. The topology of Type I edges is a linearly traceable, directional, and low-branching main structure.

[0092] Type II edges: Implicit topological paths: Type II edges connect nodes that have statistical dependencies between non-adjacent versions, representing the "cross-generational impact of hyperparameters on subsequent performance." These edges constitute a high-dimensional causal topology in the network, used to capture historical change paths that appear disconnected but actually have performance propagation effects. The topology of Type II edges is a sparse, nonlinear, and highly correlated complementary structure.

[0093] Dual topology fusion features:

[0094] The first type of edge reflects the time-series dependency of the model iteration;

[0095] The second type of edge supplements the parameter influence dependency in model evolution;

[0096] Together, they form a dual-topological dependency network that supports both sequential evolutionary analysis and cross-generational dependency backtracking, enabling graph-based modeling of the global state during model training.

[0097] S2, Cross-generational anomaly path tracing: Taking the dual topological dependency network as input, utilizing the topological connectivity of the first type of edge constraints, the graph neural network propagates node state changes on the second type of edges, and outputs the cross-generational anomaly propagation path of the evaluation index anomaly.

[0098] S21, Graph Neural Network Initialization: Normalize the evaluation metric value of each version node in the dual-topology dependency network and initialize it as a state vector: Wherein, if the state of a node satisfies: This version is considered to be in an initial abnormal state, where, Let represent the initial state value of the i-th node, the normalized index score, and β be the anomaly detection threshold (set to 0.8). β is used to identify version nodes with significantly abnormal evaluation index values ​​in the initial state. Typically, the evaluation index is normalized to the [0,1] interval, and a high value close to the upper limit is selected as the anomaly detection threshold. In this invention, if key indicators such as model accuracy and F1 score are significantly higher than the normal distribution average, such as an average accuracy of around 0.7, there may be problems such as abnormal overfitting or data leakage. Therefore, setting 0.8 as the detection starting point helps to mark suspicious version nodes early and intervene in abnormal path tracing.

[0099] The evaluation index value of each version node in the dual topology dependency network is normalized. The main purpose is to map the original scores of different model versions on different evaluation indexes to a standardized interval [0,1] so that it can be used as the initial state vector for propagation in the graph neural network, thus constructing a unified propagation starting point.

[0100] The specific normalization process is as follows:

[0101] Determine the normalization metrics: Select the main evaluation metrics (precision, recall, F1 score) as the normalization objects. If there are multiple metrics, a weighted average or priority of the main metrics can be used.

[0102] Calculate the minimum and maximum values ​​of all version nodes: Traverse all version nodes in the entire graph, calculate the minimum value (Min) and maximum value (Max) of the evaluation metric, and form a normalized interval.

[0103] For each version node's metric value x, calculate its normalized state value S. (0) :

[0104] If Max = Min, then the state of all nodes is uniformly set to 0.5 or other constants to avoid division by zero errors.

[0105] The result is used as the initial state input to the normalized value S of each node in the graph neural network. (0)∈[0,1] is used as the starting state for its graph propagation, and is used for subsequent anomaly detection and propagation analysis.

[0106] S22, Topological Constraint Propagation: In a connected subgraph composed of edges of type I, a graph neural network is used for state propagation, updating only along edges of type II. The propagation formula is as follows:

[0107]

[0108] in, Let j represent the state value of the i-th node after propagation at layer k, and j represent the nodes adjacent to the target node i in the graph neural network, belonging to the second type of edge neighbor set of node i. Inside, Let σ(·) represent the state value of node j after propagation in the (k-1)th round of the graph neural network (the node state in the previous round), and let σ(·) represent the activation function ReLU. This represents the set of all second-type edge neighbors of node i within the same first-type edge-connected subgraph. D represents the edge weight of the second type of edge from node j to i, i.e., the absolute value of the partial derivative. ii D jj These are the diagonal elements of the degree matrix for nodes i and j, used for normalization to prevent highly connected nodes from dominating the propagation process.

[0109] The core idea of ​​the topological constraint propagation scheme is to use a graph neural network to perform anomalous state propagation analysis only within a subgraph consisting of version evolution paths (i.e., first-type edges), but the propagation itself only occurs along parameter influence paths (i.e., second-type edges). This design achieves a combination of structural boundary control and influence path modeling.

[0110] Specifically, we first determine which model versions belong to the same continuous evolutionary branch. These versions form a structurally connected subgraph. Then, we analyze the cross-generational influence relationships between parameters only within this subgraph. During propagation, the state of each version node does not evolve independently but is influenced by its second-type edge neighbors. The state is dynamically updated based on the degree of anomaly of these neighbors.

[0111] The propagation process is iterated in multiple rounds. After each round of updates, the state value of the node tends to reflect the cumulative impact of abnormal propagation in the dependent paths it is connected to. In order to avoid certain highly connected nodes having a dominant influence on the propagation results, the incoming state is structurally normalized to make the propagation more balanced and more interpretable.

[0112] This ensures that anomaly propagation analysis is conducted only within subgraphs with evolutionary continuity, while also uncovering deep parameter dependencies that lead to anomaly propagation paths, providing dynamic and accurate state data for subsequent anomaly backtracking.

[0113] S23, Abnormal Path Backtracking: When a target version node v t After propagation, the following conditions are met: If an anomaly propagation occurs, it is considered to have occurred across generations, triggering a path tracing process. γ is the propagation anomaly threshold, used to determine whether a node is significantly affected by an anomaly state after multiple rounds of state propagation. In this invention, the state propagation path is limited to the second-type edge neighbors within the first-type edge connected domain. If the state of the propagation endpoint node ultimately exceeds 0.85, it indicates that it is significantly driven by historical implicit dependency paths, requiring backtracking diagnosis. This value is usually set higher than β to ensure the "purity" of the propagation impact and avoid misjudging minor state fluctuations caused by local disturbances, as detailed below:

[0114] S231. Find the predecessor version node with the greatest contribution: Select the predecessor node v that contributes the most to the target node in the last round of propagation. p ;

[0115] S232. Locating the starting version node v of the state mutation source s Backtracking along the second type of edge to the first state mutation node satisfies: Here, k is used to identify key nodes that change significantly during the propagation process, η is the number of propagation rounds, and η represents the threshold for determining state abrupt changes. This indicates that in the k-th round of state propagation, node v s The state value (i.e., the output of the k-th layer of the graph neural network), This indicates that in the k-th round of state propagation, node v s The state value is set to 0.3. η is used to detect whether a node changes drastically in two consecutive rounds of state propagation in order to identify the "state mutation source node". In this invention, the propagation of the graph neural network has a certain smoothness. The state value of a normal node does not change much between adjacent propagation rounds, usually <0.1, while key nodes on the abnormal propagation path will show significant transitions. Setting the mutation threshold to 0.3 can effectively distinguish key turning points in the propagation path and support the accuracy requirements of abnormal backtracking and localization.

[0116] S233. Generate an anomaly propagation path: v s →…→v p →v t And further backtrack along the first type of edge to v s The previous version, as a recoverable rollback version. v t This indicates that an abnormal target version node has been detected. pIndicates the propagation of v t The most affected precursor node, v s This represents the starting version node where the first state change occurs during the backtracking process.

[0117] All state propagation occurs only within the first type of edge-connected subgraph; the second type of edges serve as state propagation paths, used only to capture parameter influence relationships, and do not expand the connected domain; the state values ​​and state changes of abnormal nodes are automatically generated by GNN propagation, without the need for manual specification.

[0118] The core objective of the anomaly path backtracking scheme is to automatically analyze the source of the anomaly and generate an interpretable propagation path when a model version is identified as potentially affected by anomalies after state propagation, providing a basis for rollback. Its principle is based on the propagation characteristics of states between nodes in a graph neural network. When a version node becomes anomalous after propagation, it first determines which historical versions influenced it through dependency paths (second-type edges), and identifies the predecessor node with the greatest influence in these paths, considering it the main entry point of the anomaly. Next, it traces the dependency path backward from this key predecessor node to find nodes where the state underwent drastic changes during propagation, considering these nodes as the earliest source of anomaly propagation. By identifying this "state mutation source," the starting version of the anomaly can be located. Finally, based on the first-type edges, i.e., the model evolution path, it backtracks from the mutation source node to its previous stable version, which is considered a reliable state unaffected by the anomaly and can be used as the rollback target.

[0119] The occurrence of anomalous states is rarely isolated, but rather the result of the accumulation of parameter dependencies during the evolutionary process. By combining topological structure with state evolution dynamics, the intergenerational transmission chain of anomalous states can be traced, enabling automatic source tracing and precise rollback.

[0120] S3, Topology-driven rollback decision: When an abnormal cross-generational propagation path is detected, backtrack along the first type of edge to the starting version node of the abnormal cross-generational propagation path, load the full model file of the previous version of the starting version node from the model storage address, and lock its hyperparameter configuration group in the no-code training environment to achieve automatic version rollback.

[0121] S31, using the starting version node v of the exception propagation path output by S2. s , i.e., a state mutation node, is regarded as the source version of the propagation of the abnormal chain.

[0122] S32, Rollback target location: from source node v s Backtracking along the first type of edge (model evolution path), traversing the sequence of model version nodes, during the backtracking process, for each pair of adjacent version nodes, checking the hyperparameter inheritance relationship between versions in turn, if a certain consecutive version pair (v i ,vi-1 The hyperparameter inheritance rate of R satisfies the condition defined and calculated in S1: inherit If ≥80%, then this version will be paired with the later version node v. i As the target version node for rollback v target This operation ensures that the selected rollback version has configuration continuity that is highly consistent with the current model, thereby preserving training results to the greatest extent and isolating anomalies.

[0123] S33, Access the rollback target version node v target The model is stored at a specific location; the complete model file is loaded from immutable object storage and deployed to the no-code training and runtime environment.

[0124] Access the model's storage address URI;

[0125] Load the model binary file;

[0126] Create a no-code container environment;

[0127] Deploy the model to the runtime environment and complete the loading process.

[0128] The deployment process eliminates the need for manual script writing, model structure configuration, or training process reconstruction. It directly calls the model files and parameter configurations saved in the rollback target node, enabling one-click loading and automatic execution, thus improving recovery response speed. This is particularly suitable for the time-sensitive requirements of exception handling in online business scenarios. The no-code environment centrally controls deployment logic and resource allocation, avoiding configuration errors or version inconsistencies caused by manual developer operations. The no-code platform can directly parse the node structure in the graph, automatically identifying fields such as model storage address and hyperparameter configuration groups, achieving structured integration with graph information. This allows "graph-driven rollback paths" to execute seamlessly within the platform, supporting automatic closed-loop graph management.

[0129] Loading a complete model file from immutable object storage means retrieving a complete model file from a read-only, unmodifiable data storage location to restore the model's state. "Immutable object storage" refers to a storage mechanism where data, once written, cannot be modified or deleted. This is commonly found in Alibaba Cloud OSS and AWS S3's "version locking" or "archive storage" modes, used to ensure the integrity and tamper-proof nature of historical model files. "Loading" means reading model data (including model structure and weights) from this storage system. "Full model file" refers to the complete file of the entire model, not just a part (such as parameters or configuration only), including all trained weights, network structure, and other content.

[0130] S34, Parameter Locking Deployment: After model loading is complete, perform a three-step parameter locking operation on the current model environment to ensure that the deployment process cannot be tampered with:

[0131] Parameter injection: v target The node's hyperparameter configuration group is injected into the runtime environment;

[0132] Interface freeze: Close all runtime parameter modification entry points and disable parameter tuning APIs;

[0133] Version snapshot generation: Generate snapshot records by combining model hash values ​​and parameter states, serving as complete traceable credentials for rollback versions.

[0134] This invention encompasses any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of this invention. To provide the public with a thorough understanding of this invention, specific details are described in detail in the following preferred embodiments; however, those skilled in the art will fully understand the invention even without these details. Furthermore, to avoid unnecessary misunderstanding of the essence of this invention, well-known methods, processes, procedures, components, and circuits are not described in detail.

[0135] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A graph-based management method for the training and version iteration process of AI models, characterized in that, Includes the following steps: S1, use the model version meta-information of each model as nodes; generate the first type of edge based on the hyperparameter inheritance rate between adjacent model versions, generate the second type of edge based on the implicit dependency propagation of non-adjacent model versions, and construct a dual topological dependency network based on nodes, the first type of edge and the second type of edge. The second type of edge generation rules include: For non-adjacent versions and Define the implicit dependency propagation condition. If the implicit dependency propagation condition is satisfied, then create a second type of edge with edge weights. The implicit dependency propagation condition is defined as follows: if there is a hyperparameter adjustment in an earlier version between non-adjacent model versions, and it affects the evaluation index of the subsequent version, even if there is no direct inheritance relationship between the two, there is still an implicit dependency propagation path. The graph structure is analyzed by graph neural network, and the sensitivity of the target evaluation index to historical hyperparameters is calculated during backpropagation. S2, taking the dual topological dependency network as input, utilizing the topological connectivity of the first type of edge constraints, propagating node state changes on the second type of edges through a graph neural network, and outputting the cross-generational anomaly propagation path of the evaluation index anomaly; Specifically, it includes the following stages: Topology constraint propagation phase: State propagation analysis is carried out in the connected subgraph of the model evolution path. Within the scope of the connected subgraph, nodes transmit state information through the parameter influence path. Each node updates its state based on the state of its dependent nodes and the dependency strength. Abnormal path backtracking phase: After the topological constraint propagation phase is completed, the state values ​​of all nodes are traversed to obtain the target node that exceeds the propagation anomaly threshold. That is, the target node is determined to be affected by an anomaly, and the backtracking mechanism is initiated. The propagation process of the state propagation analysis is carried out in multiple rounds. In each round, the node state is updated to simulate the propagation of abnormal states in terms of structure and parameters. The model evolution path is a first type of edge, and the parameter influence path is a second type of edge. S3, when a cross-generational anomaly propagation path is detected, backtrack along the first type of edge to the starting version node of the cross-generational anomaly propagation path, load the full model file of the previous version of the starting version node from the model storage address, and lock its hyperparameter configuration group in the no-code training environment to achieve automatic version rollback.

2. The graph-based management method for AI model training and version iteration processes according to claim 1, characterized in that, The model version metadata includes the training dataset fingerprint, hyperparameter configuration group, evaluation metric values, and model storage address, among which: Training dataset fingerprint: An identifier obtained by hashing the training data content; Hyperparameter configuration group: Stores all key-value pairs of adjustable parameters; Evaluation metrics: The numerical values ​​of the model on the validation set; Model storage address: A URI path pointing to the model binary file in immutable storage; The training dataset fingerprint, hyperparameter configuration group, evaluation metric value, and model save address are created as node attributes for each model version.

3. The graph-based management method for AI model training and version iteration processes according to claim 1, characterized in that, The first type of edge generation rules include: For adjacent versions and Calculate the hyperparameter inheritance rate. When the hyperparameter inheritance rate is greater than the first threshold, in and Create first-class edges with edge weights. The calculation of hyperparameter inheritance rate involves comparing the hyperparameter configuration sets of adjacent model versions, counting the number of elements in the intersection and union of the hyperparameter key name sets of the two versions, and obtaining the inheritance rate percentage based on the ratio of the intersection size to the union size.

4. The graph-based management method for AI model training and version iteration processes according to claim 1, characterized in that, S2 also includes graph neural network initialization: The evaluation index value of each node in the dual topology-dependent network is normalized into an initial state vector. An anomaly detection threshold is preset. When the state value of a node exceeds the anomaly detection threshold, it is marked as an initial abnormal node.

5. The graph-based management method for AI model training and version iteration processes according to claim 1, characterized in that, The backtracking process of the backtracking mechanism includes: Identify the historical predecessor node that has the greatest impact on the current abnormal node; Starting from the predecessor node, trace back along the dependency path to find the source node with the most drastic state change, which is regarded as the starting point of the anomaly propagation; Continue tracing back along the evolution path to the previous version of the source node, identify it as a safe and recoverable version, and use it for the rollback operation; The backtracking path based on the aforementioned backtracking mechanism is the intergenerational anomaly propagation path.

6. The graph-based management method for AI model training and version iteration process according to claim 1, characterized in that, S3 specifically includes: S31, Extract the starting version node of the anomaly propagation path; S32, First type of edge reverse backtracking: Starting from the starting version node, traverse the version node sequence in reverse along the first type of edge until the first consecutive version pair that satisfies the first type of edge generation rule conditions, then take the later version node in the consecutive version pair as the rollback target version node. S33 accesses the model storage address of the rollback target version node, loads the full model file from immutable storage into the no-code training environment, and completes automatic deployment in the no-code training environment.

7. The graph-based management method for AI model training and version iteration process according to claim 6, characterized in that, The S3 also includes parameter locking deployment, specifically including: Inject the hyperparameter configuration group of the target version node to be rolled back into the model runtime environment; Freeze all hyperparameter modification interfaces; Generate a version snapshot including model hash value and parameter lock, and complete the automatic rollback.

Citation Information

Patent Citations

  • Intelligent document updating processing method and system

    CN120471039A

  • Algorithm plug-in collaborative development method and system based on multi-version control

    CN120491956A