Deep exploration AI reasoning method and system
By generating dynamic calculation graphs and combining adaptive pruning with semantic weights and permission levels, the problem of low synergy between inference depth and resource efficiency is solved, and efficient and stable execution in complex tasks is achieved.
Patent Information
- Application Number
- CN202510459650.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-29
AI Technical Summary
In the existing intelligent in-depth exploration AI inference methods, the synergy between inference depth and resource efficiency is low, the accuracy of dynamic calculation graph complexity evaluation is insufficient, and the coupling of permission management and complexity control is weak, resulting in unbalanced resource allocation and incomplete inference.
By generating dynamic calculation maps, risk assessment and optimization processing are performed, adaptive pruning is performed in combination with semantic weights, authority levels and reinforcement learning, optimized calculation maps are generated, and elastic resource scheduling is implemented through multi-objective models and real-time monitoring to ensure efficient execution of inference tasks.
It realizes the coordination of resource allocation optimization and permission control in complex logical reasoning tasks, ensuring efficient and stable execution of inference tasks, avoiding resource waste and system crashes.
Smart Images

Figure CN120387515A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a deep exploration AI inference method and system. Background Art
[0002] The intelligent deep exploration AI inference method is a dedicated artificial intelligence inference method for complex logical inference tasks. Its core function is to dynamically generate a computational graph containing complex logics such as loop iteration and multi-branch structure, and execute inference tasks based on a real-time resource scheduling and permission grading mechanism. It is applied to high-complexity inference scenarios such as natural language processing, image recognition, and multi-modal data analysis.
[0003] The intelligent deep exploration AI inference method can generate a dynamic computational graph based on the input inference task of the user, analyze the dependency relationship and data flow between nodes, and optimize computing resources through a dynamic scheduling algorithm. For example, a user permission grading strategy is introduced to differentially constrain the execution complexity of tasks at different permission levels.
[0004] However, in the dynamically generated inference computational graph, the evaluation accuracy of the dynamic computational graph complexity is insufficient, and the coupling between permission management and complexity control is weak, resulting in low synergy between inference depth and resource efficiency. Summary of the Invention
[0005] This application provides a deep exploration AI inference method and system to solve the problem of low synergy between inference depth and resource efficiency.
[0006] In a first aspect, this application provides a deep exploration AI inference method, including:
[0007] Generating a dynamic computational graph based on the inference task to be performed, where the dynamic computational graph is used to represent the logical structure and data flow of the inference task to be performed;
[0008] Performing a risk assessment on the dynamic computational graph to generate a risk prediction result, where the risk assessment at least includes complexity assessment, memory occupancy assessment, and data transfer volume assessment;
[0009] Performing an optimization process on the dynamic computational graph based on the risk prediction result to generate an optimized computational graph, where the optimized computational graph is used to constrain the inference computational depth and the range of inference resource allocation.
[0010] In some feasible embodiments, the generating the dynamic computational graph includes:
[0011] Analyzing the logical structure and data flow of the inference task to be performed to generate an initial computational graph, where the initial computational graph includes nodes, and the nodes at least include one or more combinations of data nodes, logical nodes, and computational nodes;
[0012] Based on the semantic embedding model, perform semantic weight calculation on the nodes and edges to generate a weighted dynamic computational graph, where the edges are generated by the nodes;
[0013] Perform topological optimization processing on the weighted dynamic computational graph to obtain the dynamic computational graph.
[0014] In some feasible embodiments, the performing semantic weight calculation includes:
[0015] Obtain the number of text segments of the task to be inferred;
[0016] Based on the task to be inferred, match the text segments of the nodes;
[0017] Based on the semantic embedding model, perform semantic weight calculation on the nodes and edges through the number of text segments, text segments, and normalization function to generate node semantic weights, where the node semantic weights include one or more combinations of the data nodes, logical nodes, and computational nodes.
[0018] In some feasible embodiments, the performing topological optimization processing on the weighted dynamic computational graph to obtain the dynamic computational graph includes:
[0019] Perform loop termination probability analysis on the loop control nodes to generate a loop termination probability distribution, where the loop control nodes are the nodes in the weighted dynamic computational graph for managing and executing loop iteration logic;
[0020] Based on the loop termination probability distribution and the reverse edge weight adjustment algorithm, calculate the dynamic weight value, where the dynamic weight value is the weight value of the reverse logical dependency edge of the loop control nodes;
[0021] Update the topological structure of the weighted dynamic computational graph according to the dynamic weight value to obtain the dynamic computational graph.
[0022] In some feasible embodiments, the performing risk assessment on the dynamic computational graph to generate a risk prediction result includes:
[0023] Perform time-consuming calculation on the nodes to generate a node time-consuming vector;
[0024] Based on the node time-consuming vector, adjust the weight coefficient through a reinforcement learning algorithm to construct a risk assessment model, where the weight coefficient is the weight coefficient of the risk assessment model;
[0025] Through the risk assessment model, perform iteration number prediction on the loop subgraphs in the dynamic computational graph to generate a risk prediction result.
[0026] In some feasible embodiments, the constructing a risk assessment model includes:
[0027] Obtain historical execution data,
[0028] Based on the historical execution data, construct a reinforcement learning state matrix;
[0029] Update the weight coefficients through a reinforcement learning algorithm and the update parameters, and construct a risk assessment model based on the updated weight coefficients.
[0030] In some feasible embodiments, the performing optimization processing on the dynamic computation graph based on the risk prediction result to generate an optimized computation graph includes:
[0031] Obtain the user privilege level;
[0032] Calculate the average complexity of the dynamic computation graph;
[0033] Calculate the survival probability of the nodes according to the user privilege level and the average complexity;
[0034] Generate a pruning threshold based on the survival probability;
[0035] Perform path sorting on the nodes below the pruning threshold to generate a sorted list;
[0036] Perform pruning processing on the dynamic computation graph according to the sorted list;
[0037] Obtain a lightweight approximation function and insert the lightweight approximation function at the breakpoint, where the breakpoint is the breakpoint position in the dynamically computed graph after pruning.
[0038] In some feasible embodiments, the obtaining the lightweight approximation function includes:
[0039] Perform function type processing on the sorted list to obtain an initial function structure, where the initial function structure is a polynomial approximation function or a rule simplification function;
[0040] Perform function adaptation processing on the survival probability to obtain a parameter configuration, where the parameter configuration is a low-order fast approximation parameter or a high-hit-rate cache index;
[0041] Perform function constraint processing based on the average complexity to obtain a function combination, where the function combination includes an error-tolerant interpolation function or a stable-hit cache query function;
[0042] Perform fusion processing on the initial function structure, the parameter configuration, and the function combination to generate a lightweight approximation function.
[0043] In some feasible embodiments, the method further includes:
[0044] Construct a multi-objective optimization model based on the optimized computational graph;
[0045] Solve the multi-objective optimization model by the Lagrange multiplier method to generate a resource mapping table;
[0046] Perform resource allocation on the optimized computational graph based on the resource mapping table and obtain the resource consumption status;
[0047] Preset a trigger condition;
[0048] If the resource consumption status meets the trigger condition, trigger a hierarchical resource recovery mechanism, and the hierarchical resource recovery mechanism includes forced suspension processing and execution precision degradation processing.
[0049] In a second aspect, the present application also provides a deep exploration AI inference system, including:
[0050] A dynamic graph construction module, configured to generate a dynamic computational graph based on a task to be inferred, where the dynamic computational graph is used to represent the logical structure and data flow of the task to be inferred;
[0051] A risk prediction module, configured to perform risk assessment on the dynamic computational graph to generate a risk prediction result, and the risk assessment includes at least complexity assessment, memory occupancy assessment, and data transmission volume assessment;
[0052] A dynamic graph optimization module, configured to perform optimization processing on the dynamic computational graph based on the risk prediction result to generate an optimized computational graph, where the optimized computational graph is used to constrain the depth of inference calculation and the range of inference resource allocation.
[0053] As can be seen from the above technical solutions, the present application provides a deep exploration AI inference method and system. The method includes: generating a dynamic computational graph based on a task to be inferred, where the dynamic computational graph is used to represent the logical structure and data flow of the task to be inferred, then performing risk assessment on the dynamic computational graph to generate a risk prediction result, and the risk assessment includes at least complexity assessment, memory occupancy assessment, and data transmission volume assessment; performing optimization processing on the dynamic computational graph based on the risk prediction result to generate an optimized computational graph, where the optimized computational graph is used to constrain the depth of inference calculation and the range of inference resource allocation. The method converts natural language input into a semantically weighted dynamic computational graph, identifies resource risks through node-level complexity assessment; combines permission levels with reinforcement learning for adaptive pruning to generate an optimized computational graph; and implements elastic resource scheduling based on a multi-objective model and real-time monitoring to solve the problem of low coordination between inference depth and resource efficiency. Description of the Drawings
[0054] To more clearly illustrate the technical solutions of this application, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0055] Figure 1 It is a schematic flowchart of the deep exploration AI inference method provided by the embodiment of this application;
[0056] Figure 2 It is a schematic flowchart of the dynamic computation graph generation provided by the embodiment of this application;
[0057] Figure 3 It is a schematic flowchart of the risk assessment model construction provided by the embodiment of this application;
[0058] Figure 4 It is a schematic structural diagram of the deep exploration AI inference system provided by the embodiment of this application. Detailed implementation manners
[0059] The following will describe the embodiments in detail, and the examples are shown in the accompanying drawings. When the following description involves the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following embodiments do not represent all the implementation manners consistent with this application. They are only examples of the systems and methods consistent with some aspects of this application detailed in the claims.
[0060] The intelligent deep exploration AI inference method generates a logical structure by parsing user input, analyzes the node dependency relationship, and combines a dynamic scheduling algorithm to achieve optimal resource allocation and improve the processing efficiency of complex tasks.
[0061] Exemplarily, taking the legal consultation scenario as an example, when processing a request to analyze the validity of the contract liquidated damages clause according to the law, a computation graph containing multi-layer nested logic needs to be constructed. First, parse the legal provisions to generate data nodes, that is, the liquidated damages ratio and the actual loss value, then establish logical nodes, that is, the rule for judging excessive liquidated damages, and then generate computation nodes, that is, the formula for the liquidated damages adjustment range. Such a computation graph contains recursive loops and multi-condition branches, and the resource allocation needs to be dynamically adjusted to ensure the integrity of the inference.
[0062] In this scenario, the decoupling of permission management and complexity control causes unbalanced resource allocation. The critical computation nodes of high-privilege users are interrupted due to insufficient memory, while low-privilege users can occupy resources by constructing complex logic. The root cause of these problems lies in the lack of in-depth understanding and adaptive processing ability of the dynamic characteristics of the computation graph, relying on preset rules and being unable to perceive the semantic weight and resource consumption trend in real time, resulting in the dilemma of coordinated optimization of inference depth and resource efficiency.
[0063] In some embodiments, the permission level is dynamically bound to the complexity of the computational graph. For example, a permission complexity mapping table is established, stipulating that certain users can process computational graphs with a complexity index ≤ 5. Although this method enhances the pertinence, it relies on predefined complexity metrics and cannot reflect the dynamic changes of the computational graph in real time.
[0064] When the update of legal provisions leads to a change in the inference logic, the predefined complexity index may become invalid, triggering a new imbalance in resource allocation.
[0065] In some other embodiments, through a real-time pruning strategy, that is, by monitoring the number of nodes and the number of loop layers to dynamically adjust resources. For example, reducing the floating-point operation precision from FP32 to FP16 to release memory, but this rule-based adjustment has double defects. The pruning rules rely on manual experience and it is difficult to cover all scenarios. The precision degradation may lead to a decrease in the accuracy of legal inference results. Especially in cases of liquidated damages adjustment involving precise numerical calculations, the precision loss may directly affect the legality of the judgment.
[0066] That is to say, whether it is a static threshold or a rule-based dynamic adjustment, essentially it depends on a preset strategy and cannot perceive the semantic weight and resource consumption trend of the computational graph in real time.
[0067] For example, in the legal inference scenario, this limitation is manifested as incomplete inference due to the pruning strategy and system crashes caused by overly aggressive resource allocation.
[0068] In summary, the synergy between the inference depth and resource efficiency is low.
[0069] To solve the problem of low synergy between the inference depth and resource efficiency, this application provides a deep exploration AI inference method. By converting the user's natural language input into a dynamic computational graph with semantic weights, it accurately represents the logical structure and data flow of the inference task. It also performs a node-level composite complexity assessment on the computational graph to identify potential resource consumption risk points in advance. Combining the user permission level and the risk prediction result, it uses intelligent algorithms such as reinforcement learning to perform adaptive pruning on the computational graph to generate an optimized computational graph that not only meets the permission requirements but also effectively reduces resource risks. Finally, based on a multi-objective optimization model and real-time monitoring technology, it performs elastic resource scheduling on the optimized computational graph to ensure that in a complex environment, the inference tasks can all be executed efficiently and stably, while achieving the best balance between the load and the task execution efficiency.
[0070] As Figure 1 shown, the method includes the following steps:
[0071] S100: Generate a dynamic computational graph based on the task to be inferred.
[0072] The task to be inferred is a specific problem submitted by the user that requires in-depth logical analysis. For example, in natural language processing, there is a semantic parsing task, such as analyzing the sentiment tendency of a text segment and generating a summary; in computer vision, there is an object detection task, such as identifying the vehicle type and its location in an image.
[0073] Such tasks contain complex logical structures and dynamic data flows, and it is necessary to analyze their internal logical relationships and transform them into executable inference paths. The input forms of the tasks to be inferred include text, images, or voices. For example, the user inputs a text instruction through an interactive interface or uploads an image file to trigger the recognition process.
[0074] The dynamic computation graph is used to represent the logical structure and data flow of the task to be inferred. The dynamic computation graph is a real-time generated computational logic topological structure, including data nodes, logical nodes, and computational nodes, and each node is connected by directed edges. The data nodes correspond to input text segments or image features, the logical nodes represent the branch control of the inference path, and the computational nodes are specific arithmetic operations.
[0075] Exemplarily, in an industrial quality inspection task, the data nodes store the original images collected by the production line cameras, the logical nodes define the defect detection rules, and the computational nodes execute the image segmentation algorithm. The connection relationship of the edges reflects the dependency order of the task logic. For example, the original image is connected to defect localization through an edge and then flows to dimension measurement.
[0076] The dynamic nature of the dynamic computation graph is reflected in the real-time adjustment of the topological relationship between nodes and edges according to task requirements, and the quantification of the importance differences of nodes through semantic weights.
[0077] To enable the dynamic computation graph to fully express the task logic and avoid wasting resources on redundant paths. As Figure 2 shown, in some embodiments, after obtaining the task to be inferred, analyze the logical structure and data flow of the task to be inferred to generate an initial computation graph, the initial computation graph contains nodes, and the nodes include at least one or more combinations of data nodes, logical nodes, and computational nodes; based on a semantic embedding model, perform semantic weight calculation on the nodes and edges to generate a weighted dynamic computation graph.
[0078] The initial computation graph is the primary form of the dynamic computation graph, generated by a task parsing engine, containing unoptimized nodes and edges, and there are redundant paths or inefficient connections.
[0079] For example, in an image recognition task, the initial computation graph may contain multiple repeated feature extraction nodes, such as edge detection at different levels, or redundant data preprocessing branches, such as multiple resizing operations. The generation of the initial computation graph depends on task semantic parsing techniques. For example, in a text task, the subject-verb-object structure is extracted through dependency parsing to form a preliminary dependency relationship between data nodes.
[0080] The semantic embedding model is used to perform semantic weight calculation, mapping the text fragments associated with nodes to quantized weight values. The semantic embedding model converts each text fragment into a high-dimensional vector and then identifies key semantic features through an attention mechanism.
[0081] Semantic weights are parameters that quantify the importance of nodes and edges, determining the priority of resource allocation in the computation graph. The weight calculation is based on the semantic embedding model's understanding of task semantics. For example, in natural language processing, the model extracts semantic vectors of keywords, such as sentiment tendency vectors of "positive" and "negative", and combines them with a normalization function to generate weight values.
[0082] Nodes with high weight values have a significant impact on the inference result, such as keyword nodes in a sentiment classification task; nodes with low weight values correspond to auxiliary operations, such as punctuation processing.
[0083] Semantic weights are updated in real time. For example, in an image recognition task, the node weights are adjusted according to the real-time feature extraction results to ensure that high-influence features, such as vehicle contours, are preferentially allocated computing resources.
[0084] To address the anti-noise problem of semantic weight calculation, in some embodiments, the number of text fragments of the to-be-inferred task is obtained; based on the to-be-inferred task, the text fragments of the node are matched; based on the semantic embedding model, semantic weight calculation is performed on the node and the edge through the number of text fragments, the text fragments, and the normalization function to generate node semantic weights, where the node semantic weights include one or more combinations of the data node, the logic node, and the computing node.
[0085] Semantic weight calculation is performed through the following formula:
[0086]
[0087] Where W(v i ) is the semantic weight of node (v i ), n is the number of text fragments, σ is the normalization function, S k is the kth text fragment associated with the node, and TF-IDF is to extract key information.
[0088] Exemplarily, after the user submits a task to be inferred, identifies the vehicles in the video and counts the quantity, the task parsing engine parses the logical structure of the video frame sequence to generate an initial computational graph. The initial computational graph includes data nodes, i.e., video frames, logical nodes, i.e., loop control for processing each frame, and computational nodes, i.e., vehicle detection models. The semantic embedding model extracts the text descriptions associated with the nodes, i.e., "frame decoding" and "feature extraction", to generate semantic weights. For example, the weight value of the vehicle detection node is relatively high, while the weight value of the frame rate adjustment node is relatively low.
[0089] When the text segments associated with the nodes contain interfering vocabulary, the synergistic effect of the number of text segments and the normalization function can effectively suppress the influence of noise weights, making the semantic weights of key nodes higher than those of auxiliary operation nodes, providing an accurate basis for pruning, and solving the anti-noise problem of semantic weight calculation.
[0090] Through weight calculation, a weighted dynamic computational graph is obtained. The weighted dynamic computational graph is a computational graph assigned with semantic weights, and the weight values of its nodes and edges affect the optimization strategy. For example, in a sentiment analysis task, the edge between the high-weight sentiment classification node and the downstream summary generation node is strengthened to form an efficient connection; the low-weight redundant cleaning node is marked as an object to be pruned.
[0091] The weighted dynamic computational graph can be generated by semantic weight assignment and weight matrix storage. For example, the node weights and their associated edges are recorded through a graph database for the topology optimization module to call.
[0092] After generating the weighted computational graph, topological optimization processing is performed on the weighted dynamic computational graph to obtain a dynamic computational graph. Topological optimization processing is a process of adjusting the structure of the weighted dynamic computational graph, that is, enhancing the relevance of the critical path and reducing redundancy.
[0093] The node connections are dynamically adjusted according to the semantic weights. For example, the edge connection strength of high-weight nodes is increased, or the redundant edges of low-weight nodes are deleted. In a loop iteration scenario, the weights of the reverse edges of the loop control nodes are corrected. For example, after predicting the loop termination probability, the weights of the invalid iteration paths are reduced. The output of topological optimization is an optimized dynamic computational graph with a compact structure and prominent critical path. For example, in an image recognition task, only the high-weight feature extraction and classification paths are retained, and the repeated scaling operations are deleted.
[0094] Exemplarily, the weighted dynamic computational graph enters the topological optimization stage, enhancing the connection strength of high-weight nodes, such as vehicle detection, and deleting low-weight nodes, such as repeated frame decoding. In the optimized dynamic computational graph, the vehicle detection node is directly connected to the classification node to form a path.
[0095] Some methods rely on static rules for resource prediction of loop structures and cannot adapt to fluctuations in the number of iterations in dynamic task scenarios. For example, in a video stream analysis task, the number of iterations for loop processing a frame sequence may fluctuate significantly due to changes in video length or complexity.
[0096] In some embodiments, to accurately identify and optimize loop control nodes in a computational graph and dynamically adjust the weights of their reverse logical dependency edges, loop termination probability analysis is performed on the loop control nodes to generate a loop termination probability distribution.
[0097] Based on the loop termination probability distribution and a reverse edge weight adjustment algorithm, a dynamic weight value is calculated, where the dynamic weight value is the weight of the reverse logical dependency edge of the loop control node.
[0098] According to the dynamic weight value, the topological structure of the weighted dynamic computational graph is updated to obtain a dynamic computational graph.
[0099] Among them, a loop control node is the core unit in a dynamic computational graph that manages loop iteration logic and includes logical nodes with conditional judgment and process control functions. For example, in a video stream processing task, a loop control node is used to determine whether to continue processing the next frame, and its termination condition can be based on a frame number threshold or a feature convergence state.
[0100] The generation of a node is through the recognition of loop logic by a task parsing engine. For example, in a natural language multi-turn dialogue task, after parsing the "continue to ask" intention of the user input, a loop control node is generated to manage the dialogue turns. The structure of a loop control node includes condition evaluation (e.g., determining whether the loop termination condition is met), iteration counting (e.g., recording the current turn), and resource monitoring (e.g., providing real-time feedback on memory occupancy).
[0101] The loop termination probability distribution is the probability prediction result of the loop control node terminating iteration in advance and depends on historical task data statistics and real-time state analysis. For example, in an image recognition task, the historical average number of iterations for a certain loop branch is ten times, but in the current task, the feature convergence speed is relatively fast, and it is predicted that the iteration may terminate after five times.
[0102] The calculation of the probability distribution can be achieved through a Bayesian model or a Markov chain. For example, according to the feature change trend in the previous three iterations, the subsequent termination probability is dynamically corrected. The output form of the probability distribution is a list of discrete probability values. For example, the termination probability for the fifth iteration is 30%, and for the sixth iteration is 50%, which is called by the reverse edge weight adjustment algorithm.
[0103] Based on the loop termination probability distribution and a reverse edge weight optimization algorithm, the weights of the reverse logical dependency edges of the loop control node are calculated to generate a reverse edge weight matrix, and the dynamic weight value is calculated by the following formula:
[0104]
[0105] Among them, W reverse (e ij ) is the weight value of the reverse logical dependency edge e ij N iter is the predicted value of the maximum number of iterations, N break is the number of iterations at which the loop terminates in historical statistics, p break (k) is the termination probability of the k-th iteration extracted from the loop termination probability distribution.
[0106] Perform multi-objective fusion on semantic weights (node importance from natural language parsing), resource-sensitive weights (real-time calculated CPU / GPU consumption rate), and permission weights (access coefficients mapped by user levels) to generate a composite weight matrix.
[0107] In some embodiments, a permission attenuation coefficient can be embedded in the weight calculation to dynamically scale the weights of reverse edges according to the user level. The weights of reverse edges of low-permission tasks are restricted to a safe interval, forcibly constraining the activatable loop depth and branch complexity.
[0108] For example, physically delete the reverse edges with weights lower than the hard threshold to interrupt invalid logical dependencies, perform soft isolation on the edges in the transition interval, retain metadata but freeze resource allocation, and aggregate high-weight forward edges to form composite nodes to optimize the calculation path.
[0109] The semantic weight calculation of the dynamic computational graph and the topological optimization complement each other. Among them, the semantic weights enable key nodes to obtain resource supply, and the topological optimization suppresses the resource consumption of high-fluctuation paths. Through the synergistic effect of the loop termination probability distribution modeling and the reverse edge weight adjustment algorithm, it is possible to dynamically correct the loop path weights according to historical behavior patterns and real-time states to generate an adjusted dynamic computational graph.
[0110] S200: Perform a risk assessment on the dynamic computational graph to generate a risk prediction result.
[0111] The risk assessment at least includes complexity assessment, memory occupancy assessment, and data transfer volume assessment. The risk prediction result obtained through the risk assessment is a quantitative assessment of the resource consumption risk in the dynamic computational graph, including three dimensions: computational complexity, memory occupancy, and data transfer volume.
[0112] Exemplarily, in the scenario of smart city traffic scheduling, the risk prediction result can identify that the computational complexity of the "real-time road condition prediction model" node is high risk, while the memory occupancy of the "historical data cache" node is medium risk. When generating the risk prediction result, historical execution data and real-time monitoring metrics can be combined. For example, the fluctuation range of the iteration times of the prediction loop subgraph is predicted through a reinforcement learning model and mapped to the resource consumption level.
[0113] In some methods, the evaluation of complexity uses a static model, that is, based on predefined calculation formulas or empirical thresholds. When facing a dynamically generated inference computation graph, it is difficult to adapt to the fluctuations in node execution time and task scenario changes. For example, in a real-time video stream processing task, the decoding time of frames with different resolutions varies significantly. The static model may underestimate the processing time of high-resolution frames, resulting in unbalanced resource allocation.
[0114] To solve the above problems, as Figure 3 shown, in some embodiments, the execution time of the node is calculated to generate a node execution time vector;
[0115] Based on the node execution time vector, the weight coefficient is adjusted through a reinforcement learning algorithm to construct a risk assessment model, where the weight coefficient is the weight coefficient of the risk assessment model;
[0116] Through the risk assessment model, the iteration times of the loop subgraph in the dynamic computation graph are predicted to generate a risk prediction result.
[0117] The node execution time vector is a parameter that quantifies the execution time of each node in the dynamic computation graph. The node execution time vector is a multi-dimensional vector, including computational complexity, memory occupancy, data transfer volume, etc.
[0118] The generation of the node execution time vector is through the real-time monitoring and historical data analysis of the node. For example, in an image classification task, the execution time vector of the feature extraction node may include the number of floating-point operations (such as 1e6 times), the memory occupancy (such as 2GB), and the data transfer scale with the downstream node (such as 10MB per second).
[0119] Exemplarily, the execution time of each node in the dynamic computation graph is calculated. For example, the actual execution time of the frame decoding node fluctuates due to video resolution (such as 50ms for 1080p frames and 30ms for 720p frames), and its execution time vector is updated in real time.
[0120] The adjusted weight coefficient can identify the influence intensity of key path nodes on the loop termination probability, and then construct an iteration number prediction model with clear causal relationships. The weight coefficient establishes a mathematical mapping relationship between node time consumption and resource risk. Through exploring the optimal weight combination, reinforcement learning enables the risk assessment model to simultaneously meet the dual constraints of low latency (shortening the iteration cycle) and high stability (avoiding memory overrun), achieving the optimization of resource efficiency and security.
[0121] The reinforcement learning algorithm is the logic for dynamically adjusting the weight coefficient of the risk assessment model. By interacting with the environment, that is, continuously optimizing the model parameters based on the actual task execution results. The implementation of the reinforcement learning algorithm includes state matrix construction and policy update. For example, encoding historical task data into a state matrix as the input for model training.
[0122] In some embodiments, to construct a risk assessment model, it includes: obtaining historical execution data, and based on the historical execution data, constructing a reinforcement learning state matrix;
[0123] Updating the weight coefficient through the reinforcement learning algorithm and the update parameters, so as to construct a risk assessment model based on the updated weight coefficient.
[0124] Historical execution data is the data for model training and optimization. Historical execution data may include structured information such as memory occupancy logs. Store these data through a time series database and establish indexes according to dimensions such as task type, time period, and hardware configuration.
[0125] The reinforcement learning state matrix is a structured container for storing historical task execution data. Its rows represent different task instances, and its columns correspond to dimensions such as node time consumption, resource consumption, and task results. When constructing the reinforcement learning state matrix, execution records in similar scenarios can be extracted from historical data as training samples.
[0126] The reinforcement learning algorithm is the Q-learning algorithm. The Q-learning algorithm dynamically adjusts the weight coefficient of the risk assessment model according to the difference between the current task state and historical data. For example, when detecting an increase in the high-resolution frame ratio, the algorithm increases the weight coefficient of computational complexity, making the model pay more attention to the floating-point operation volume rather than memory occupancy.
[0127] The Q-learning algorithm updates the weight coefficient through the following formula:
[0128]
[0129] where, Δα is the updated weight parameter, η is the preset learning rate, R actual is the actual task completion reward value, R predict is the model prediction reward value, T totalis the total task - related quantity, is the partial derivative.
[0130] It can be understood that the preset learning rate, the actual task - completion reward value, the model - prediction reward value, and the total task - related quantity are collectively referred to as update parameters.
[0131] The adjusted weight coefficient is used to construct a risk - assessment model. Through the dynamically adjusted risk - assessment model, the maximum number of iterations of the loop sub - graph is predicted. For example, the initial prediction is 100 complete iterations, but it is corrected to 80 times according to the real - time weight coefficient, and 20% redundant resources are reserved to cope with sudden loads. The risk - prediction result guides subsequent optimization strategies, such as simplifying the model of high - time - consuming violation - detection nodes (such as reducing the input resolution), or asynchronously processing low - priority nodes (such as logging).
[0132] And the risk - prediction result is calculated by the following formula:
[0133]
[0134] where N iter is the risk - prediction result, ∈ is the preset error tolerance, δ is the loop - state deviation value, p break is the loop - termination probability based on historical data statistics.
[0135] Through the dynamic weight adjustment of the node - time - consumption vector and the reinforcement - learning model, the risk assessment can be adapted to the task requirements in real time. For example, when processing high - resolution videos, the weight coefficient of the decoding node is automatically increased to improve its resource - allocation priority, and at the same time, the prediction error is corrected by combining historical task data. This dynamic adjustment mechanism can maintain high - precision evaluation in complex and variable inference scenarios, avoiding resource waste or task interruption caused by model rigidity.
[0136] S300: Based on the risk - prediction result, perform optimization processing on the dynamic computational graph to generate an optimized computational graph.
[0137] The optimized computational graph is used to constrain the depth of inference calculation and the range of inference - resource allocation. The optimized computational graph is the structure after permission - aware pruning and lightweight - function insertion on the original dynamic computational graph. The generation of the optimized computational graph needs to balance the calculation depth and resource constraints. Excessive pruning may lead to logical missing, while retaining too many nodes will cause resource overload.
[0138] To solve the problem of the disconnection between permission control and computational - graph optimization, in some embodiments, the user - permission level is obtained. The user - permission level is a hierarchical identifier of the user's operation permissions and can be a predefined or dynamically allocated security policy. The generation of the permission level depends on user identity authentication, and the permission level affects the pruning threshold and resource - allocation upper limit of the computational graph.
[0139] Then, calculate the average complexity of the dynamic computation graph. According to the user permission level and the average complexity, calculate the survival probability of nodes. The node survival probability is the probability value that each node in the dynamic computation graph is retained. Its calculation is based on the permission level and the average complexity of the computation graph, and the survival probability is calculated through the following formula:
[0140]
[0141] where, P keep (v i ) is the survival probability of node v i , L is the user permission level, L max is the preset highest permission level, W base is the preset weight reference value, C avg is the average complexity of the dynamic computation graph, C(v i ) is the complexity of node v i .
[0142] The generation of the survival probability fuses the permission coefficient and the complexity weight. For example, in low-permission tasks, the complexity weight is amplified, so that more high-risk nodes are marked as objects to be pruned.
[0143] Based on the survival probability, generate a pruning threshold. The pruning threshold is the critical probability value that determines whether a node is retained. Sort the paths of the nodes below the pruning threshold to generate a sorted list. According to the sorted list, perform pruning processing on the dynamic computation graph. Obtain a lightweight approximation function and insert the lightweight approximation function at the breakpoint, where the breakpoint is the breakpoint position in the dynamic computation graph after pruning processing, to generate an optimized computation graph.
[0144] Exemplarily, based on the survival probability distribution, generate a pruning threshold of 0.4. For nodes with a survival probability lower than 0.4, such as 3D reconstruction nodes, perform path centrality scoring and sorting. The 3D reconstruction nodes have a low score and are marked as pruning objects. After pruning, insert a lightweight approximation function at the breakpoint.
[0145] To obtain the lightweight approximation function, in some embodiments, it includes:
[0146] Perform function type processing on the sorted list to obtain an initial function structure, where the initial function structure is a polynomial approximation function or a rule simplification function;
[0147] Perform function adaptation processing on the survival probability to obtain a parameter configuration, where the parameter configuration is a low-order fast approximation parameter or a high-hit-rate cache index;
[0148] Based on the average complexity, perform function constraint processing to obtain a function combination, which includes an error-tolerant interpolation function or a cache query function with stable hits;
[0149] Perform fusion processing on the initial function structure, parameter configuration, and function combination to generate a lightweight approximation function.
[0150] The initial function structure can be selected from a preset function library according to the type of pruning node and the task scenario. For example, in an image processing task, if the pruning node is high-resolution feature extraction, select a polynomial approximation function as the initial structure. This function fits the input-output mapping relationship of the original deep learning model through a low-order polynomial, reducing the computational complexity from O(n 2 ) to O(n).
[0151] In the financial risk control scenario, for a trading rule matching node, select a rule simplification function to replace complex model reasoning through a predefined if-then logic chain. The selection of the initial function structure preferentially uses approximation functions for computationally intensive nodes and simplification functions for rule-intensive nodes.
[0152] Based on the survival probability value of the pruning node, dynamically configure function parameters to adapt to different privilege levels and resource constraints. For example, in a medical image analysis task, configure low-order fast approximation parameters, adopt a third-order polynomial and a fixed-step sampling strategy, and compress the single-inference time from 120 ms to 25 ms. The parameter configuration rule library stores parameter combinations corresponding to different survival probability intervals to ensure real-time and efficient adaptation.
[0153] To ensure that the output error of the lightweight function is within an acceptable range, double-constraint verification can be performed. For example, an error-tolerant interpolation function and a cache query function with stable hits are used synchronously.
[0154] Fuse the initial function structure, parameter configuration, and constraint conditions in multiple dimensions, such as through structure-parameter fusion or constraint-condition embedding.
[0155] Among them, for structure-parameter fusion, in an autonomous driving path planning task, a polynomial approximation function (initial structure) is combined with low-order fast approximation parameters (configuration) to form a constrained fast trajectory prediction function.
[0156] Through privilege-aware survival probability calculation and lightweight approximation function insertion, the pruning process can not only constrain the resource consumption of low-privilege tasks but also ensure the integrity of high-privilege tasks. This dynamic grading mechanism achieves a flexible balance between privilege control and task quality, avoiding resource waste or function loss.
[0157] In some embodiments, to predict changes in resource requirements in advance based on resource consumption trends and task queue conditions and make scheduling adjustments in advance, the method further includes: planning a resource allocation strategy for the optimization computation graph based on a multi-objective optimization model, and dynamically adjusting resource allocation by monitoring the system load status in real time.
[0158] Specifically, based on the optimization computation graph, a multi-objective optimization model is constructed. The multi-objective optimization model is a framework that balances load and task efficiency, defining multiple optimization objectives (such as minimizing task latency and maximizing resource utilization) and constraint conditions (such as memory capacity and permission level). The construction of the multi-objective optimization model is based on the node resource requirements of the optimization computation graph and the real-time system state.
[0159] Dynamically adjust the upper limit of the memory constraint of the multi-objective optimization model according to the user permission level to generate a permission-sensitive resource allocation strategy; among them, the constraint condition is calculated according to the following formula:
[0160] Mem prak ≤M total *(1 + λL);
[0161] Among them, Mem prak is the peak memory usage, M total is the total memory capacity, and λ is a preset expansion coefficient.
[0162] Then, solve the multi-objective optimization model by the Lagrange multiplier method to generate a resource mapping table. The solution of the model transforms the multi-objective into a single-objective problem with constraints by the Lagrange multiplier method, for example, dealing with permission grading constraints by introducing slack variables. The generation of the resource mapping table depends on the load data collected by the real-time monitoring module. When the GPU utilization rate exceeds 85%, the model automatically adjusts the mapping strategy to migrate some computing tasks to idle edge nodes.
[0163] Then, perform resource allocation on the optimization computation graph based on the resource mapping table, and obtain the resource consumption status and a preset trigger condition. If the resource consumption status meets the trigger condition, trigger a hierarchical resource recovery mechanism, and the hierarchical resource recovery mechanism is used to handle resource overlimit scenarios.
[0164] In some embodiments, calculate a comprehensive resource occupancy rate index based on the task status after resource allocation to generate a resource overlimit alarm. The comprehensive resource occupancy rate index is as follows:
[0165]
[0166] Among them, R(t) is the comprehensive resource occupancy rate, Used GPU (t) is the amount of resources already used by the GPU at time t, Total GPU is the total amount of resources of the GPU, UsedRAM (t) is the amount of resources already used in the memory at time t, and Total RAM is the total amount of resources in the memory, and μ is a preset weight coefficient.
[0167] When the comprehensive resource occupancy rate exceeds the preset occupancy threshold, the tasks of low-privilege users are forcibly suspended. The preset occupancy threshold is used to trigger emergency resource recovery; dynamic precision degradation processing is performed based on the computing requirements of high-privilege user tasks to generate degraded computing tasks. The dynamic precision degradation processing is used to ensure the basic execution of high-priority tasks.
[0168] The method provided in this embodiment solves the problem of quantifying the importance of nodes by dynamically assigning semantic weights to the computational graph, expanding the complexity evaluation from a single index to multiple dimensions such as semantic relevance and historical iteration patterns; the coupled calculation of permission levels and survival probabilities breaks the static threshold limit, and even if the node complexity is high, its complete execution can still be ensured based on the permission weights.
[0169] Based on the above-mentioned deep exploration AI inference method, as Figure 4 shown, some embodiments of the present application further provide a deep exploration AI inference system 100, including:
[0170] A dynamic graph construction module 110, configured to generate a dynamic computational graph based on the task to be inferred, where the dynamic computational graph is used to represent the logical structure and data flow of the task to be inferred;
[0171] A risk prediction module 120, configured to perform a risk assessment on the dynamic computational graph to generate a risk prediction result, where the risk assessment at least includes complexity assessment, memory occupancy assessment, and data transfer volume assessment;
[0172] A dynamic graph optimization module 130, configured to perform an optimization process on the dynamic computational graph based on the risk prediction result to generate an optimized computational graph, where the optimized computational graph is used to constrain the depth of inference calculation and the range of inference resource allocation.
[0173] In some embodiments, it further includes a policy scheduling module 140, configured to perform elastic resource scheduling on the optimized computational graph based on a multi-objective optimization model, and adjust the resource allocation policy by real-time monitoring of the resource consumption status. The resource allocation policy is used to balance the system load and the task execution efficiency.
[0174] The present application provides a deep exploration AI inference method and system. The method includes: generating a dynamic computation graph based on a task to be inferred, where the dynamic computation graph is used to represent the logical structure and data flow of the task to be inferred, and then performing a risk assessment on the dynamic computation graph to generate a risk prediction result. The risk assessment at least includes complexity assessment, memory occupancy assessment, and data transmission volume assessment; based on the risk prediction result, performing an optimization process on the dynamic computation graph to generate an optimized computation graph, where the optimized computation graph is used to constrain the depth of inference calculation and the range of inference resource allocation. The method converts natural language input into a semantically weighted dynamic computation graph, and identifies resource risks through node-level complexity assessment; performs adaptive pruning in combination with permission levels and reinforcement learning to generate an optimized computation graph; implements elastic resource scheduling based on a multi-objective model and real-time monitoring to solve the problem of low coordination between inference depth and resource efficiency.
[0175] For the similar parts between the embodiments provided in the present application, reference can be made to each other. The specific embodiments provided above are only several examples under the general concept of the present application, and do not constitute a limitation on the protection scope of the present application. For those skilled in the art, any other embodiments extended based on the solution of the present application without creative efforts fall within the protection scope of the present application.
Claims
1. A deep exploration AI reasoning method, characterized by: include: Generate a dynamic computation graph based on the task to be inferred, wherein the dynamic computation graph is used to represent the logical structure and data flow of the task to be inferred; Performing a risk assessment on the dynamic computation graph to generate a risk prediction result, wherein the risk assessment includes at least a complexity assessment, a memory usage assessment, and a data transmission volume assessment; Based on the risk prediction result, the dynamic computation graph is optimized to generate an optimized computation graph, where the optimized computation graph is used to constrain the inference computation depth and the inference resource allocation range.
2. The deep exploration AI inference method according to claim 1, wherein Generating a dynamic computation graph includes: Analyze the logical structure and data flow of the task to be inferred, and generate an initial computation graph, wherein the initial computation graph includes nodes, and the nodes include at least one or more combinations of data nodes, logic nodes, and computation nodes; Based on the semantic embedding model, performing semantic weight calculation on the nodes and edges to generate a weighted dynamic computation graph, wherein the edges are generated through the nodes; Topology optimization is performed on the weighted dynamic computation graph to obtain the dynamic computation graph.
3. The deep exploration AI reasoning method according to claim 2, characterized in that: The performing of semantic weight calculation includes: Obtaining the number of text segments for the task to be inferred; Matching the text segment of the node based on the task to be inferred; Based on the semantic embedding model, semantic weight calculation is performed on the nodes and edges through the number of text fragments, text fragments, and normalization function to generate node semantic weights, where the node semantic weights include one or more combinations of the data nodes, logic nodes, and computing nodes.
4. The deep exploration AI inference method according to claim 2, wherein The performing topology optimization on the weighted dynamic computation graph to obtain the dynamic computation graph includes: Performing loop termination probability analysis on loop control nodes to generate a loop termination probability distribution, wherein the loop control nodes are nodes in the weighted dynamic computation graph used to manage and execute loop iteration logic; Calculating a dynamic weight value based on the loop termination probability distribution and the reverse edge weight adjustment algorithm, where the dynamic weight value is the weight value of the reverse logical dependency edge of the loop control node; The topological structure of the weighted dynamic calculation graph is updated according to the dynamic weight value to obtain a dynamic calculation graph.
5. The deep exploration AI reasoning method according to claim 4, characterized in that: The performing risk assessment on the dynamic computation graph to generate a risk prediction result includes: Performing time consumption calculation on the loop control node to generate a node time consumption vector; Based on the node time consumption vector, a weight coefficient is adjusted by a reinforcement learning algorithm to construct a risk assessment model, where the weight coefficient is a weight coefficient of the risk assessment model; The risk assessment model is used to predict the number of iterations of the cyclic subgraph in the dynamic computation graph to generate a risk prediction result.
6. The deep exploration AI inference method according to claim 5, wherein, The risk assessment model is constructed, including: Get historical execution data; constructing a reinforcement learning state matrix based on the historical execution data; The weight coefficients are updated by a reinforcement learning algorithm and updated parameters, so as to construct a risk assessment model based on the updated weight coefficients.
7. The deep exploration AI reasoning method according to claim 5, characterized in that: The step of performing optimization processing on the dynamic calculation graph based on the risk prediction result to generate an optimized calculation graph includes: Get user permission level; Calculating the average complexity of the dynamic computation graph; Calculate the survival probability of the node according to the user permission level and the average complexity; Generate a pruning threshold based on the survival probability; Perform path sorting on the nodes below the pruning threshold to generate a sorted list; Perform pruning processing on the dynamic computational graph according to the sorted list; Obtain a lightweight approximation function and insert the lightweight approximation function at breakpoints to generate an optimized computational graph, where the breakpoints are the breakpoint positions in the dynamic computational graph after pruning processing; 8. The deep exploration AI inference method according to claim 7, wherein The obtaining of the lightweight approximation function includes: Perform function type processing on the sorted list to obtain an initial function structure, where the initial function structure is a polynomial approximation function or a rule simplification function; Perform function adaptation processing on the survival probability to obtain a parameter configuration, where the parameter configuration is a low-order fast approximation parameter or a high-hit-rate cache index; Based on the average complexity, perform function constraint processing to obtain a function combination, where the function combination includes an error-tolerant interpolation function or a stable-hit cache query function; Perform fusion processing on the initial function structure, the parameter configuration, and the function combination to generate a lightweight approximation function; 9. The deep exploration AI inference method according to claim 1, wherein The method further includes: Construct a multi-objective optimization model based on the optimized computational graph; Solve the multi-objective optimization model by the Lagrange multiplier method to generate a resource mapping table; Perform resource allocation on the optimized computational graph based on the resource mapping table and obtain the resource consumption status; Preset a trigger condition; If the resource consumption status meets the trigger condition, trigger a hierarchical resource recycling mechanism, where the hierarchical resource recycling mechanism includes forced suspension processing and execution precision degradation processing; 10. A deep exploration AI inference system, characterized in that, For performing the deep exploration AI inference method according to any one of claims 1-9, including: A dynamic graph construction module, configured to generate a dynamic computational graph based on a to-be-inferred task, where the dynamic computational graph is used to represent the logical structure and data flow of the to-be-inferred task; A risk prediction module, configured to perform risk assessment on the dynamic computational graph to generate a risk prediction result, where the risk assessment at least includes complexity assessment, memory occupancy assessment, and data transmission volume assessment; A dynamic graph optimization module, configured to perform optimization processing on the dynamic computational graph based on the risk prediction result to generate an optimized computational graph, where the optimized computational graph is used to constrain the inference calculation depth and the inference resource allocation range;
Citation Information
Cited By
Inference acceleration optimization method and system applied to intelligent dialogue large model
CN120725158A
Model reasoning scheduling system
CN121078050A
Graph structure-based machine learning Pipeline generation and deployment method and system
CN121303272A
Computational graph optimization method, system and device supporting large model reasoning
CN121390320A
Computational graph optimization methods, systems, and devices supporting large model inference
CN121390320B