Dynamic reasoning path optimization method based on neural architecture search
Through neural architecture search and dynamic inference path optimization methods, the neural network architecture is adjusted in real time, solving the problem of waste of computing resources and energy consumption of traditional static neural networks in complex and changing environments, and achieving efficient and low-energy inference path optimization.
Patent Information
- Application Number
- CN202510666314.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-26
AI Technical Summary
When traditional static neural network architecture faces complex and changeable input data and dynamic operating environments, there are problems such as wasting computing resources, difficulty in adapting to dynamic environments and high energy consumption.
The dynamic inference path optimization method based on neural architecture search is adopted, and the neural architecture search algorithm design, dynamic inference path construction, model training and optimization are used to monitor the input data and system resource status in real time, dynamically adjust the inference path, build a deformable network topology, and use model pruning and quantization technology to compress the model.
It effectively solves the waste of computing resources, improves the inference speed, adapts to dynamic environments, reduces energy consumption, and ensures the efficiency and accuracy of inference.
Abstract
Description
Technical Field
[0001] The present invention relates to the field of neural network architecture technology, and specifically to a dynamic reasoning path optimization method based on neural architecture search. Background Art
[0002] Neural network architecture refers to the overall structure and organization of a neural network, which determines how the neural network processes input data, learns, and generates output.
[0003] In current AI model reasoning applications, traditional static neural network architectures often struggle to achieve efficient reasoning when faced with complex and changing input data and dynamic operating environments. This presents the following challenges:
[0004] 1. Waste of computing resources: Using a complete, complex neural network for inference on simple input data consumes excessive computing resources and time. For example, in image recognition tasks, for simple images with distinct features (such as a single geometric shape on a solid background), using a large image recognition model (such as ResNet-152) for inference will result in unnecessary computations on the numerous convolutional and fully connected layers in the model, resulting in wasted computing resources and slower inference speeds.
[0005] 2. Difficulty adapting to dynamic environments: In real-world applications, conditions such as data distribution, scale, and hardware resources can change dynamically. For example, during promotional periods on e-commerce platforms, user visits and the amount of data generated can increase significantly. Traditional static recommendation models are unable to dynamically adjust their inference paths based on real-time traffic and data changes, potentially leading to delayed recommendations, decreased accuracy, and a failure to meet user needs.
[0006] 3. High energy consumption: In mobile devices and edge computing scenarios, limited battery capacity and computing resources place stringent demands on model energy consumption. Using a fixed neural network architecture for inference can lead to excessive energy consumption for tasks that don't require high-complexity processing. For example, in smartwatch health monitoring applications, using complex deep learning models for simple heart rate data trend analysis can consume excessive power and shorten device battery life.
[0007] Based on the above, a dynamic reasoning path optimization method based on neural architecture search is invented. Summary of the Invention
[0008] To solve the above technical problems, according to one aspect of the present invention, the present invention provides the following technical solutions:
[0009] The dynamic reasoning path optimization method based on neural architecture search includes the following specific steps:
[0010] S1, neural architecture search algorithm design: first define the search space, then select the search strategy, and then set the constraints;
[0011] S2, dynamic reasoning path construction: first analyze the input data features, then guide the reasoning path based on the knowledge graph, and then use the model to predict the reasoning path based on the input data features and the knowledge graph;
[0012] S3, model training and optimization: first perform joint training, then perform model compression and acceleration.
[0013] As a preferred solution of the dynamic reasoning path optimization method based on neural architecture search described in the present invention, the specific steps of S1 are as follows:
[0014] S11, search space definition: construct a search space containing multiple network structure elements and connection methods;
[0015] S12, search strategy selection: using reinforcement learning and evolutionary algorithm strategies to find the optimal network architecture in the search space;
[0016] S13, constraint setting: clarify the constraints of computing resources, inference time requirements, and accuracy targets.
[0017] As a preferred solution of the dynamic reasoning path optimization method based on neural architecture search described in the present invention, the specific steps of S2 are as follows:
[0018] S21, input data feature analysis: preprocessing and feature analysis of input data;
[0019] S22, reasoning path decision model: Construct a reasoning path decision model to predict the optimal reasoning path based on the characteristics of the input data;
[0020] S23, dynamic path switching mechanism: First, it monitors changes in input data and system resource status in real time. Then, when it detects significant changes in input data characteristics or fluctuations in system resources, it triggers the dynamic path switching mechanism and promptly adjusts the inference path to ensure efficient and accurate inference.
[0021] S24, spatiotemporal dynamic network topology deformation: Construct a dynamically deformable network topology structure, allowing the network to adjust node connection methods and data flow paths in real time according to the dynamic changes and spatial characteristics of time series data during inference.
[0022] As a preferred solution of the dynamic reasoning path optimization method based on neural architecture search described in the present invention, the specific steps of S21 are as follows:
[0023] S211, data preprocessing: first clean the input data, then normalize the data;
[0024] S212, Image Data Feature Extraction: Use a pre-trained deep neural network to extract high-level semantic features of the image. First, input the image into the pre-trained network, and then obtain the output of the middle layer or the last layer of the network as the feature representation of the image;
[0025] S213, text data feature extraction: treat the text as a collection of words, ignore the order of the words, count the frequency of each word in the text, and construct a feature vector;
[0026] S214, feature screening: select the most representative features based on their relevance to the target task and remove redundant and irrelevant features;
[0027] S215, feature dimensionality reduction: When there are too many features, in order to reduce the amount of calculation and avoid the dimensionality disaster, the features can be reduced in dimension;
[0028] S216, feature encoding: For non-numeric data, convert it into numeric data for subsequent processing and analysis;
[0029] S217, Feature Fusion: When processing multimodal data or multiple types of features, features from different sources can be fused to fully utilize the comprehensive information of the data.
[0030] As a preferred solution of the dynamic reasoning path optimization method based on neural architecture search described in the present invention, the specific steps of S22 are as follows:
[0031] S221, select model architecture: select the corresponding model architecture based on task complexity and data characteristics;
[0032] S222, prepare training data: first collect data related to the application scenario and annotate it according to input and output requirements, then divide the data into training set, validation set and test set;
[0033] S223, data preprocessing: performing unified preprocessing operations on the training set, validation set, and test set data;
[0034] S224, model initialization: initializing the model parameters according to the selected model architecture;
[0035] S225, training model: input the training set data into the model, calculate the model output through forward propagation, and compare it with the labeled correct output to calculate the loss function value. Then, use the backpropagation algorithm to update the model parameters according to the loss function, and continuously adjust the model parameters to gradually reduce the loss function value and optimize the model performance;
[0036] S226, data input: In practical applications, the input data after feature analysis is input into the inference path decision model;
[0037] S227, decision output: The model analyzes and calculates the input data and outputs the corresponding reasoning path decision result.
[0038] As a preferred solution of the dynamic reasoning path optimization method based on neural architecture search described in the present invention, the specific steps of S23 are as follows:
[0039] S231, real-time monitoring: real-time monitoring of changes in input data and system resource status;
[0040] S232, trigger condition judgment: first set the trigger threshold, then perform condition judgment;
[0041] S233, candidate path screening: first build a path library, then determine the screening conditions, and then perform path screening;
[0042] S234, Decision Model Selection: For the selected candidate paths, different decision-making methods are used to determine the final switching path. In simple scenarios, a direct selection can be made based on preset priorities, such as prioritizing paths with low computing resource consumption and satisfactory accuracy. In complex scenarios, models such as decision trees and neural networks can be used for decision-making.
[0043] S235, path switching operation: After determining the final switching path, first pause the currently executing reasoning task and save the intermediate state of the task. Then, reconfigure system resources according to the new reasoning path, load the corresponding reasoning model and data, and start a new reasoning task to ensure the continuity of the reasoning process and the consistency of the data.
[0044] As a preferred solution of the dynamic reasoning path optimization method based on neural architecture search described in the present invention, the specific steps of S232 are as follows:
[0045] S2321, set trigger threshold: according to different monitoring objects, pre-set reasonable trigger threshold;
[0046] S2322, condition judgment: compare the real-time monitoring data with the set trigger threshold, and use the rule engine or condition judgment algorithm to check whether the monitoring data meets the trigger condition one by one.
[0047] As a preferred solution of the dynamic reasoning path optimization method based on neural architecture search described in the present invention, the specific steps of S233 are as follows:
[0048] S2331, Path Library Construction: Based on different application scenarios and reasoning requirements, a candidate reasoning path library is pre-built. The path library contains multiple reasoning path solutions. Each solution records the reasoning model, computing resource requirements, expected reasoning time, and accuracy information in detail.
[0049] S2332, screening condition determination: Based on the triggering condition and the current system state, the screening conditions for candidate paths are determined. If the switching is triggered by insufficient system resources, paths with lower computing resource requirements are prioritized. If the switching is triggered by a change in input data, an inference path that is more suitable for the current data characteristics is selected.
[0050] S2333, path screening execution: use the screening algorithm to screen the candidate paths in the path library, exclude paths that do not meet the conditions, and retain the candidate path set that meets the screening conditions.
[0051] As a preferred solution of the dynamic reasoning path optimization method based on neural architecture search described in the present invention, the specific steps of S3 are as follows:
[0052] S31, Joint Training: Jointly train the dynamic architecture and model parameters. During the training process, not only the model parameters are optimized, but also the dynamic inference path selection strategy is optimized, so that the model can adaptively select the optimal path for inference under different inputs;
[0053] S32, Model Compression and Acceleration: Use model pruning and quantization techniques to compress and accelerate the model.
[0054] Compared with existing technologies:
[0055] 1. To solve the problem of wasted computing resources: The present invention effectively solves the problem of wasted computing resources through the design of neural architecture search algorithms and the construction of dynamic inference paths. In the search space definition, a variety of network structure elements and connection methods are constructed, and combined with the search strategies of reinforcement learning or evolutionary algorithms, network architectures that are adapted to different input data can be explored. When encountering simple input data, the path decision model will select lightweight neural network branches for inference based on the results of input data feature analysis, avoiding the use of a large number of unnecessary convolutional layers and fully connected layers in large image recognition models, greatly reducing computing resource consumption and significantly improving inference speed.
[0056] 2. Addressing the problem of difficulty in adapting to dynamic environments: During the construction of dynamic inference paths, the present invention can monitor changes in input data and system resource status in real time. Once it detects a surge in user visits or a substantial increase in data volume during a major promotion on an e-commerce platform, or fluctuations in hardware resources (such as a decrease in available memory), the dynamic path switching mechanism will be triggered. Through the trained path decision model, the inference path will be adjusted in a timely manner to ensure that the recommendation system can still provide recommendation services to users quickly and accurately under high traffic.
[0057] 3. Regarding energy consumption issues: The present invention compresses and accelerates the trained model by adopting model pruning, quantization and other technologies during the model training and optimization stages. This can reduce the model's storage occupancy and computational complexity without affecting the model's accuracy, thereby reducing the high energy consumption caused by the model to a certain extent. DETAILED DESCRIPTION
[0058] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be described in further detail below.
[0059] The present invention provides a dynamic reasoning path optimization method based on neural architecture search, which includes the following specific steps:
[0060] S1, neural architecture search algorithm design: first define the search space, then select the search strategy, and then set the constraints;
[0061] The specific steps of S1 are as follows:
[0062] S11, search space definition: construct a search space containing multiple network structure elements and connection methods;
[0063] S12, search strategy selection: using reinforcement learning and evolutionary algorithm strategies to find the optimal network architecture in the search space;
[0064] S13, constraint setting: clarify the constraints of computing resources, inference time requirements and accuracy targets;
[0065] S2, dynamic reasoning path construction: first analyze the input data features, then guide the reasoning path based on the knowledge graph, and then use the model to predict the reasoning path based on the input data features and the knowledge graph;
[0066] The specific steps of S2 are as follows:
[0067] S21, input data feature analysis: preprocessing and feature analysis of input data;
[0068] The specific steps of S21 are as follows:
[0069] S211, data preprocessing: first clean the input data, then normalize the data;
[0070] S212, Image Data Feature Extraction: Use a pre-trained deep neural network to extract high-level semantic features of the image. First, input the image into the pre-trained network, and then obtain the output of the middle layer or the last layer of the network as the feature representation of the image;
[0071] S213, text data feature extraction: treat the text as a collection of words, ignore the order of the words, count the frequency of each word in the text, and construct a feature vector;
[0072] S214, feature screening: select the most representative features based on their relevance to the target task and remove redundant and irrelevant features;
[0073] S215, feature dimensionality reduction: When there are too many features, in order to reduce the amount of calculation and avoid the dimensionality disaster, the features can be reduced in dimension;
[0074] S216, feature encoding: For non-numeric data, convert it into numeric data for subsequent processing and analysis;
[0075] S217, Feature Fusion: When processing multimodal data or multiple types of features, it is possible to fuse features from different sources to fully utilize the comprehensive information of the data;
[0076] S22, reasoning path decision model: Construct a reasoning path decision model to predict the optimal reasoning path based on the characteristics of the input data;
[0077] The specific steps of S22 are as follows:
[0078] S221, select model architecture: select the corresponding model architecture based on task complexity and data characteristics;
[0079] S222, prepare training data: first collect data related to the application scenario and annotate it according to input and output requirements, then divide the data into training set, validation set and test set;
[0080] S223, data preprocessing: performing unified preprocessing operations on the training set, validation set, and test set data;
[0081] S224, model initialization: initializing the model parameters according to the selected model architecture;
[0082] S225, training model: input the training set data into the model, calculate the model output through forward propagation, and compare it with the labeled correct output to calculate the loss function value. Then, use the backpropagation algorithm to update the model parameters according to the loss function, and continuously adjust the model parameters to gradually reduce the loss function value and optimize the model performance;
[0083] S226, data input: In practical applications, the input data after feature analysis is input into the inference path decision model;
[0084] S227, decision output: The model analyzes and calculates the input data and outputs the corresponding reasoning path decision result;
[0085] S23, dynamic path switching mechanism: First, it monitors changes in input data and system resource status in real time. Then, when it detects significant changes in input data characteristics or fluctuations in system resources, it triggers the dynamic path switching mechanism and promptly adjusts the inference path to ensure efficient and accurate inference.
[0086] The specific steps of S23 are as follows:
[0087] S231, real-time monitoring: real-time monitoring of changes in input data and system resource status;
[0088] S232, trigger condition judgment: first set the trigger threshold, then perform condition judgment;
[0089] The specific steps of S232 are as follows:
[0090] S2321, set trigger threshold: according to different monitoring objects, pre-set reasonable trigger threshold;
[0091] S2322, conditional judgment: compare the real-time monitoring data with the set trigger threshold, and use the rule engine or conditional judgment algorithm to check whether the monitoring data meets the trigger condition one by one;
[0092] S233, candidate path screening: first build a path library, then determine the screening conditions, and then perform path screening;
[0093] The specific steps of S233 are as follows:
[0094] S2331, Path Library Construction: Based on different application scenarios and reasoning requirements, a candidate reasoning path library is pre-built. The path library contains multiple reasoning path solutions. Each solution records the reasoning model, computing resource requirements, expected reasoning time, and accuracy information in detail.
[0095] S2332, screening condition determination: Based on the triggering condition and the current system state, the screening conditions for candidate paths are determined. If the switching is triggered by insufficient system resources, paths with lower computing resource requirements are prioritized. If the switching is triggered by a change in input data, an inference path that is more suitable for the current data characteristics is selected.
[0096] S2333, path screening execution: using a screening algorithm to screen candidate paths in the path library, excluding paths that do not meet the conditions, and retaining a set of candidate paths that meet the screening conditions;
[0097] S234, Decision Model Selection: For the selected candidate paths, different decision-making methods are used to determine the final switching path. In simple scenarios, a direct selection can be made based on preset priorities, such as prioritizing paths with low computing resource consumption and satisfactory accuracy. In complex scenarios, models such as decision trees and neural networks can be used for decision-making.
[0098] S235, path switching operation: After determining the final switching path, the currently executing reasoning task is first paused and the intermediate state of the task is saved. Then, the system resources are reconfigured according to the new reasoning path, the corresponding reasoning model and data are loaded, and a new reasoning task is started to ensure the continuity of the reasoning process and the consistency of the data;
[0099] S24, spatiotemporal dynamic network topology deformation: Constructing a dynamically deformable network topology structure, allowing the network to adjust node connections and data flow paths in real time according to the dynamic changes and spatial characteristics of time series data during inference;
[0100] S3, model training and optimization: first perform joint training, then perform model compression and acceleration;
[0101] The specific steps of S3 are as follows:
[0102] S31, Joint Training: Jointly train the dynamic architecture and model parameters. During the training process, not only the model parameters are optimized, but also the dynamic inference path selection strategy is optimized, so that the model can adaptively select the optimal path for inference under different inputs;
[0103] S32, Model Compression and Acceleration: Use model pruning and quantization techniques to compress and accelerate the model, remove unimportant connections and parameters in the model, and convert parameter data types from high-precision to low-precision. Without affecting the model's accuracy, this reduces the model's storage usage and computational complexity, further improving inference efficiency.
[0104] Although the present invention has been described above with reference to embodiments, various modifications may be made thereto and equivalent components may be substituted without departing from the scope of the present invention. In particular, as long as there are no structural conflicts, the various features of the embodiments disclosed herein may be combined with each other in any manner, and the omission of an exhaustive description of such combinations in this specification is solely for the sake of space and resource conservation. Therefore, the present invention is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
Claims
1. A dynamic reasoning path optimization method based on neural architecture search, characterized by: The specific steps are as follows: S1, neural architecture search algorithm design: first define the search space, then select the search strategy, and then set the constraints; S2, dynamic reasoning path construction: first analyze the input data features, then guide the reasoning path based on the knowledge graph, and then use the model to predict the reasoning path based on the input data features and the knowledge graph; S3, model training and optimization: first perform joint training, then perform model compression and acceleration.
2. The dynamic reasoning path optimization method based on neural architecture search according to claim 1 is characterized in that: The specific steps of S1 are as follows: S11, search space definition: construct a search space containing multiple network structure elements and connection methods; S12, search strategy selection: using reinforcement learning and evolutionary algorithm strategies to find the optimal network architecture in the search space; S13, constraint setting: clarify the constraints of computing resources, inference time requirements, and accuracy targets.
3. The method for dynamic reasoning path optimization based on neural architecture search according to claim 1, characterized in that: The specific steps of S2 are as follows: S21, input data feature analysis: preprocessing and feature analysis of input data; S22, reasoning path decision model: Construct a reasoning path decision model to predict the optimal reasoning path based on the characteristics of the input data; S23, dynamic path switching mechanism: First, it monitors changes in input data and system resource status in real time. Then, when it detects significant changes in input data characteristics or fluctuations in system resources, it triggers the dynamic path switching mechanism and adjusts the inference path in a timely manner to ensure efficient and accurate inference. S24, spatiotemporal dynamic network topology deformation: Construct a dynamically deformable network topology structure, allowing the network to adjust node connection methods and data flow paths in real time according to the dynamic changes and spatial characteristics of time series data during inference.
4. The method for dynamic reasoning path optimization based on neural architecture search according to claim 3, characterized in that: The specific steps of S21 are as follows: S211, data preprocessing: first clean the input data, then normalize the data; S212, Image Data Feature Extraction: Use a pre-trained deep neural network to extract high-level semantic features of the image. First, input the image into the pre-trained network, and then obtain the output of the middle layer or the last layer of the network as the feature representation of the image; S213, text data feature extraction: treat the text as a collection of words, ignore the order of the words, count the frequency of each word in the text, and construct a feature vector; S214, feature screening: select the most representative features based on their relevance to the target task and remove redundant and irrelevant features; S215, feature dimensionality reduction: When there are too many features, in order to reduce the amount of calculation and avoid the dimensionality disaster, the features can be reduced in dimension; S216, feature encoding: For non-numeric data, convert it into numeric data for subsequent processing and analysis; S217, Feature Fusion: When processing multimodal data or multiple types of features, features from different sources can be fused to fully utilize the comprehensive information of the data.
5. The method for dynamic reasoning path optimization based on neural architecture search according to claim 3, characterized in that: The specific steps of S22 are as follows: S221, select model architecture: select the corresponding model architecture based on task complexity and data characteristics; S222, prepare training data: first collect data related to the application scenario and annotate it according to input and output requirements, then divide the data into training set, validation set and test set; S223, data preprocessing: performing unified preprocessing operations on the training set, validation set, and test set data; S224, model initialization: initializing the model parameters according to the selected model architecture; S225, training model: input the training set data into the model, calculate the model output through forward propagation, and compare it with the labeled correct output to calculate the loss function value. Then, use the backpropagation algorithm to update the model parameters according to the loss function, and continuously adjust the model parameters to gradually reduce the loss function value and optimize the model performance; S226, data input: In practical applications, the input data after feature analysis is input into the inference path decision model; S227, decision output: The model analyzes and calculates the input data and outputs the corresponding reasoning path decision result.
6. The method for dynamic reasoning path optimization based on neural architecture search according to claim 3, characterized in that: The specific steps of S23 are as follows: S231, real-time monitoring: real-time monitoring of changes in input data and system resource status; S232, trigger condition judgment: first set the trigger threshold, then perform condition judgment; S233, candidate path screening: first build a path library, then determine the screening conditions, and then perform path screening; S234, decision model selection: for the selected candidate paths, different decision methods are used to determine the final switching path; S235, path switching operation: After determining the final switching path, first pause the currently executing reasoning task and save the intermediate state of the task. Then, reconfigure system resources according to the new reasoning path, load the corresponding reasoning model and data, and start a new reasoning task to ensure the continuity of the reasoning process and the consistency of the data.
7. The method for dynamic reasoning path optimization based on neural architecture search according to claim 6, characterized in that: The specific steps of S232 are as follows: S2321, set trigger threshold: according to different monitoring objects, pre-set reasonable trigger threshold; S2322, condition judgment: compare the real-time monitoring data with the set trigger threshold, and use the rule engine or condition judgment algorithm to check whether the monitoring data meets the trigger condition one by one.
8. The method for dynamic reasoning path optimization based on neural architecture search according to claim 6, characterized in that: The specific steps of S233 are as follows: S2331, Path Library Construction: Based on different application scenarios and reasoning requirements, a candidate reasoning path library is pre-built. The path library contains multiple reasoning path solutions. Each solution records the reasoning model, computing resource requirements, expected reasoning time, and accuracy information in detail. S2332, screening condition determination: Determine the screening conditions of the candidate paths based on the triggering conditions and the current system state; If the switch is triggered by insufficient system resources, the path with lower computing resource requirements will be prioritized. If the switch is triggered by changes in input data, the inference path that is more suitable for the current data characteristics will be selected. S2333, path screening execution: use the screening algorithm to screen the candidate paths in the path library, exclude paths that do not meet the conditions, and retain the candidate path set that meets the screening conditions.
9. The method for dynamic reasoning path optimization based on neural architecture search according to claim 1, characterized in that: The specific steps of S3 are as follows: S31, Joint Training: Jointly train the dynamic architecture and model parameters. During the training process, not only the model parameters are optimized, but also the dynamic inference path selection strategy is optimized, so that the model can adaptively select the optimal path for inference under different inputs; S32, Model Compression and Acceleration: Use model pruning and quantization techniques to compress and accelerate the model.
Citation Information
Patent Citations
System and method for machine learning architecture for multi-task learning with dynamic neural networks
CA3178364A1
AI intelligent data processing method for fire fighting field
CN119558394A
Neural network architecture searching method based on hierarchical structure information of image component
CN119886265A
Neural network architecture search method based on existing knowledge reasoning generation structure and coefficient
CN119940484A
Neural architecture search system and search method
US20230385603A1
Cited By
Heterogeneous graph neural network-based search model establishment method and system and application
CN120764639A
Model structure searching method and device, computer equipment and storage medium
CN121787480A
Model structure searching method and device, computer device and storage medium
CN121787480B
Multi-stage lightweight natural fire remote sensing image inference method based on feature complexity self-adaptation
CN122223571B