Hardware-aware transformer architecture search for intelligent ball making visual inspection
Patent Information
- Application Number
- CN202611030976.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-13
- Publication Date
- 2026-09-25
AI Technical Summary
[0007]本发明的目的在于提供一种面向智能造球视觉检测的硬件感知Transformer架构搜索方法,以解决视觉Transformer模型在工业现场端侧部署时难以同时满足生球粒径识别精度、超规格物料检测实时性、峰值显存占用和功耗约束的问题,并提高搜索得到的模型在目标端侧设备上的运行稳定性
(1)本发明将技术对象限定为智能造球工业视觉检测,并将视觉Transformer架构搜索过程与车间侧边缘计算设备的推理延迟、峰值显存占用和功耗约束相结合,能够得到更适合工业现场端侧设备部署的检测模型。
Smart Images

Figure CN122821097A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of industrial visual inspection, intelligent manufacturing, edge intelligence, artificial intelligence model compression and neural network architecture search technology. Specifically, it relates to a hardware-aware Transformer architecture search method for intelligent ball-making visual inspection, which is suitable for real-time detection and positioning of raw balls, oversized materials and abnormal material distribution areas in industrial images of the ball-making site on workshop-side edge computing nodes or industrial edge servers. Background Technology
[0002] Intelligent pelletizing processes typically require acquiring high-resolution industrial images at locations such as the pelletizing tray exit, conveyor belt transfer points, and material transfer nodes. These images are then used to detect and locate green pellet size distribution, oversized materials, abnormally large materials, or areas with abnormal material distribution. These targets often exhibit significant scale variations, mutual occlusion, complex backgrounds, dust interference, and dynamic material tumbling in the images. Relying solely on manual sampling or transmitting all images back to a remote server for processing is easily limited by detection frequency, transmission bandwidth, and response latency. Therefore, performing real-time or near-real-time inference on workshop-side edge computing nodes or industrial edge servers has high practical value.
[0003] Visual Transformers possess strong global modeling capabilities, enhancing the recognition of raw pellet boundaries, out-of-specification materials, and abnormal regions against complex material backgrounds. However, visual Transformers typically involve numerous parameters and complex self-attention calculations, which can lead to issues such as excessive inference latency, high peak memory usage, and significant power consumption fluctuations when directly deployed on industrial edge devices. Furthermore, workshop-side edge computing devices are constrained by computing power, heat dissipation, power consumption, and real-time production scheduling. Therefore, a model design approach that solely pursues detection accuracy is insufficient to meet the requirements of intelligent pelletizing visual inspection tasks.
[0004] Existing hardware-aware neural network architecture search methods typically incorporate both model accuracy and hardware cost into the search objective to find a network structure suitable for the target hardware. While these methods reduce the workload of manually designing model structures, in intelligent ball-making industrial vision inspection scenarios, the actual inference cost is not only related to the number of model layers, channels, attention heads, and input resolution, but also affected by factors such as industrial camera image resolution, ambient lighting, dust interference, material movement, edge computing node load, and production scheduling status. If only a single hardware cost prediction value is used to evaluate candidate models, it is difficult to reflect the reliability of the prediction results, potentially leading to significant performance fluctuations in the deployed application of the searched model.
[0005] Furthermore, existing visual Transformer compression or search methods often employ a two-stage process of "search first, then distillation," where a lightweight student network is first obtained through search, followed by offline distillation using a high-precision teacher model. This approach leaves candidate student subnetworks in the search phase lacking intermediate feature supervision from the teacher model, and the search strategy relies primarily on the validation accuracy obtained from short-cycle training. This can easily underestimate candidate structures with potential expressive power that have not yet fully converged. Simultaneously, the number of layers, attention heads, and feature dimensions of the student subnetwork dynamically change during the search process, leading to inconsistencies in layer and dimension between the student subnetwork and the fixed teacher model. This increases the difficulty of coupling the distillation mechanism with the search process.
[0006] Therefore, for the industrial visual inspection task of intelligent ball making, there is a need to provide a visual Transformer architecture search method that can evaluate the hardware cost of the target edge device, take into account the uncertainty of hardware cost prediction, and introduce teacher model supervision synchronously during the search iteration process, so as to obtain a lightweight inspection model that can be stably deployed on the workshop edge computing node or industrial edge server. Summary of the Invention
[0007] The purpose of this invention is to provide a hardware-aware Transformer architecture search method for visual inspection of intelligent ball-making equipment. This method addresses the challenges of simultaneously meeting constraints on green ball size recognition accuracy, real-time detection of out-of-specification materials, peak memory usage, and power consumption when deploying visual Transformer models on industrial edge devices. It also improves the operational stability of the searched model on the target edge device. The method takes intelligent ball-making industrial image training data, validation data, and the hardware constraints of the target edge device as input, and outputs a Pareto-optimal visual Transformer detection model that satisfies the edge hardware constraints.
[0008] To achieve the above objectives, the technical solution of the present invention is as follows: A hardware-aware Transformer architecture search method for intelligent ball-making visual inspection includes the following steps: Step S1: Obtain intelligent ball-making industrial visual inspection data and edge hardware constraints.
[0009] Acquire training images, verification images, and their target annotations in the intelligent pelletizing scenario; the training images and verification images are industrial images of the pelletizing site captured by industrial cameras deployed at the pelletizing tray outlet, belt transfer point, or material transfer node; the target annotations include the location boxes and category labels of raw pellet targets, oversized material targets, abnormally large material targets, or areas with abnormal material distribution.
[0010] Obtain the hardware constraints of the target edge device; the target edge device is an edge computing node or industrial edge server deployed at the pelletizing site or workshop, and the hardware constraints include at least inference latency constraints. Peak memory usage constraints and power consumption constraints ;in, This represents the maximum allowable latency for a single image inference operation. This indicates the maximum allowed video memory usage during the inference process. This indicates the upper limit of power consumption allowed under industrial field deployment conditions.
[0011] This step provides the data foundation for training, validating, and deploying constraint judgments.
[0012] Step S2: Construct a visual Transformer search space for industrial visual inspection of intelligent ball-making and determine candidate student subnetworks.
[0013] The visual Transformer search space includes the number of backbone network layers, embedding dimension per layer, number of attention heads, image patch size, multilayer perceptron scaling, number of detection head feature fusion layers, and input image scaling. Then, based on the current search strategy, a set of architecture parameters is selected from the visual Transformer search space, and this set of architecture parameters determines the current candidate student sub-network. The candidate student sub-network takes the industrial image of the pelletizing site to be detected as input, and after processing by the visual Transformer feature extraction module and the detection head, outputs at least one detection result; each detection result includes the target category (e.g., undeveloped pellet target, oversized material target, or abnormal region), target location bounding box coordinates, and corresponding confidence score.
[0014] This step yields the candidate student subnetwork in the current search iteration.
[0015] Step S3: Extract the architecture-hardware joint encoding of the candidate student subnetwork and predict the hardware cost based on random forest.
[0016] For the current candidate student subnetwork extraction architecture - hardware joint coding : in, The network topology features include the number of layers, embedding dimension, number of attention heads, image patch size, and number of detection head feature fusion layers. The operator attribute features are represented, including self-attention type, normalization operator type, activation function type, and feature fusion operator type; This indicates the hardware characteristics of the target device, including available video memory, memory bandwidth, peak computing power, and power consumption limits. The input is a hardware cost proxy prediction model, which is a regression prediction model built on random forest. Its training samples include the architecture-hardware joint encoding of multiple calibrated candidate student subnetworks, as well as the inference latency, peak memory usage and power consumption measured on the target device or equivalent calibration platform. The model input is the architecture-hardware joint encoding of the current candidate student subnetwork, and the output is the predicted mean and predicted standard deviation of each hardware cost dimension.
[0017] In one implementation, a random forest includes A decision tree. For hardware cost evaluation dimensions. , No. The output of each decision tree is denoted as Hardware cost evaluation dimensions Belongs to set ,gather At least including inference delay Peak memory usage and power consumption ,when ≤ , ≤ and ≤ When both conditions are met, the candidate student subnetwork is determined to satisfy the end-side hardware constraints. Random forests in dimensionality... The average estimated hardware cost on for: in, This represents the estimated hardware cost of the current candidate student subnetwork on the target device.
[0018] Random forest in dimensionality The predicted standard deviation for: in, This indicates the degree of dispersion of prediction results of different decision trees for the same candidate student subnetwork in a random forest, and is used to characterize the uncertainty of hardware cost prediction.
[0019] This step yields the mean hardware cost and prediction uncertainty of the candidate student subnetwork.
[0020] Step S4: Use the uncertainty of prediction to correct the hardware cost penalty term and update the architecture search strategy.
[0021] Train the current candidate student subnetwork and obtain the verification accuracy using the verification images. The verification accuracy The mean precision, recall, or weighted detection metric used in small object detection tasks can be used to represent the result, and normalized to [the appropriate value]. arrive The interval. Based on the interval obtained in step S3. and Calculate the hardware cost penalty item : in, Indicates the dimensions of hardware cost evaluation The constraint threshold on, Representing dimensions The cost penalty weight, This represents the amplification factor for prediction uncertainty. This represents the confidence penalty weight. Since both hardware cost bias and prediction standard deviation in the formula are expressed as... Normalization is performed on the hardware cost penalty term. It is a dimensionless quantity. When Exceed When any corrected hardware cost exceeds the corresponding constraint threshold, the candidate student subnetwork is determined to not meet the edge hardware constraints, and a corresponding penalty is added to the composite reward function to reduce the probability that the candidate student subnetwork will be sampled or selected as the final model in subsequent searches; when the prediction standard deviation Greater than the preset uncertainty threshold If the prediction stability of the hardware cost dimension is insufficient, the hardware cost penalty term corresponding to the candidate student subnetwork is increased.
[0022] Based on verification accuracy and hardware cost penalty items Calculate the compound reward function : in, This is used to evaluate the overall performance of the current student subnetwork in terms of small target detection accuracy and edge deployment stability. The sampler agent, based on... The search strategy is updated to make subsequent sampled student subnetworks tend to have higher detection accuracy, lower hardware cost, and lower prediction uncertainty.
[0023] This step yields the updated search strategy.
[0024] Step S5: Simultaneously perform multi-level cross-layer collaborative distillation during the search iteration.
[0025] In each search iteration, the pre-trained visual Transformer detection model is used as the teacher model, and the currently sampled candidate network is used as the student sub-network. Let the total number of layers in the teacher model be... The total number of student subnetwork layers is For students Determine the corresponding teacher level : in, Indicates with student level The teacher layer index corresponding to the network depth ratio and Used to restrict the teacher layer index to to Within the specified range, avoid index out-of-bounds errors. Create a student-level index. Teacher level index The dynamic hierarchical mapping relationship, where Indicates the relationship with the first The teacher layer index corresponding to each student layer in terms of network depth ratio; a cross-layer distillation set composed of multiple mapping pairs. This is used to determine the student and teacher layers involved in feature alignment. For the set... Each student level An adaptive projection layer is set at the student feature output end. Student feature map Projected onto teacher feature map The same dimension is used to eliminate dimensional differences between the two.
[0026] Calculate the cross-entropy loss based on the detection output of the student subnetwork. ; Calculate the feature map alignment loss based on the projected student feature map and the corresponding teacher feature map. : in, Indicates the student subnetwork number Feature maps output by the layer The teacher model is represented by the first... Feature maps output by the layer.
[0027] Calculate the spatial attention alignment loss based on the spatial attention graphs of the student subnetwork and the teacher model. : in, Indicates the student subnetwork number Spatial attention map of layers, The teacher model is represented by the first... Spatial attention map of the layer, This represents the distillation temperature coefficient. The spatial attention alignment loss is calculated based on the spatial attention distributions of the candidate student subnetwork and the teacher model, causing the spatial attention region of the candidate student subnetwork to move closer to the spatial attention region of the teacher model.
[0028] The three types of losses are combined into a total synergistic loss. : in, This represents the feature map alignment loss weights. This represents the weights of the spatial attention alignment loss. Utilizing... Train the current student subnetwork so that the candidate student subnetworks can obtain feature supervision from the teacher model during the search phase.
[0029] This step yields the student subnetwork after teacher-supervised training and its verification accuracy.
[0030] Step S6: Select the Pareto optimal model and complete the industrial field end-side deployment.
[0031] When the preset search stopping condition is reached, the accuracy will be verified. Estimated average hardware cost Prediction standard deviation The Pareto-optimal visual Transformer detection model is selected based on hardware constraints of the target edge device. The search stopping conditions include reaching the maximum number of search rounds, the Pareto front ceasing updates after several consecutive rounds, or the candidate student subnetwork satisfying constraints on inference latency, peak memory usage, and power consumption during prediction. After converting the selected model into an inference format supported by the target edge device, it is deployed to workshop-side edge computing nodes or industrial edge servers for real-time detection and localization of raw pellet targets, oversized material targets, abnormally large material targets, or areas with abnormal material distribution in industrial images of the pelletizing site.
[0032] Compared with the prior art, the present invention has the following beneficial effects: (1) This invention limits the technical object to intelligent ball-making industrial visual inspection, and combines the visual Transformer architecture search process with the inference latency, peak memory usage and power consumption constraints of the workshop-side edge computing device to obtain a detection model that is more suitable for the deployment of end-side devices in the industrial field.
[0033] (2) This invention outputs the mean of the estimated hardware cost simultaneously through random forest. and the predicted standard deviation Furthermore, the prediction standard deviation is introduced into the hardware cost penalty term, which imposes an additional penalty on candidate student subnetworks with high prediction uncertainty, thereby reducing the risk of fluctuations in the running cost of candidate models on the target end device.
[0034] (3) The present invention performs multi-level cross-layer collaborative distillation simultaneously during the search iteration process, so that the student sub-network can obtain intermediate feature supervision of the teacher model in the search stage, which helps to improve the lightweight visual Transformer model's ability to represent raw ball boundaries, oversized materials and complex material backgrounds.
[0035] (4) This invention solves the problem of inconsistent layer number and feature dimension between teacher model and student subnetwork by dynamic hierarchical mapping and adaptive projection layer, so that cross-layer distillation can adapt to the dynamically changing candidate structure during the search process.
[0036] (5) The present invention outputs a Pareto candidate visual Transformer detection model that achieves a trade-off between prediction hardware cost and detection accuracy. It can be used for deployment of workshop-side edge computing nodes or industrial edge servers to improve the real-time performance and stability of green ball size identification, out-of-specification material detection and anomaly warning in intelligent ball making scenarios. Attached Figure Description
[0037] Figure 1 This is a schematic diagram of the overall process of the visual Transformer hardware-aware architecture search method provided by the present invention.
[0038] Figure 2 This is a schematic diagram of the multi-level cross-layer collaborative distillation and feature alignment mechanism provided by the present invention. Detailed Implementation
[0039] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings and technical solutions.
[0040] like Figure 1 As shown, in the intelligent pelletizing industrial vision inspection scenario, training and validation data are first established based on the pelletizing site inspection task. The training and validation data include industrial images of the pelletizing site and target annotations. Targets can be unprocessed pellets, oversized materials, abnormally large materials, or areas with abnormal material distribution. The system simultaneously reads the hardware constraints of the target-side devices, including inference latency constraints. Peak memory usage constraints and power consumption constraints .in, This represents the maximum allowable latency for a single image inference operation. This indicates the maximum allowed video memory usage during the inference process. This indicates the upper limit of power consumption allowed under industrial field deployment conditions.
[0041] A visual Transformer search space is constructed. The search space includes the number of layers, embedding dimension, number of attention heads, image patch size, multilayer perceptron scaling, number of detection head feature fusion layers, and input image scaling. The sampler agent generates architectural actions from the search space, forming the current student sub-network accordingly. The current student sub-network receives industrial images of the ball-making site and outputs the target category, target location, and confidence score. This part corresponds to... Figure 1 The search space, sampler agent, and student subnetwork within it.
[0042] The current candidate student subnetwork is jointly encoded using architecture and hardware to obtain... .in, correspond Figure 1 The topological features in the image can include the number of layers, embedding dimension, number of attention heads, image patch size, and number of detection head feature fusion layers. correspond Figure 1 The operator features in the data can include self-attention type, normalization operator type, activation function type, and feature fusion operator type; correspond Figure 1 The hardware parameters can include available video memory, memory bandwidth, peak computing power, and power consumption limits.
[0043] Will After inputting the random forest proxy model, the random forest proxy model outputs the mean hardware cost. and the predicted standard deviation This section corresponds to Figure 1 Feature encoding extraction, random forest proxy model, hardware cost mean, and prediction standard deviation are discussed.
[0044] like Figure 2 As shown, in each iteration of the search process, the system simultaneously performs multi-level cross-layer collaborative distillation. The teacher model can be a high-precision visual Transformer detection model trained on intelligent ball-making industrial visual inspection data, and the student sub-network is the candidate model currently sampled by the sampler agent. Since the number of layers, embedding dimension, and number of attention heads of the student sub-network will change with the search action, there is usually no fixed one-to-one hierarchical relationship between the teacher model and the student sub-network.
[0045] Based on the total number of layers in the teacher model Total number of student subnetwork layers Establish a dynamic hierarchical mapping. For the student level... The system determines the corresponding teacher level. And multiple student layers are combined into a cross-layer distillation set. When the student subnetwork is shallow, multiple student layers can correspond to intermediate layers selected proportionally to the depth of the teacher model; when the student subnetwork is deep, the number of student layers participating in distillation can be increased to maintain the coverage of teacher feature supervision in the shallow, intermediate, and deep layers of the network. This part corresponds to... Figure 2 The teacher model, student subnetwork, and dynamic hierarchical mapping in the text.
[0046] For each corresponding student layer and teacher layer, an adaptive projection layer is set at the output of the intermediate layer corresponding to the student sub-network. Adaptive projection layer It can consist of linear layers, non-linear activation functions, and normalization layers, and is used to process student feature maps. Projected onto teacher feature map Consistent dimensionality. This setting allows feature map alignment loss to be computed within the same feature space even if the embedding dimension or number of channels in the current student subnetwork differs from that of the teacher model. This part corresponds to... Figure 2 The adaptive projection layer and feature dimension alignment path in the process.
[0047] Calculate three types of losses. The first type is cross-entropy loss. The first type is used to constrain the detection output of the student subnetwork for the category and location of industrial vision inspection targets; the second type is feature map alignment loss. The first category is used to constrain the alignment of the projected student feature map with the teacher feature map in the numerical space; the third category is spatial attention alignment loss. This is used to constrain the student subnetwork and the teacher model to maintain consistency in spatial attention distribution. The three types of losses are summarized into the overall collaborative loss. This data is then used to train the current student subnetwork. The trained student subnetwork achieves validation accuracy on the validation images. The accuracy of this verification will then participate Figure 1 The calculation of the composite reward function. This part corresponds to... Figure 2 In , , and .
[0048] Through the above mechanism Figure 2 The teacher model, dynamic hierarchical mapping, adaptive projection layer, student subnetwork, and three types of losses form a synergistic relationship: dynamic hierarchical mapping solves the problem of inconsistent numbers of teacher and student layers; adaptive projection layer solves the problem of inconsistent feature dimensions; and feature map alignment loss and spatial attention alignment loss guide the student subnetwork to learn effective representations of the teacher model from both intermediate features and attention distribution perspectives. Therefore, Figure 2 The distillation process shown is similar to Figure 1The search feedback processes shown are interconnected, forming a closed loop of "sampling candidate structures, predicting hardware costs, distilling and training student networks, calculating compound rewards, and updating sampling strategies".
[0049] This embodiment uses intelligent pelletizing industrial visual inspection as an application scenario to illustrate the visual Transformer hardware perception architecture search method of the present invention. The targets to be detected include raw pellets, oversized material targets, abnormally large material targets, and areas with abnormal material distribution. The target-side device is a workshop-side edge computing node or an industrial edge server, with deployment constraints set as follows: single-frame inference latency not exceeding 38.0ms, peak video memory usage not exceeding 620.0MB, and power consumption not exceeding 11.0W; device hardware characteristics include 4096MB of available video memory, 68GB / s memory bandwidth, and a peak computing power of 1.35 TOPS.
[0050] First, a visual Transformer search space is constructed. This search space includes parameters such as the number of backbone network layers, embedding dimension, number of attention heads, image patch size, multilayer perceptron scaling, number of detection head feature fusion layers, and input image scaling scale. In each round of the search, the sampler generates candidate student subnetworks and encodes the candidate structures as architecture-hardware joint features, which include network topology features, operator attribute features, and edge device hardware features.
[0051] Then, the architecture-hardware joint features are input into the random forest hardware cost surrogate model. The random forest outputs the predicted mean and standard deviation of the candidate student subnetworks in terms of inference latency, peak memory usage, and power consumption. In this embodiment, the hardware cost penalty term is not only determined based on the deviation between the predicted mean and the constraint threshold, but also introduces the prediction standard deviation as an uncertainty penalty, so that candidate structures with large fluctuations in hardware cost prediction are additionally penalized in the composite reward function.
[0052] Simultaneously, a pre-trained high-precision visual Transformer detection model is introduced as the teacher model during the search iteration process to perform multi-level cross-layer collaborative distillation on the current candidate student sub-networks. For student sub-networks with dynamically changing layers, a dynamic hierarchical mapping corresponding to the network depth ratio is established based on the number of student and teacher layers, and an adaptive projection layer is set at the output of the student's intermediate layer to make the student feature dimensions consistent with the teacher's feature dimensions. During training, the detection output loss, feature map alignment loss, and spatial attention alignment loss are comprehensively calculated, so that the student sub-network can obtain intermediate feature supervision from the teacher model during the search phase.
[0053] Once the search reaches the stopping condition, Pareto candidate models are selected based on the verification detection score, the mean of hardware cost prediction, the standard deviation of hardware cost prediction, and edge hardware constraints. In this embodiment, the parameters of a candidate model obtained using the method of this invention are: 6 backbone network layers, 96 embedding dimensions, 8 attention heads, 16 image patch size, 4 multilayer perceptron expansion ratio, 3 detection head feature fusion layers, and an input image scaling scale of 640. This candidate model has a detection score of 0.7911, a mean prediction inference latency of 28.73ms, a prediction standard deviation of 5.94ms, a mean prediction peak memory usage of 205.9MB, a prediction standard deviation of 13.8MB, and a mean prediction power consumption of 5.60W, with a prediction standard deviation of 0.45W. Even considering the robustness constraints after taking into account the prediction standard deviation, this candidate model still satisfies the edge inference latency, peak memory usage, and power consumption constraints.
[0054] To verify the effectiveness of this invention, its method was compared with two comparative strategies. The highest detection score obtained by the precision-only search strategy was 0.8553, but the prediction latency of its optimal candidate model reached 1413.62ms, far exceeding the real-time inference constraints of industrial field edge devices, and the robust deployability rate was 0, indicating that simply pursuing detection accuracy easily leads to heavy models that cannot be deployed. The highest detection score of the mean hardware-aware search strategy was 0.7736, with an optimal latency of 37.35ms and a robust deployability rate of 81.2%, but its average over-constraint risk was still 22.12%, indicating that considering only the mean hardware cost may still ignore the deployment risk caused by prediction fluctuations. In contrast, the detection score of the method of this invention was 0.7911, the optimal latency was 28.73ms, the robust deployability rate was increased to 90.0%, and the average over-constraint risk was reduced to 1.81%. Therefore, this invention can significantly improve the deployment stability of the visual Transformer model on industrial field edge devices while maintaining high industrial vision detection capabilities.
Claims
1. A hardware-aware Transformer architecture search method for intelligent ball-making visual inspection, characterized in that, Includes the following steps: Step S1: Obtain intelligent ball-making industrial visual inspection data and edge hardware constraints; Acquire training images, verification images, and their target annotations in the intelligent pelletizing scenario; the training images and verification images are industrial images of the pelletizing site captured by industrial cameras deployed at the pelletizing tray outlet, belt transfer point, or material transfer node; the target annotations include the location boxes and category labels of raw pellet targets, oversized material targets, abnormally large material targets, or areas with abnormal material distribution; Obtain the hardware constraints of the target edge device; the target edge device is an edge computing node or industrial edge server deployed at the pelletizing site or workshop, and the hardware constraints include at least inference latency constraints. Peak memory usage constraints and power consumption constraints ;in, This represents the maximum allowable latency for a single image inference operation. This indicates the maximum allowed video memory usage during the inference process. This indicates the upper and lower allowable power consumption limits under industrial field deployment conditions. Step S2: Construct a visual Transformer search space for industrial visual inspection of intelligent ball-making, and determine candidate student subnetworks; The visual Transformer search space includes the number of backbone network layers, embedding dimension per layer, number of attention heads, image patch size, multilayer perceptron scaling, number of detection head feature fusion layers, and input image scaling scale. Then, based on the current search strategy, a set of architecture parameters is selected from the visual Transformer search space, and this set of architecture parameters determines the current candidate student sub-network. The candidate student sub-network takes an industrial image of the pelletizing site as input, and after processing by the visual Transformer feature extraction module and the detection head, outputs at least one detection result. Each detection result includes the target category (e.g., unsized pellet target, out-of-specification material target, or abnormal region), target location bounding box coordinates, and corresponding confidence level. Step S3: Extract the architecture-hardware joint encoding of the candidate student subnetwork and predict the hardware cost based on random forest; For the current candidate student subnetwork extraction architecture - hardware joint coding ;Will The input is a hardware cost proxy prediction model, which is a regression prediction model built on random forest. Its training samples include the architecture-hardware joint encoding of multiple calibrated candidate student subnetworks, as well as the inference latency, peak memory usage and power consumption measured on the target device or equivalent calibration platform. The model input is the architecture-hardware joint encoding of the current candidate student subnetwork, and the output is the predicted mean and predicted standard deviation of each hardware cost dimension. Step S4: Use the uncertainty of prediction to correct the hardware cost penalty term and update the architecture search strategy; Train the current candidate student subnetwork and obtain the verification accuracy using the verification images. The verification accuracy The mean precision, recall, or weighted detection metric used in small object detection tasks can be used to represent the result, and normalized to [the appropriate value]. arrive The range; based on the average estimated hardware cost obtained in step S3. and the predicted standard deviation Calculate the hardware cost penalty item ; Based on verification accuracy and hardware cost penalty items Calculate the compound reward function : in, This is used to evaluate the overall performance of the current student subnetwork in terms of small target detection accuracy and edge deployment stability. Step S5: Simultaneously perform multi-level cross-layer collaborative distillation during the search iteration; In each search iteration, the pre-trained visual Transformer detection model is used as the teacher model, and the currently sampled candidate network is used as the student subnetwork; let the total number of layers in the teacher model be . The total number of student subnetwork layers is For students Determine the corresponding teacher level : in, Indicates with student level The teacher layer index corresponding to the network depth ratio and Used to restrict the teacher layer index to to Within the scope; establish a student-level index Teacher level index The dynamic hierarchical mapping relationship, where Indicates the relationship with the first The teacher layer index corresponding to each student layer in terms of network depth ratio; a cross-layer distillation set composed of multiple mapping pairs. This is used to determine the student and teacher layers involved in feature alignment; for the set Each student level An adaptive projection layer is set at the student feature output end. Student feature map Projected onto teacher feature map The same dimension is used to eliminate dimensional differences between the two. Calculate the cross-entropy loss based on the detection output of the student subnetwork. ; Calculate the feature map alignment loss based on the projected student feature map and the corresponding teacher feature map. ; Calculate the spatial attention alignment loss based on the spatial attention graphs of the student subnetwork and the teacher model. ; The three types of losses are combined into a total synergistic loss. : in, This represents the feature map alignment loss weights. Represent the spatial attention alignment loss weights; utilize Train the current student subnetwork so that the candidate student subnetworks can obtain feature supervision from the teacher model during the search phase; Step S6: Select the Pareto optimal model and complete the industrial field end-side deployment; When the preset search stopping condition is reached, the accuracy will be verified. Estimated average hardware cost Prediction standard deviation The system uses hardware constraints on the target edge device to select the Pareto-optimal visual Transformer detection model. The search stopping conditions include reaching the maximum number of search rounds, the Pareto front no longer being updated after several consecutive rounds, or the candidate student subnetwork satisfying inference latency, peak memory usage, and power consumption constraints during prediction. After converting the selected model into an inference format supported by the target edge device, it is deployed to the workshop-side edge computing node or industrial edge server for real-time detection and location of raw pellet targets, oversized material targets, abnormally large material targets, or abnormally distributed material areas in industrial images of the pelletizing site.
2. The hardware-aware Transformer architecture search method for intelligent ball-forming visual inspection according to claim 1, characterized in that, In step S3, architecture-hardware joint coding as follows: in, The network topology features include the number of layers, embedding dimension, number of attention heads, image patch size, and number of detection head feature fusion layers. The operator attribute features are represented, including self-attention type, normalization operator type, activation function type, and feature fusion operator type; This indicates the hardware characteristics of the target device, including available video memory, memory bandwidth, peak computing power, and power consumption limits.
3. The hardware-aware Transformer architecture search method for intelligent ball-forming visual inspection according to claim 1, characterized in that, In step S3, the random forest includes Decision tree; for hardware cost evaluation dimensions , No. The output of each decision tree is denoted as Hardware cost evaluation dimensions Belongs to set ,gather At least including inference delay Peak memory usage and power consumption ,when ≤ , ≤ and ≤ Simultaneously, upon establishment, it is determined that the candidate student subnetwork satisfies the end-side hardware constraints; the random forest in dimensionality... The average estimated hardware cost on for: in, This represents the estimated hardware cost of the current candidate student subnetwork on the target device. Random forest in dimensionality The predicted standard deviation for: in, This indicates the degree of dispersion of prediction results of different decision trees for the same candidate student subnetwork in a random forest, and is used to characterize the uncertainty of hardware cost prediction.
4. The hardware-aware Transformer architecture search method for intelligent ball-forming visual inspection according to claim 1, characterized in that, In step S4, the hardware cost penalty item The calculation is as follows: in, Indicates the dimensions of hardware cost evaluation The constraint threshold on, Representing dimensions The cost penalty weight, This represents the amplification factor for prediction uncertainty. Indicates the confidence penalty weight; when Exceed When any corrected hardware cost exceeds the corresponding constraint threshold, the candidate student subnetwork is determined to not meet the edge hardware constraints, and a corresponding penalty is added to the composite reward function to reduce the probability that the candidate student subnetwork will be sampled or selected as the final model in subsequent searches; when the prediction standard deviation Greater than the preset uncertainty threshold If the prediction stability of the hardware cost dimension is insufficient, the hardware cost penalty term corresponding to the candidate student subnetwork is increased.
5. The hardware-aware Transformer architecture search method for intelligent ball-forming visual detection according to claim 1, characterized in that, In step S5, Feature map alignment loss The calculation is as follows: in, Indicates the student subnetwork number Feature maps output by the layer The teacher model is represented by the first... Feature maps output by the layer; Spatial attention alignment loss The calculation is as follows: in, Indicates the student subnetwork number Spatial attention map of the layer, The teacher model is represented by the first... Spatial attention map of the layer, This represents the distillation temperature coefficient.