Practical training method of unmanned aerial vehicle assembling and debugging practical training teaching platform based on image recognition
By building a high-definition teaching screen and a dedicated installation box, combining a lightweight SSD network model and an intelligent error correction mechanism, the problems of insufficient hardware and insufficient guidance in traditional drone training and teaching are solved, and efficient identification and error correction of drone micro components are achieved, teaching quality and efficiency are improved, and personalized learning needs are met.
Patent Information
- Application Number
- CN202510530325.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-12
AI Technical Summary
In traditional drone practical training and teaching, the hardware facilities are insufficient, the guidance methods are lacking in intelligence, and the evaluation methods are not objective enough, resulting in uneven teaching quality and low learning efficiency for students, making it difficult to meet personalized needs.
Build a hardware architecture based on high-definition teaching screen and special installation box, adopt a lightweight SSD network model for image recognition, combine intelligent error correction and guidance mechanisms to achieve efficient recognition of micro components of drone and sub-pixel-level positioning, dynamically generate error correction guidance, integrate teaching and practical training functions, and optimize the platform user experience.
It realizes accurate identification and error correction of drone micro components, improves teaching quality and efficiency, students can correct mistakes in a timely manner, teachers can adjust teaching strategies based on data, and students can understand the learning progress themselves and meet personalized needs.
Smart Images

Figure CN120472729A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of unmanned aerial vehicle (UAV) practical training and teaching platforms, and in particular to a practical training method for an UAV assembly and debugging practical training and teaching platform based on image recognition. Background Art
[0002] In today's era of rapid advancement in drone technology, education and training in drone-related fields have become increasingly important. Drone assembly and commissioning are core skills in this field, and the quality of practical training directly impacts students' future careers in the industry. However, traditional drone practical training currently faces numerous challenges that need to be addressed.
[0003] First, in terms of teaching hardware facilities, many training sites lack comprehensive and targeted equipment. Commonly, the display resolution of teaching equipment is insufficient, resulting in poor playback quality during instructional videos. Students struggle to discern the details of drone assembly and debugging, such as the installation location and method of small components. Furthermore, there's a lack of scientific and rational planning for the storage and access of drone components. The lack of dedicated installation boxes for orderly storage makes finding components during practical training time-consuming and laborious, reducing learning efficiency and easily causing damage or loss. Furthermore, the lack of high-definition cameras above the training stations prevents students from clearly recording the assembly process. This not only hinders teachers' subsequent detailed analysis of student operations, but also prevents students from providing effective video footage to review their own processes and identify any errors. Second, regarding the instructional approach, existing teaching methods largely rely on on-site, one-on-one instruction from teachers. However, due to limited teacher attention, teachers can't simultaneously focus on every student's operational details, resulting in some students' errors not being promptly identified and corrected. This becomes even more challenging with large numbers of students, leading to uneven teaching quality. Furthermore, traditional instructional methods lack intelligent tools, preventing accurate analysis and real-time feedback based on students' real-time operational data, and thus failing to meet the needs of personalized instruction. Furthermore, traditional methods rely primarily on subjective teacher judgment to assess student learning outcomes, lacking objective and accurate data support. This makes it difficult to quantitatively assess the standardization and accuracy of student operations during drone assembly and debugging. This makes it difficult for teachers to fully understand students' learning progress and weaknesses, preventing them from adjusting their teaching strategies accordingly. Furthermore, it hinders students from accurately understanding their learning status and developing appropriate study plans.
[0004] To sum up, the existing UAV practical training teaching has obvious deficiencies in hardware facilities, guidance methods, evaluation means, combination of teaching and training, and platform usage experience. A new practical training teaching platform is urgently needed to solve these problems, so as to improve the quality and efficiency of UAV practical training teaching and cultivate professional talents that are more adaptable to the needs of industry development. Summary of the Invention
[0005] To solve the above problems, the present invention proposes a practical training method for a UAV assembly and debugging training teaching platform based on image recognition. The specific steps are as follows:
[0006] Step 1: Build the hardware architecture of the practical training platform. Install a high-definition teaching screen as the core equipment of the teaching institution at the training site to play the drone assembly and debugging teaching video. Equip the training site with a dedicated installation box to place the disassembled drone parts in an orderly manner, providing students with practical objects. Deploy high-definition cameras above the training stations to ensure clear footage of the entire drone assembly process.
[0007] Step 2: Build an image recognition software system. Use a lightweight SSD network model to establish a multi-scale feature fusion detection framework. Combined with the characteristics of drone micro-component installation scenarios, optimize the prior box parameters and loss function to achieve efficient recognition and sub-pixel positioning of drone micro-components in camera images, enabling the system to accurately obtain component categories, spatial coordinates, and confidence data.
[0008] Step 3: Establish an intelligent error correction and guidance mechanism. This performs a sequential logic match between the component data output by the SSD network and the preset standard assembly sequence. Based on the matching results, error correction instructions are dynamically generated. When deviations in the trainee's assembly operation are detected, a real-time annotation prompt is triggered. The error location and correct assembly method are intuitively displayed to the trainee via a teaching screen or other terminal device, helping the trainee to correct the operation promptly.
[0009] Step 4: Integrate teaching and practical training functions. Teaching videos are played on the teaching screen, and students select drone components from the installation box and assemble them according to the video content, achieving a close integration of teaching and practical training. The system records and analyzes the entire training process and automatically generates learning reports to help students understand their learning progress and weak points, while also providing data support for teachers to adjust teaching strategies.
[0010] Step 5: Optimize the platform experience, simplify the platform operation process, design a user-friendly interactive interface to facilitate students and teachers to quickly get started, regularly maintain and upgrade the platform hardware equipment to ensure the normal operation of the teaching screen and camera, iteratively optimize the image recognition software system, continuously improve its recognition accuracy and error correction response speed, and continuously improve the overall platform usage efficiency and learning effect.
[0011] 2. The training method for the UAV assembly and debugging training teaching platform based on image recognition according to claim 1 is characterized in that: the image recognition software system constructed in step 2 can be expressed as:
[0012] Step 2.1 Data acquisition and preprocessing
[0013] Before building the system, we first collected image data of drone micro-components from the training site to ensure that the data covered different backgrounds, lighting conditions, and viewing angles. We also used a high-resolution camera to record the assembly process and annotated key scenes for component disassembly, assembly, and verification.
[0014] Annotate the bounding box of the component for each image in the format (x min ,y min ,x max ,y max ) form; To facilitate model training and normalize the image, its mathematical expression is
[0015]
[0016] Among them, I norm is the normalized image value, I is the original image grayscale or RGB value, I min , I max are the minimum and maximum pixel values of the image respectively;
[0017] Step 2.2 Lightweight SSD network model
[0018] A lightweight SSD network is used as a detector, and feature maps of multiple scales are used to detect objects of different sizes. The features extracted in each layer of the network are represented as
[0019] f i =Conv i (I norm )
[0020] Among them, f i Represents the i-th layer feature map, Conv i Represents the convolution operation processing of the i-th layer; in SSD, features are extracted from multiple layers and the low-level fine features are fused with the high-level semantic features; the specific fusion formula can be expressed as
[0021]
[0022] Among them, F is the fused feature, N is the total number of fused feature layers, and w i The weight coefficients of each layer's features ensure that features of different scales can effectively complement each other, thereby improving the detection of drone micro-components. To meet the needs of real-time detection, the lightweight solution uses MobileNet as the basic network and builds an SSD detection head on this basis. It replaces traditional convolution with depthwise separable convolution to reduce the amount of computation and parameters.
[0023] Step 2.3. Prior box parameter optimization
[0024] The prior box is used to predict the target position in SSD, and the default boxes of multiple scales and aspect ratios are preset; for each feature map grid, a prior box is generated at the center position, and its parameters are defined as
[0025] b=(x,y,w,h)
[0026] Where b is the candidate box for target detection, x, y are the coordinates of the box center, w, h are the width and height, and w, h are calculated based on the given scale s and aspect ratio r:
[0027]
[0028] Scale setting: set the relative scale range of the prior frame to s∈[0.1,0.3] to meet the detection requirements of small-sized components; aspect ratio setting: select aspect ratio r∈{0.5,1.0,2.0} to cover the shapes of long, wide and approximately square components; the optimized prior frame parameters are recorded as b′, and the intersection over union (IoU) is used in model training to evaluate the matching degree between the frame and the real target; for each prior frame, the IoU calculation formula is
[0029]
[0030] Among them, g is the true target box. When the IoU exceeds the set threshold of 0.5, the prior box is considered to match the target;
[0031] Step 2.4. Loss function design and optimization
[0032] In SSD, the loss function of target detection combines positioning loss and confidence loss, and the total loss function L(x,c,l,g) is defined as
[0033]
[0034] Among them, x is the sample, c is the sample prediction category, L conf (x,c) is the classification loss, using softmax cross entropy loss; L loc (l,g) is the positioning loss, using Smooth L1 loss to calculate the deviation between the predicted bounding box l and the true box g; N is the number of positive samples, and α is the balance coefficient;
[0035] The SmoothL1 loss function is used to solve the problem that the L1 loss is too smooth under small deviations. The formula is
[0036]
[0037] Among them, x is the input variable of the Smooth L1 loss function. For each matched prior box, its positioning loss is
[0038]
[0039] Among them, Pos is the index of all positive samples that match the real target, cx is the horizontal coordinate of the center point, cy is the vertical coordinate of the center point, w is the width, and h is the height. The value of the m-th dimension of the i-th positive sample in the model prediction box, is the corresponding value of the i-th real target frame after "encoding"; to achieve sub-pixel positioning, the offset of the prior frame is encoded; let the center of the prior frame be (x a ,y a ) and size (w a ,h a ), the center of the real box is (x, y) and the size is (w, h), and the encoding formula is
[0040]
[0041] in, and The center position and size offset of the target box relative to the prior box are respectively. In this way, small displacements and scaling can be significantly reflected during the training process, thereby achieving higher-precision positioning;
[0042] Step 2.5. Model training and sub-pixel positioning implementation
[0043] To improve the robustness of the model, the training data is rotated, scaled, and flipped. A multi-stage learning rate schedule is used to ensure full convergence of the model. At the same time, an online difficult sample mining strategy is used to balance the ratio of positive and negative samples.
[0044] To ensure sub-pixel positioning accuracy, an interpolation algorithm is used to make fine adjustments based on the bounding box prediction; bilinear interpolation is used to perform quadratic fitting on the predicted position, which is mathematically expressed as
[0045]
[0046] Among them, f(x,y) is the fine position estimate after interpolation, I(x i ,y i ) represents the value of the surrounding pixels, w ij is the weight associated with the position offset; based on the preliminary positioning results, this method can achieve fine position adjustment through the gradient information of image details to improve positioning accuracy;
[0047] During the training process, the loss changes and detection accuracy are monitored in real time through the validation set, and the prior box parameters and loss function weights are fine-tuned; non-maximum suppression is performed on the prediction results, and the formula is
[0048]
[0049] Among them, NMS(b i ,b j ) is non-maximum suppression, b i and b j Represents two predicted bounding boxes to be compared. This step ensures that the model can stably output high-confidence and high-precision detection results;
[0050] Step 2.6 Deployment and real-time detection process
[0051] Deploy the trained model on a server with real-time processing capabilities; connect the camera to capture images in real time, normalize them through the pre-processing module, and then pass them into the SSD network for inference;
[0052] The real-time detection process is expressed as:
[0053] Step 2.6.1 Image acquisition: The camera continuously acquires images of the assembly scene;
[0054] Step 2.6.2 Image preprocessing: normalize each frame of image;
[0055] Step 2.6.3 Model Inference: Use the trained SSD model to calculate the category prediction c of each prior box i and position offset l i ;
[0056] Step 2.6.4 Post-processing: Decode the bounding box output by the model and apply NMS to obtain the final detection result;
[0057] Step 2.6.5 Sub-pixel positioning fine-tuning: Use the interpolation method to optimize the sub-pixel position of the preliminary bounding box results and output the final component position and classification results.
[0058] 3. The training method for the UAV assembly and debugging training teaching platform based on image recognition according to claim 1 is characterized in that the establishment of the intelligent error correction and guidance mechanism in step 3 can be expressed as follows:
[0059] Step 3.1 Data input and standard assembly sequence definition
[0060] Step 3.1.1 SSD network output data
[0061] The SSD model detection results obtained from step 2 include the category, position, size, and sub-pixel positioning adjustment information of each detection frame; the detected component data set is recorded as
[0062] F={(c i ,x i ,y i ,w i ,hi ,Δx i ,Δy i )|i=1,2,...,N}
[0063] Among them, c i is the component category predicted by the i-th detection box, x i ,y i is the center position of the prediction box, w i ,h i is the width and height of the prediction box, Δx i ,Δy i is the sub-pixel offset obtained after fine-tuning through bilinear interpolation, and N is the total number of detection results output by the SSD model in a single detection;
[0064] Step 3.1.2 Preset standard assembly sequence
[0065] Establish a standard assembly sequence, denoted as
[0066] S={s j |j=1,2,...,M}
[0067] Among them, in the jth standard step, each standard step s j Includes: Standard component categories based on assembly templates Standard position Standard size Assembly order mark order j , M is the total number of steps;
[0068] Step 3.2 Sequence matching and error quantification
[0069] Step 3.2.1 Part Category Matching
[0070] For each detection result i and the corresponding step j of the standard sequence, the category matching condition is first determined:
[0071]
[0072] Where i is the index of the i-th detection box in the detection result set F, j is the index of the j-th step in the standard assembly sequence S, and c i is the part category predicted by the i-th detection box, is the expected component category of the jth step in the standard assembly sequence. If it does not match, it can be judged as an assembly error;
[0073] Step 3.2.2 Position error calculation
[0074] When the component category is matched, the deviation between the detected component and the standard position is calculated; with the help of the final position (x' i,y' i ) defines the Euclidean distance error:
[0075]
[0076] Among them, (x i ,y i ) is the standard center coordinate required for the jth step in the standard assembly sequence, (x * j ,y * j ) is the standard center coordinate required for the jth step in the standard assembly sequence, E i is the Euclidean distance error between the i-th detection component and the standard position; (x' i ,y' i )=(x i +Δx i ,y i +Δy i ), Δx i and Δy i is the sub-pixel offset; if E i If it exceeds the preset threshold ε, it is considered as a position deviation error;
[0077] Step 3.2.3 Dimensional error detection
[0078] Perform the same check for width and height:
[0079] E w =|w' i -w * j |,E h =|h' i -h * j |
[0080] Among them, w' i and h' i The predicted width and height of the i-th component in the detection result, w * j and h * j is the standard width and height required for the jth step in the standard assembly sequence, if E w or E h If it exceeds the set range, it is considered as a dimensional deviation error;
[0081] Step 3.2.4 Assembly sequence verification
[0082] Compare the current step order with the standard order based on the timestamp of each detection frame or the order in the actual operation jAre they consistent? If a sequence error is detected, follow step 3.3 for error correction guidance.
[0083] Step 3.3 Error correction guide generation and real-time feedback
[0084] Step 3.3.1 Calculation of error vector and correction amount
[0085] Let the error vector be
[0086] d j =(x' i -x * j ,y' i -y * j )
[0087] Among them, d j is the error vector, Based on the error direction and magnitude, a correction suggestion is generated; the specific adjustment amount is calculated using the proportional factor α:
[0088]
[0089] Among them, (x corr ,y corr ) is the corrected target position suggested by the system, and α is the scaling factor for error correction adjustment, which guides the trainees to fine-tune the component position;
[0090] Step 3.3.2 Graphical annotation and real-time prompts
[0091] Overlay error data and correction vectors onto real-time video; utilize image annotation algorithms to display the currently detected error location and correction suggestions in the student's operating area;
[0092] Step 3.3.3 Feedback closed loop mechanism
[0093] Adopting iterative feedback detection, after the trainee adjusts the operation, the system collects the position data again to form a new error vector And apply the correction formula:
[0094]
[0095] in, is the error vector at the kth iteration, indicating the deviation between the current operating position and the standard position, α is the feedback gain coefficient, is the error vector at the k+1th iteration. When the error converges to the threshold range, the correction is confirmed to be successful. This closed-loop process ensures that the operation accuracy reaches the sub-pixel level and continuously optimizes the error correction prompt.
[0096] Step 3.4 System Integration and Operation Process
[0097] Step 3.4.1 Intelligent comparison module execution
[0098] The system compares the test results with the preset standard sequence S item by item and calculates the error E using the above formula i ,E w ,E h and sequence bias;
[0099] Step 3.4.2 Instant error correction feedback
[0100] When deviations or errors are found during comparison, the error correction algorithm generates corresponding adjustment suggestions and dynamically displays prompts on the teaching screen to guide students on how to adjust the assembly operation.
[0101] Step 3.4.3 Iterative Feedback Optimization
[0102] During the student's adjustment process, the system continuously monitors the operation results and repeatedly iterates the error correction guidance until the test results meet the standard requirements; finally, the system can record and feedback historical error data to provide a basis for teachers' subsequent teaching adjustments.
[0103] The present invention provides a practical training method for a UAV assembly and debugging training teaching platform based on image recognition, which has beneficial effects. The technical effects of the present invention are:
[0104] 1. This paper uses a lightweight SSD network model to build a multi-scale feature fusion detection framework. It also optimizes the prior box parameters and loss function for drone micro-component installation scenarios, achieving efficient recognition and sub-pixel localization of drone micro-components in camera-captured images. This enables the system to accurately obtain component classification, spatial coordinates, and confidence data, providing precise data support for subsequent intelligent error correction and ensuring accurate assessment of student operations.
[0105] 2. The lightweight solution developed in this paper uses MobileNet as the foundational network and builds an SSD detection head. This replaces traditional convolution with depthwise separable convolution, reducing computational complexity and parameter requirements. This ensures real-time detection while maintaining accuracy. The system rapidly processes camera images and provides timely feedback on component identification and location results without noticeable delay, ensuring smooth practical training.
[0106] 3. This invention performs sequential logic matching of component data output by the SSD network with a preset standard assembly sequence, dynamically generating error correction instructions based on the matching results. If a student's assembly operation deviates, a real-time annotation prompt is triggered, visually displaying the error location and the correct assembly method via a teaching screen or other terminal device. This allows students to identify and correct errors immediately, preventing incorrect operations from becoming a habit, improving operational standardization and accuracy, and helping them more quickly master drone assembly and debugging skills. BRIEF DESCRIPTION OF THE DRAWINGS
[0107] Figure 1 is a flow chart of the present invention;
[0108] Figure 2 This is a diagram of the SSD network structure using MobileNet as the basic network of the present invention. DETAILED DESCRIPTION
[0109] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:
[0110] The present invention relates to a UAV training and teaching platform, which builds a hardware architecture including a high-definition screen, a dedicated box and a camera, and constructs an image recognition system based on a lightweight SSD network to achieve efficient component identification and positioning. It also establishes an intelligent error correction guidance mechanism, integrates teaching and training functions, optimizes user experience, improves teaching quality and efficiency, assists students in mastering skills, and provides data support for teaching adjustments. The invention flow chart is as follows: Figure 1 As shown, the steps of the present invention are described in detail below.
[0111] Step 1: Build the hardware architecture of the training teaching platform
[0112] At the training site, a high-definition teaching screen is installed as the core equipment of the teaching institution to play drone assembly and debugging teaching videos. A special installation box is equipped at the training site to place the disassembled drone parts in an orderly manner to provide students with practical objects. High-definition cameras are deployed above the training workstations to ensure that the entire process of students assembling drones can be clearly captured.
[0113] Step 2: Build an image recognition software system
[0114] A lightweight SSD network model is used to build a multi-scale feature fusion detection framework. Combined with the characteristics of the drone micro-component installation scenario, the prior box parameters and loss function are optimized in a targeted manner to achieve efficient recognition and sub-pixel positioning of drone micro-components in the camera capture image, enabling the system to accurately obtain component categories, spatial coordinates and confidence data.
[0115] Step 2.1 Data acquisition and preprocessing
[0116] Before building the system, image data of drone micro-components was collected from the training site, ensuring that the data covered different backgrounds, lighting conditions, and viewing angles. A high-resolution camera was used to record the assembly process, and key scenes such as component disassembly, assembly, and verification were annotated.
[0117] Annotate the bounding box of the component for each image in the format (x min ,y min ,x max ,y max ) form. To facilitate model training and normalize the image, its mathematical expression is
[0118]
[0119] Among them, I norm is the normalized image value, I is the original image grayscale or RGB value, I min , I max are the minimum and maximum pixel values of the image, respectively.
[0120] Step 2.2 Lightweight SSD network model
[0121] A lightweight SSD network is used as a detector to detect objects of different sizes using feature maps of multiple scales. The features extracted in each layer of the network are represented as
[0122] f i =Conv i (I norm )
[0123] Among them, f i Represents the i-th layer feature map, Conv i Indicates the convolution operation processing of the i-th layer. In SSD, features are extracted from multiple layers and the low-level fine features are fused with the high-level semantic features. The specific fusion formula can be expressed as
[0124]
[0125] Among them, F is the fused feature, N is the total number of fused feature layers, and w i is the weight coefficient of each layer feature, ensuring that features of different scales can effectively complement each other, thereby improving the detection effect of drone micro-components. To meet the needs of real-time detection, the lightweight solution uses MobileNet as the basic network and builds the SSD detection head on this basis. Replacing traditional convolution with depthwise separable convolution reduces the amount of computation and parameters. The SSD network structure with MobileNet as the basic network is shown in the figure below. Figure 2 shown.
[0126] Step 2.3. Prior box parameter optimization
[0127] The prior box is used to predict the target position in SSD, and the default boxes of multiple scales and aspect ratios are preset. For each feature map grid, a prior box is generated at the center position, and its parameters are defined as
[0128] b=(x,y,w,h)
[0129] Where b is the candidate box for target detection, x, y are the coordinates of the box center, w, h are the width and height, and w, h are calculated based on the given scale s and aspect ratio r:
[0130]
[0131] Scale setting: Set the relative scale range of the prior frame to s∈[0.1,0.3] to meet the detection requirements of small-sized components. Aspect ratio setting: Select aspect ratio r∈{0.5,1.0,2.0} to cover the shape of components that are long, wide, and approximately square. The optimized prior frame parameters are recorded as b′. The intersection over union (IoU) is used in model training to evaluate the degree of match between the frame and the real target. For each prior frame, the IoU calculation formula is:
[0132]
[0133] Among them, g is the true target box. When the IoU exceeds the set threshold of 0.5, the prior box is considered to match the target.
[0134] Step 2.4. Loss function design and optimization
[0135] In SSD, the loss function of target detection combines positioning loss and confidence loss, and the total loss function L(x,c,l,g) is defined as
[0136]
[0137] Among them, x is the sample, c is the sample prediction category, L conf (x,c) is the classification loss, using softmax cross entropy loss; L loc (l,g) is the positioning loss, and the Smooth L1 loss is used to calculate the deviation between the predicted bounding box l and the true box g; N is the number of positive samples, and α is the balance coefficient.
[0138] The SmoothL1 loss function is used to solve the problem that the L1 loss is too smooth under small deviations. The formula is
[0139]
[0140] Among them, x is the input variable of the Smooth L1 loss function. For each matched prior box, its positioning loss is
[0141]
[0142] Among them, Pos is the index of all positive samples that match the real target, cx is the horizontal coordinate of the center point, cy is the vertical coordinate of the center point, w is the width, and h is the height. The value of the m-th dimension of the i-th positive sample in the model prediction box, is the corresponding value of the i-th real target box after "encoding". In order to achieve sub-pixel positioning, the offset of the prior box is encoded. Let the center of the prior box be (x a ,y a ) and size (w a ,h a ), the center of the real box is (x, y) and the size is (w, h), and the encoding formula is
[0143]
[0144] in, and They are the center position and size offset of the target box relative to the prior box, respectively. In this way, small displacements and scaling can be significantly reflected during the training process, thereby achieving higher precision positioning.
[0145] Step 2.5. Model training and sub-pixel positioning implementation
[0146] To improve model robustness, the training data is rotated, scaled, and flipped. A multi-stage learning rate schedule is used to ensure sufficient model convergence. Furthermore, an online difficult sample mining strategy is employed to balance the ratio of positive and negative samples.
[0147] To ensure sub-pixel positioning accuracy, an interpolation algorithm is used to fine-tune the bounding box prediction. The predicted position is quadratically fitted using bilinear interpolation, which is mathematically expressed as
[0148]
[0149] Among them, f(x,y) is the fine position estimate after interpolation, I(x i ,y i ) represents the value of the surrounding pixels, w ij is the weight associated with the position offset. Based on the preliminary positioning results, this method can achieve fine-tuning of the position through the gradient information of image details, thereby improving positioning accuracy.
[0150] During the training process, the loss change and detection accuracy are monitored in real time through the validation set, and the prior box parameters and loss function weights are fine-tuned. Non-maximum suppression is performed on the prediction results, and the formula is
[0151]
[0152] Among them, NMS(b i ,b j ) is non-maximum suppression, b i and b j Represents two predicted bounding boxes to be compared. This step ensures that the model can stably output high-confidence and high-precision detection results.
[0153] Step 2.6 Deployment and real-time detection process
[0154] The trained model is deployed on a server with real-time processing capabilities. A camera is connected to capture real-time images, which are normalized by the preprocessing module and then passed to the SSD network for inference.
[0155] The real-time detection process of step 2 is expressed as:
[0156] Step 2.6.1 Image acquisition: The camera continuously acquires images of the assembly scene.
[0157] Step 2.6.2 Image preprocessing: Normalize each frame of image.
[0158] Step 2.6.3 Model Inference: Use the trained SSD model to calculate the category prediction c of each prior box i and position offset l i .
[0159] Step 2.6.4 Post-processing: Decode the bounding box output by the model and apply NMS to obtain the final detection result.
[0160] Step 2.6.5 Sub-pixel positioning fine-tuning: Use the interpolation method to optimize the sub-pixel position of the preliminary bounding box results and output the final component position and classification results.
[0161] Step 3: Establish an intelligent error correction and guidance mechanism
[0162] The component data output by the SSD network is matched with the preset standard assembly sequence through timing logic, and error correction instructions are dynamically generated based on the matching results. When deviations in the trainee's assembly operation are detected, the real-time annotation prompt function is triggered, and the error location and correct assembly method are intuitively displayed to the trainee through the teaching screen or other terminal devices, assisting the trainee in correcting the operation in a timely manner.
[0163] Step 3.1 Data input and standard assembly sequence definition
[0164] Step 3.1.1 SSD network output data
[0165] The SSD model detection results obtained from step 2 include the category, position, size, and sub-pixel positioning adjustment information of each detection frame. The detected component data set is recorded as
[0166] F={(c i ,x i ,y i ,w i ,h i ,Δx i ,Δy i )|i=1,2,...,N}
[0167] Among them, c i is the component category predicted by the i-th detection box, x i ,y i is the center position of the prediction box, w i ,h i is the width and height of the prediction box, Δx i ,Δy i is the sub-pixel offset obtained by bilinear interpolation fine-tuning, and N is the total number of detection results output by the SSD model in a single detection.
[0168] Step 3.1.2 Preset standard assembly sequence
[0169] Establish a standard assembly sequence, denoted as
[0170] S={s j |j=1,2,...,M}
[0171] Among them, in the jth standard step, each standard step s j Includes: Standard component categories based on assembly templates Standard position Standard size Assembly order mark order j , M is the total number of steps.
[0172] Step 3.2 Sequence matching and error quantification
[0173] Step 3.2.1 Part Category Matching
[0174] For each detection result i and the corresponding step j of the standard sequence, the category matching condition is first determined:
[0175]
[0176] Where i is the index of the i-th detection box in the detection result set F, j is the index of the j-th step in the standard assembly sequence S, and c i is the part category predicted by the i-th detection box, is the expected component category of the jth step in the standard assembly sequence. If it does not match, it can be judged as an assembly error.
[0177] Step 3.2.2 Position error calculation
[0178] When the component category is matched, the deviation between the detected component and the standard position is calculated. With the help of the final position (x' i ,y' i ) defines the Euclidean distance error:
[0179]
[0180] Among them, (x i ,y i ) is the standard center coordinate required for the jth step in the standard assembly sequence, (x * j ,y * j ) is the standard center coordinate required for the jth step in the standard assembly sequence, E i is the Euclidean distance error between the i-th detection component and the standard position; (x' i ,y' i )=(x i +Δx i ,y i +Δy i ), Δx i and Δy i is the sub-pixel offset; if E i If it exceeds the preset threshold ε, it is considered as a position deviation error.
[0181] Step 3.2.3 Dimensional error detection
[0182] Perform the same check for width and height:
[0183] E w =|w' i -w * j |,E h =|h' i -h * j |
[0184] Among them, w' i and h' i The predicted width and height of the i-th component in the detection result, w * j and h * j is the standard width and height required for the jth step in the standard assembly sequence, if E w or E hIf it exceeds the set range, it is regarded as a dimensional deviation error.
[0185] Step 3.2.4 Assembly sequence verification
[0186] Compare the current step order with the standard order based on the timestamp of each detection frame or the order in the actual operation j If a sequence error is detected, proceed to step 3.3 for error correction guidance.
[0187] Step 3.3 Error correction guide generation and real-time feedback
[0188] Step 3.3.1 Calculation of error vector and correction amount
[0189] Let the error vector be
[0190] d j =(x' i -x * j ,y' i -y * j )
[0191] Among them, d j is the error vector, Based on the error direction and magnitude, a correction suggestion is generated. The specific adjustment amount is calculated using the proportional factor α:
[0192]
[0193] Among them, (x corr ,y corr ) is the corrected target position recommended by the system, and α is the proportional factor of the error correction adjustment, which guides the students to fine-tune the component position.
[0194] Step 3.3.2 Graphical annotation and real-time prompts
[0195] The error data and correction vectors are superimposed on the real-time video. Using the image annotation algorithm, the current error location and correction suggestions are displayed in the student's operating area.
[0196] Step 3.3.3 Feedback closed loop mechanism
[0197] Adopting iterative feedback detection, after the trainee adjusts the operation, the system collects position data again to form a new error vector And apply the correction formula:
[0198]
[0199] in, is the error vector at the kth iteration, indicating the deviation between the current operating position and the standard position, α is the feedback gain coefficient, is the error vector at the k+1th iteration. Correction is considered successful when the error converges to within the threshold. This closed-loop process ensures sub-pixel accuracy and continuously optimizes error correction prompts.
[0200] Step 3.4 System Integration and Operation Process
[0201] Step 3.4.1 Intelligent comparison module execution
[0202] The system compares the test results with the preset standard sequence S item by item and calculates the error E using the above formula i ,E w ,E h and sequence bias.
[0203] Step 3.4.2 Instant error correction feedback
[0204] When the comparison finds deviations or errors, the error correction algorithm generates corresponding adjustment suggestions and dynamically displays prompt information through the teaching screen to guide students on how to adjust the assembly operation.
[0205] Step 3.4.3 Iterative Feedback Optimization
[0206] The system continuously monitors the student's progress and iterates error correction instructions until the test results meet the standard requirements. Ultimately, the system records and provides feedback on historical error data, providing a basis for teachers to make subsequent teaching adjustments.
[0207] Step 4: Integrate teaching and practical training functions
[0208] Teaching videos are played on the teaching screen, and students select drone components from the installation box for assembly training according to the video content, achieving a close integration of teaching and practical training. The system records and analyzes the students' training process throughout, and automatically generates learning reports to help students understand their own learning progress and weak links, while providing data support for teachers to adjust their teaching strategies.
[0209] Step 5: Optimize the platform experience
[0210] Simplify the platform operation process and design a user-friendly interactive interface to facilitate students and teachers to quickly get started. Regularly maintain and upgrade the platform hardware equipment to ensure the normal operation of the teaching screen and camera, iteratively optimize the image recognition software system, continuously improve its recognition accuracy and error correction response speed, and continuously improve the overall platform usage efficiency and learning effects.
[0211] The above description is merely a preferred embodiment of the present invention and does not constitute any other form of limitation to the present invention. Any modification or equivalent variation based on the technical essence of the present invention shall still fall within the scope of protection claimed by the present invention.
Claims
1. The practical training method of the UAV assembly and debugging teaching platform based on image recognition has the following specific steps, which are characterized by: Step 1: Build the hardware architecture of the practical training platform. Install a high-definition teaching screen as the core equipment of the teaching institution at the training site to play the drone assembly and debugging teaching video. Equip the training site with a dedicated installation box to place the disassembled drone parts in an orderly manner, providing students with practical objects. Deploy high-definition cameras above the training stations to ensure clear footage of the entire drone assembly process. Step 2: Build an image recognition software system. Use a lightweight SSD network model to establish a multi-scale feature fusion detection framework. Combined with the characteristics of drone micro-component installation scenarios, optimize the prior box parameters and loss function to achieve efficient recognition and sub-pixel positioning of drone micro-components in camera images, enabling the system to accurately obtain component categories, spatial coordinates, and confidence data. Step 3: Establish an intelligent error correction and guidance mechanism. This performs a sequential logic match between the component data output by the SSD network and the preset standard assembly sequence. Based on the matching results, error correction instructions are dynamically generated. When deviations in the trainee's assembly operation are detected, a real-time annotation prompt is triggered. The error location and correct assembly method are intuitively displayed to the trainee via a teaching screen or other terminal device, helping the trainee to correct the operation promptly. Step 4: Integrate teaching and practical training functions. Teaching videos are played on the teaching screen, and students select drone components from the installation box and assemble them according to the video content, achieving a close integration of teaching and practical training. The system records and analyzes the entire training process and automatically generates learning reports to help students understand their learning progress and weak points, while also providing data support for teachers to adjust teaching strategies. Step 5: Optimize the platform experience, simplify the platform operation process, design a user-friendly interactive interface to facilitate students and teachers to quickly get started, regularly maintain and upgrade the platform hardware equipment to ensure the normal operation of the teaching screen and camera, iteratively optimize the image recognition software system, continuously improve its recognition accuracy and error correction response speed, and continuously improve the overall platform usage efficiency and learning effect.
2. The training method for the UAV assembly and debugging training teaching platform based on image recognition according to claim 1 is characterized by: The image recognition software system constructed in step 2 can be expressed as: Step 2.1 Data acquisition and preprocessing Before building the system, we first collected image data of drone micro-components from the training site to ensure that the data covered different backgrounds, lighting conditions, and viewing angles. We also used a high-resolution camera to record the assembly process and annotated key scenes for component disassembly, assembly, and verification. Annotate the bounding box of the component for each image in the format (x min ,y min ,x max ,y max ) form; To facilitate model training and normalize the image, its mathematical expression is Among them, I norm is the normalized image value, I is the original image grayscale or RGB value, I min , I max are the minimum and maximum pixel values of the image respectively; Step 2.2 Lightweight SSD network model A lightweight SSD network is used as a detector, and feature maps of multiple scales are used to detect objects of different sizes. The features extracted in each layer of the network are represented as f i =Conv i (I norm ) Among them, f i Represents the feature map of layer i, Conv i Represents the convolution operation processing of the i-th layer; in SSD, features are extracted from multiple layers and the low-level fine features are fused with the high-level semantic features; the specific fusion formula can be expressed as Among them, F is the fused feature, N is the total number of fused feature layers, and w i The weight coefficients of each layer's features ensure that features of different scales can effectively complement each other, thereby improving the detection of drone micro-components. To meet the needs of real-time detection, the lightweight solution uses MobileNet as the basic network and builds an SSD detection head on this basis. It replaces traditional convolution with depthwise separable convolution to reduce the amount of computation and parameters. Step 2.
3. Prior box parameter optimization The prior box is used to predict the target position in SSD, and the default boxes of multiple scales and aspect ratios are preset; for each feature map grid, a prior box is generated at the center position, and its parameters are defined as b=(x,y,w,h) Where b is the candidate box for target detection, x, y are the coordinates of the box center, w, h are the width and height, and w, h are calculated based on the given scale s and aspect ratio r: Scale setting: set the relative scale range of the prior frame to s∈[0.1,0.3] to meet the detection requirements of small-sized components; aspect ratio setting: select aspect ratio r∈{0.5,1.0,2.0} to cover the shapes of long, wide and approximately square components; the optimized prior frame parameters are recorded as b′, and the intersection over union (IoU) is used in model training to evaluate the matching degree between the frame and the real target; for each prior frame, the IoU calculation formula is Among them, g is the true target box. When the IoU exceeds the set threshold of 0.5, the prior box is considered to match the target; Step 2.
4. Loss function design and optimization In SSD, the loss function of target detection combines positioning loss and confidence loss, and the total loss function L(x,c,l,g) is defined as Among them, x is the sample, c is the sample prediction category, L conf (x,c) is the classification loss, using softmax cross entropy loss; L loc (l,g) is the positioning loss, using Smooth L1 loss to calculate the deviation between the predicted bounding box l and the true box g; N is the number of positive samples, and α is the balance coefficient; The SmoothL1 loss function is used to solve the problem that the L1 loss is too smooth under small deviations. The formula is Among them, x is the input variable of the Smooth L1 loss function. For each matched prior box, its positioning loss is Among them, Pos is the index of all positive samples that match the real target, cx is the horizontal coordinate of the center point, cy is the vertical coordinate of the center point, w is the width, and h is the height. The value of the m-th dimension of the i-th positive sample in the model prediction box, is the corresponding value of the i-th real target frame after "encoding"; to achieve sub-pixel positioning, the offset of the prior frame is encoded; let the center of the prior frame be (x a ,y a ) and size (w a ,h a ), the center of the real box is (x, y) and the size is (w, h), and the encoding formula is in, and The center position and size offset of the target box relative to the prior box are respectively. In this way, small displacements and scaling can be significantly reflected during the training process, thereby achieving higher-precision positioning; Step 2.
5. Model training and sub-pixel positioning implementation To improve the robustness of the model, the training data is rotated, scaled, and flipped. A multi-stage learning rate schedule is used to ensure full convergence of the model. At the same time, an online difficult sample mining strategy is used to balance the ratio of positive and negative samples. To ensure sub-pixel positioning accuracy, an interpolation algorithm is used to make fine adjustments based on the bounding box prediction; bilinear interpolation is used to perform quadratic fitting on the predicted position, which is mathematically expressed as Among them, f(x,y) is the fine position estimate after interpolation, I(x i ,y i ) represents the value of the surrounding pixels, w ij is the weight associated with the position offset; based on the preliminary positioning results, this method can achieve fine position adjustment through the gradient information of image details to improve positioning accuracy; During the training process, the loss changes and detection accuracy are monitored in real time through the validation set, and the prior box parameters and loss function weights are fine-tuned; non-maximum suppression is performed on the prediction results, and the formula is Among them, NMS(b i ,b j ) is non-maximum suppression, b i and b j Represents two predicted bounding boxes to be compared. This step ensures that the model can stably output high-confidence and high-precision detection results; Step 2.6 Deployment and real-time detection process Deploy the trained model on a server with real-time processing capabilities; connect the camera to capture images in real time, normalize them through the pre-processing module, and then pass them into the SSD network for inference; The real-time detection process is expressed as: Step 2.6.1 Image acquisition: The camera continuously acquires images of the assembly scene; Step 2.6.2 Image preprocessing: normalize each frame of image; Step 2.6.3 Model Inference: Use the trained SSD model to calculate the category prediction c of each prior box i and position offset l i ; Step 2.6.4 Post-processing: Decode the bounding box output by the model and apply NMS to obtain the final detection result; Step 2.6.5 Sub-pixel positioning fine-tuning: Use the interpolation method to optimize the sub-pixel position of the preliminary bounding box results and output the final component position and classification results.
3. The training method for the UAV assembly and debugging training teaching platform based on image recognition according to claim 1 is characterized by: The establishment of an intelligent error correction and guidance mechanism in step 3 can be expressed as follows: Step 3.1 Data input and standard assembly sequence definition Step 3.1.1 SSD network output data The SSD model detection results obtained from step 2 include the category, position, size, and sub-pixel positioning adjustment information of each detection frame; the detected component data set is recorded as F={(c i ,x i ,y i ,w i ,h i ,Δx i ,Δy i )|i=1,2,...,N} Among them, c i is the component category predicted by the i-th detection box, x i ,y i is the center position of the prediction box, w i ,h i is the width and height of the prediction box, Δx i ,Δy i is the sub-pixel offset obtained after fine-tuning through bilinear interpolation, and N is the total number of detection results output by the SSD model in a single detection; Step 3.1.2 Preset standard assembly sequence Establish a standard assembly sequence, denoted as S={s j |j=1,2,...,M} Among them, in the jth standard step, each standard step s j Includes: Standard component categories based on assembly templates Standard position Standard size Assembly order mark order j , M is the total number of steps; Step 3.2 Sequence matching and error quantification Step 3.2.1 Part Category Matching For each detection result i and the corresponding step j of the standard sequence, the category matching condition is first determined: Where i is the index of the i-th detection box in the detection result set F, j is the index of the j-th step in the standard assembly sequence S, and c i is the part category predicted by the i-th detection box, is the expected component category of the jth step in the standard assembly sequence. If it does not match, it can be judged as an assembly error; Step 3.2.2 Position error calculation When the component category is matched, the deviation between the detected component and the standard position is calculated; with the help of the final position (x' i ,y' i ) defines the Euclidean distance error: Among them, (x i ,y i ) is the standard center coordinate required for the jth step in the standard assembly sequence, (x * j ,y * j ) is the standard center coordinate required for the jth step in the standard assembly sequence, E i is the Euclidean distance error between the i-th detection component and the standard position; (x' i ,y' i )=(x i +Δx i ,y i +Δy i ), Δx i and Δy i is the sub-pixel offset; if E i If it exceeds the preset threshold ε, it is considered as a position deviation error; Step 3.2.3 Dimensional error detection Perform the same check for width and height: Among them, w' i and h' i The predicted width and height of the i-th component in the detection result, w * j and h * j is the standard width and height required for the jth step in the standard assembly sequence, if E w or E h If it exceeds the set range, it is considered as a dimensional deviation error; Step 3.2.4 Assembly sequence verification Compare the current step order with the standard order based on the timestamp of each detection frame or the order in the actual operation j Are they consistent? If a sequence error is detected, follow step 3.3 for error correction guidance. Step 3.3 Error correction guide generation and real-time feedback Step 3.3.1 Calculation of error vector and correction amount Let the error vector be d j =(x' i -x * j ,y' i -y * j ) Among them, d j is the error vector, Based on the error direction and magnitude, a correction suggestion is generated; the specific adjustment amount is calculated using the proportional factor α: Among them, (x corr ,y corr ) is the corrected target position suggested by the system, and α is the scaling factor for error correction adjustment, which guides the trainees to fine-tune the component position; Step 3.3.2 Graphical annotation and real-time prompts Overlay error data and correction vectors onto real-time video; utilize image annotation algorithms to display the currently detected error location and correction suggestions in the student's operating area; Step 3.3.3 Feedback closed loop mechanism Adopting iterative feedback detection, after the trainee adjusts the operation, the system collects position data again to form a new error vector And apply the correction formula: in, is the error vector at the kth iteration, indicating the deviation between the current operating position and the standard position, α is the feedback gain coefficient, is the error vector at the k+1th iteration. When the error converges to the threshold range, the correction is confirmed to be successful. This closed-loop process ensures that the operation accuracy reaches the sub-pixel level and continuously optimizes the error correction prompt. Step 3.4 System Integration and Operation Process Step 3.4.1 Intelligent comparison module execution The system compares the test results with the preset standard sequence S item by item and calculates the error E using the above formula i ,E w ,E h and sequence bias; Step 3.4.2 Instant error correction feedback When deviations or errors are found during comparison, the error correction algorithm generates corresponding adjustment suggestions and dynamically displays prompts on the teaching screen to guide students on how to adjust the assembly operation. Step 3.4.3 Iterative Feedback Optimization During the student's adjustment process, the system continuously monitors the operation results and repeatedly iterates the error correction guidance until the test results meet the standard requirements; finally, the system can record and feedback historical error data to provide a basis for teachers' subsequent teaching adjustments.
Citation Information
Cited By
Performance evaluation method and device for unmanned aerial vehicle
CN122242028A