Model training data processing method and system based on improved YOLO algorithm

Through the methods of data conversion and unified parameter definition, the insufficient adaptability of the YOLO algorithm in application expansion on embedded platforms is solved, the robustness and performance stability of model training data processing are achieved, and it adapts to changes in embedded platforms.

CN120633781APending Publication Date: 2025-09-12GUANGZHOU LINGMOU TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510716104.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

The existing YOLO algorithm's model training data processing lacks adaptability to application expansion on embedded platforms. It has problems such as fragmented parameter configuration, inability to perceive changes in external operating scenarios, lack of a feedback loop in the post-deployment model tuning mechanism, and lack of flexibility in cross-platform deployment.

Method used

By inputting a predefined data set, automatically converting it into a predefined format, extracting the original parameters of the YOLO algorithm configuration, performing evaluation and analysis, and generating a PC-side prediction and recognition model, converting it into an embedded application environment model, and updating the model configuration in combination with the original scheduling parameters, data conversion and parameter unified definition, dynamic feedback mechanism, and adaptation to changes in embedded platforms.

Benefits of technology

It improves the application expansion adaptability of the YOLO model on embedded platforms, achieves the robustness and performance stability of model training data processing, and solves the problems of parameter configuration fragmentation, insufficient perception of external scene changes, and cross-platform deployment flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633781A_ABST
    Figure CN120633781A_ABST
Patent Text Reader

Abstract

The invention discloses a model training data processing method and system based on an improved YOLO algorithm. The method relates to the technical field of YOLO algorithm data processing, and comprises the steps of data conversion and original data configuration, PC end prediction recognition model generation, embedded end prediction recognition model generation and embedded end prediction recognition model updating. The method comprises the following steps: extracting a predefined data set to identify and count original parameters configured by a YOLO algorithm; configuring initial parameters according to a YOLO algorithm, and inputting the initial parameters into a training script for data training; converting the PC end prediction identification model into an embedded application environment model; and configuring initial parameters according to the updated YOLO algorithm to reconfigure the embedded end prediction identification model. The effect of improving the application extension adaptability of the YOLO model on the embedded platform is achieved, and the problem that in the prior art, the application extension adaptability of model training data processing of a YOLO algorithm on the embedded platform is insufficient is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of YOLO algorithm data processing technology, and in particular to a model training data processing method and system based on an improved YOLO algorithm. Background Art

[0002] With the development of the Internet of Things, Industry 4.0, and smart cities, real-time object detection technology has become a core requirement in smart security, autonomous driving, drone inspection, and other fields. The proliferation of devices such as drones and smart cameras has further driven the demand for lightweight algorithms. Furthermore, the rise of edge computing and embedded devices has made embedded platforms a key enabler for algorithm implementation. The YOLO algorithm, due to its high efficiency, is widely used in scenarios such as real-time video surveillance and robot navigation. However, performance bottlenecks in resource-constrained environments need to be addressed.

[0003] The existing model training data processing method based on the improved YOLO algorithm is implemented through the following technologies: The CSPNet backbone network and FPN+PAN architecture are used to achieve multi-scale feature fusion; anchor-free detection simplifies the prediction process; Mosaic data augmentation and the CIoU loss function are used to improve model generalization; redundant channels are removed and FP32 to INT8 quantization is performed to reduce model size and computational complexity; the inference engine is optimized using TensorRT and ONNX Runtime, supporting GPU / TPU parallel computing; enhancement parameters are adjusted based on reinforcement learning to meet the needs of different training stages; and diverse training samples are synthesized to alleviate the problem of insufficient small object data.

[0004] For example, the invention patent with publication number CN118313427A discloses a method for constructing an improved Yolo network model based on experimental teaching and examinations, including: S1: collecting a training data set, cleaning, labeling, and checking the data in the data sample set, and then dividing the data sample set; S2: setting the data size and hyperparameters of each batch, performing iterative training of the model, and obtaining an improved Yolo network model; S3: performing automated testing and iterative updates on the trained improved Yolo network model, and finally obtaining an improved Yolo network model for middle school experimental teaching and examinations.

[0005] For example, the invention patent announcement with announcement number: CN113392973B discloses an FPGA-based AI chip neural network acceleration method, which includes: performing quantization training when training the YOLO network, converting the floating-point algorithm of the neural network into fixed-point, greatly reducing memory usage, improving computing speed and bandwidth, and achieving the effect of reducing power consumption; using the HLS development method to quickly generate the YOLO convolutional neural network accelerator IP core based on the Darknet framework, and at the same time transforming the convolution calculation.

[0006] However, in the process of implementing the technical solutions of the invention in the embodiments of the present application, the present application found that the above technology has at least the following technical problems:

[0007] In the existing technology, there are many difficulties in the embedded deployment process of the existing YOLO algorithm's target detection model, such as fragmented parameter configuration, no unified standard, inability to perceive changes in external operating scenarios, lack of a feedback loop in the post-deployment model tuning mechanism, lack of flexibility in cross-platform deployment, etc. There is also the problem of insufficient adaptability of the YOLO algorithm's model training data processing in the application expansion of embedded platforms. Summary of the Invention

[0008] The embodiments of the present application provide a method and system for processing model training data based on an improved YOLO algorithm, thereby solving the problem in the prior art of insufficient adaptability of model training data processing of the YOLO algorithm to application expansion on embedded platforms, thereby improving the adaptability of the YOLO model to application expansion on embedded platforms.

[0009] An embodiment of the present application provides a model training data processing method based on an improved YOLO algorithm, comprising the following steps: inputting a predefined data set, automatically converting it into a predefined format through a predefined script, and extracting the original parameters of the YOLO algorithm configuration according to the recognition statistics of the predefined data set; performing evaluation and analysis based on the original parameters of the YOLO algorithm configuration to obtain the initial parameters of the YOLO algorithm configuration, inputting a training script based on the initial parameters of the YOLO algorithm configuration to perform data training, and obtaining a PC-side prediction and recognition model; converting the PC-side prediction and recognition model into an embedded application environment model to obtain an embedded-side prediction and recognition model, running the embedded-side prediction and recognition model on an embedded application platform and collecting the original scheduling parameters; updating the initial parameters of the YOLO algorithm configuration according to the YOLO algorithm configuration update model in combination with the original scheduling parameters, and reconfiguring the embedded-side prediction and recognition model according to the updated initial parameters of the YOLO algorithm configuration.

[0010] An embodiment of the present application provides a model training data processing system based on an improved YOLO algorithm, including a data conversion and raw data configuration module, a PC-side prediction and recognition model generation module, an embedded-side prediction and recognition model generation module, and an embedded-side prediction and recognition model update module: the data conversion and raw data configuration module is used to input a predefined data set, automatically convert it into a predefined format through a predefined script, and at the same time extract the original parameters of the YOLO algorithm configuration for recognition statistics of the predefined data set; the PC-side prediction and recognition model generation module is used to perform evaluation and analysis based on the original parameters of the YOLO algorithm configuration to obtain the initial parameters of the YOLO algorithm configuration, input the training script based on the initial parameters of the YOLO algorithm configuration to perform data training, and obtain the PC-side prediction and recognition model; the embedded-side prediction and recognition model generation module is used to convert the PC-side prediction and recognition model into an embedded application environment model to obtain an embedded-side prediction and recognition model, run the embedded-side prediction and recognition model on the embedded application platform and collect the original scheduling parameters; the embedded-side prediction and recognition model update module is used to update the initial parameters of the YOLO algorithm configuration according to the YOLO algorithm configuration update model in combination with the original scheduling parameters, and reconfigure the embedded-side prediction and recognition model according to the updated initial parameters of the YOLO algorithm configuration.

[0011] 1. This method extracts the original parameters of the YOLO algorithm from a predefined dataset for statistical identification and configuration; inputs the initial parameters into a training script for data training; converts the PC-side prediction and recognition model into an embedded application environment model; and reconfigures the embedded-side prediction and recognition model based on the updated initial parameters of the YOLO algorithm. This method improves the adaptability of the YOLO model to embedded applications, resolving the existing issue of insufficient adaptability of model training data processing for embedded applications.

[0012] 2. Real-time collection of operating parameters allows for dynamic adjustment of confidence thresholds and model complexity. Based on hardware load, the system automatically reduces resolution or filters low-confidence detection boxes. This lowers the confidence threshold to improve recall when the model is training densely populated objects, while increasing it during idle time to reduce false positives. By adjusting model parameters, the system avoids hardware throttling and maintains frame rate stability, ultimately enhancing the robustness of the YOLO algorithm's model training data processing method.

[0013] 3. Calculate the model input size, batch size, and learning rate based on parameters such as the category ratio and image entropy. By presetting the weighted rules for category ratios and entropy, the initial training parameters are dynamically generated to ensure that the model structure matches the data features. This allows the algorithm to automatically reduce the batch size and enhance data perturbations when recognizing specific images during specific algorithm execution, balancing accuracy and real-time performance. Furthermore, the input size is automatically limited based on the embedded platform's memory capacity to prevent memory overflow, thereby improving the stability of the YOLO algorithm model's performance in practical applications on embedded platforms. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 Flowchart of the model training data processing method based on the improved YOLO algorithm provided in an embodiment of the present application.

[0015] Figure 2 A software call diagram showing that the predefined data set provided in the embodiment of the present application is divided into two directories: a training set and a test set.

[0016] Figure 3 A schematic diagram of software calls for converting Voc data format into Yolo data format provided in an embodiment of the present application.

[0017] Figure 4 A schematic diagram of the process of generating the embedded-end prediction and recognition model provided in an embodiment of the present application.

[0018] Figure 5 A schematic diagram of the model conversion operation flow provided in an embodiment of the present application.

[0019] Figure 6 This is the hardware connection method between the PC and the embedded application platform provided in the embodiment of the present application.

[0020] Figure 7 A schematic diagram of the process of updating the embedded-side prediction and recognition model provided in an embodiment of the present application.

[0021] Figure 8 This is a structural diagram of the model training data processing system based on the improved YOLO algorithm provided in an embodiment of the present application. DETAILED DESCRIPTION

[0022] The embodiments of the present application provide a model training data processing method and system based on an improved YOLO algorithm, thereby solving the problem in the prior art of insufficient adaptability of model training data processing of the YOLO algorithm to application expansion on embedded platforms. Through data conversion and original data configuration, PC-side prediction and recognition model generation, embedded-side prediction and recognition model generation, and embedded-side prediction and recognition model update, the effect of improving the application expansion adaptability of the YOLO model on embedded platforms is achieved.

[0023] In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0024] like Figure 1 As shown, it is a flow chart of a model training data processing method based on an improved YOLO algorithm provided in an embodiment of the present application. The method is applied to a model training data processing system based on the improved YOLO algorithm, and the method includes the following steps: inputting a predefined data set, automatically converting it into a predefined format through a predefined script, and extracting the original parameters of the predefined data set to identify and statistically configure the YOLO algorithm; performing evaluation and analysis based on the original parameters of the YOLO algorithm configuration to obtain the initial parameters of the YOLO algorithm configuration, inputting a training script based on the initial parameters of the YOLO algorithm configuration to perform data training, and obtaining a PC-side prediction and recognition model; converting the PC-side prediction and recognition model into an embedded application environment model to obtain an embedded-side prediction and recognition model, running the embedded-side prediction and recognition model on the embedded application platform and collecting the original scheduling parameters; updating the initial parameters of the YOLO algorithm configuration according to the YOLO algorithm configuration update model in combination with the original scheduling parameters, and reconfiguring the embedded-side prediction and recognition model according to the updated initial parameters of the YOLO algorithm configuration.

[0025] In this embodiment, the existing embedded deployment process for object detection models suffers from the following core pain points: fragmented parameter configuration and a lack of unified standards. The parameters required for model training, exporting, conversion, and deployment are configured separately and independently, making unified scheduling difficult. For example, the complex relationship between imgsz and batch_size, and between quantization accuracy and platform computing power, requires manual experience to set. Existing systems are inefficient and prone to large errors.

[0026] Unable to perceive changes in external operating scenarios: The current mainstream deployment process cannot dynamically adjust inference parameters based on edge site conditions (such as occlusion rate, number of targets, and frame rate fluctuations), resulting in performance degradation or energy waste.

[0027] The post-deployment model tuning mechanism lacks a feedback loop: After ONNX is converted to RKNN, if you need to adjust the conf_threshold or switch to a lightweight model structure, you need to manually redeploy it, and it cannot be automatically updated based on feedback during runtime.

[0028] Cross-platform deployment lacks flexibility: Different devices (such as Orin Nano and RK3576) have significant differences in memory layout, computing power architecture, etc., and the existing configuration mechanism is difficult to universalize and often requires manual customization and modification.

[0029] This invention is particularly suitable for the following scenarios with strong demand for edge computing deployment: industrial inspection pipelines: workpiece type, size, and occlusion change in real time, requiring continuous dynamic optimization of recognition strategies; smart security monitoring: camera deployment environment lighting and occlusion frequently change, requiring real-time performance and accuracy; unmanned retail terminals: different user behavior patterns lead to complex detection target states, requiring real-time optimization of model parameters; in-vehicle vision systems: resource-constrained, requiring automatic switching of model complexity and inference configuration based on driving status. It should be noted that the above are only examples and do not limit the scope of application of this technology.

[0030] Dataset preparation: Before training YOLOv11, prepare the training data, such as VOC2007, and then divide the VOC2007 data into two directories: training set and test set. Figure 2 As shown, it is a software call schematic diagram of the predefined data set provided in the embodiment of the present application being divided into two directories: training set and test set.

[0031] Automatically convert to predefined format through predefined scripts: After the predefined dataset is prepared, use the data / voc_2_yolo.py script to convert the Voc data format to the Yolo data format. The converted data is stored in the yolo_data directory at the same level as the original data, such as Figure 3 As shown, it is a software call diagram for converting Voc data format into Yolo data format provided by an embodiment of the present application.

[0032] Furthermore, the original parameters of the YOLO algorithm configuration are extracted from the predefined data set to identify and count them, specifically including: using the predefined tool software to parse the label file of the predefined data set, counting and extracting the total number of different category identifiers; using the predefined tool software to parse the image grayscale of the predefined data set, analyzing and calculating the two-dimensional entropy of the image; extracting the standard value of the total number of preset category identifiers, the standard value of the preset two-dimensional entropy value of the image, the PC CPU running memory capacity value, the YOLO algorithm training model scale correction factor, the model learning rate adjustment coefficient, and the historical average value of the model initial training batch size from the YOLO algorithm model training database; extracting the number of convolution layers, the number of channels, the kernel size and the total number of parameters from the preset model parameters of the YOLO algorithm through the predefined software call, coupling the number of convolution layers, the number of channels and the kernel size, and then coupling them with the total number of parameters to obtain the complexity constant of the YOLO algorithm training model.

[0033] In this embodiment, predefined tool software parses the label file of a predefined dataset, such as VOC XML, and uses the set() function in a Python script to count the total number of different class identifiers (classID). For example, the set() function is used in a Python script to remove duplicate counts of categories in all labels. This is used to influence the last layer of the model structure (output dimension), and also affects the training complexity, image resolution recommendation value, etc.

[0034] Image two-dimensional entropy, based on grayscale distribution, incorporates neighborhood spatial features (such as neighborhood mean or co-occurrence matrix). This method describes the grayscale value of a pixel i and its neighborhood mean j using a two-tuple (i, j), reflecting spatial correlation. The image is traversed, recording the grayscale value of each pixel i and its neighborhood mean j to generate a two-dimensional frequency matrix. This method can distinguish textures with the same grayscale but different spatial distributions (e.g., checkerboard and random noise).

[0035] The following is an example of obtaining the two-dimensional entropy of an image using Python:

[0036] def calculate_entropy(image):

[0037] gray=cv2.cvtColor(image,cv2.COLOR_BGR2GRAY)

[0038] hist=cv2.calcHist([gray],[0],None,

[256] ,[0,256])

[0039] prob=hist / (gray.shape[0]gray.shape[1])

[0040] entropy=-np.sum(prob np.log2(prob+1e-10))

[0041] return entropy

[0042] MAG indicates that the model initial image sets the input size.

[0043]

[0044] Indicates the rounding symbol to ensure that the image input size is an integer.

[0045] The model's initial image input size is used in the configuration parameters required for subsequent YOLO algorithm model training. Specifically, it can be applied to the model structure configuration file, for example: yolo11.yaml.

[0046] SD represents the preset image input size constant; it ensures that the input size is a multiple of the grid requirements of the subsequent YOLO algorithm and meets the grid requirements of the subsequent YOLO algorithm. It can be set to 16, 32, 64, etc. according to the specific YOLO algorithm image training requirements and can be set by relevant personnel.

[0047] CC represents the preset base image size constant; it ensures that the minimum input image size will be at least the preset base image size. It can be set by the relevant personnel based on the computing power of the embedded platform.

[0048] BS represents the total number of different category identifiers in the predefined dataset; PBS represents the standard value of the total number of preset category identifiers, which is extracted from the YOLO algorithm model training database;

[0049] The training images in the predefined dataset are numbered, X0 represents the number of training images in the predefined dataset, X represents the total number of training images in the predefined dataset, and X0 = 1, 2, ..., X.

[0050] It represents the image two-dimensional entropy value of the X0th image in the predefined data set; PSZ represents the preset image two-dimensional entropy standard value, which is extracted from the YOLO algorithm model training database.

[0051] Furthermore, an evaluation analysis is performed based on the original parameters of the YOLO algorithm configuration to obtain the initial parameters of the YOLO algorithm configuration, specifically including: performing a ratio analysis on the total number of different category identifiers and the preset standard value of the total number of category identifiers to obtain the total number score of the category identifiers; performing a ratio analysis on the different image two-dimensional entropy values ​​and the preset standard value of the two-dimensional entropy values ​​and then summing and averaging them to obtain the total number score of the two-dimensional entropy values; performing logarithmic processing on the coupling result of the total number score of the category identifiers and the total number score of the two-dimensional entropy values, and then coupling them with the preset image base size constant and the preset image setting input size constant respectively to obtain the model initial image setting input size; coupling analysis is performed on the PC CPU running memory capacity value and the batch size preset constant to obtain the maximum undetermined component of the initial batch; coupling analysis is performed on the YOLO algorithm training model complexity constant, the exponential analysis result of the model initial image setting input size and the YOLO algorithm training model scale correction factor to obtain the initial batch balanced undetermined component; performing a ratio analysis on the initial batch maximum undetermined component and the initial batch balanced undetermined component to obtain the model initial training batch size; and configuring the corresponding model parameters according to the model initial image setting input size and the model initial training batch size.

[0052] In this embodiment, after the data conversion is completed, the training parameters of the model are configured: data.yaml, default.yaml, yolo11.yaml. Among them: data.yaml: is the path of the data to be trained and the verification data, as well as the number of categories and category names; default.yaml: is the yolo11 training parameter, and the parameters of the model training can be adjusted by yourself; Yolo11.yaml: is the yolo11 model structure, and the number of categories can be modified during model training.

[0053]

[0054] Indicates the rounding symbol to ensure that the image input size is an integer.

[0055] BAT represents the initial training batch size of the model, that is, the number of samples required for each model training. The model's initial image input size is used as a configuration parameter for subsequent YOLO algorithm model training. Specifically, it can be applied to the training hyperparameter configuration file, such as default.yaml.

[0056] YC represents the dimensionless value of the PC CPU running memory capacity. The PC CPU running memory capacity is directly extracted from the corresponding manufacturer chip manual in the YOLO algorithm model training database.

[0057] YCCS represents a preset constant for the batch size, ensuring that the initial batch size meets the computing power and running resource requirements of the subsequent YOLO algorithm and meets the network requirements of the subsequent YOLO algorithm. Depending on the specific YOLO algorithm image training requirements, it can be set to 256, 512, 1024, etc., and can be set by relevant personnel.

[0058] K represents the complexity constant of the YOLO algorithm training model; it is calculated by calling the YOLO algorithm preset model parameters in the model structure file (yolov11.yaml) in Python, K = ∑ (number of convolutional layers × number of channels × kernel size) / total number of parameters; for example, the following reference values: k≈0.35 for YOLOv11n and k≈1.2 for YOLOv11x.

[0059] MG represents the scale correction factor of the YOLO algorithm training model, which is extracted from the YOLO algorithm model training database. The YOLO algorithm scale types are classified as follows: 1. Nano, minimalist structure, less than 5M parameters; 2. Small, balanced speed and accuracy, 5-10M parameters; 3. Large, high accuracy, 10-30M parameters; 4. X-Large, extra-large model. The mapping relationship between the YOLO algorithm scale type and the YOLO algorithm training model scale correction factor in the YOLO algorithm model training database is obtained by inputting the known YOLO algorithm scale type to obtain the corresponding YOLO algorithm training model scale correction factor. The specific mapping relationship examples include |nano|0.25|, |small|0.50|, |large|1.00|, and |x-large|1.50|.

[0060] Furthermore, the initial parameters of the YOLO algorithm configuration are obtained by performing an evaluation analysis based on the original parameters of the YOLO algorithm configuration, and the method also includes: obtaining the total data volume of the model training by performing statistics on the predefined data set, performing a proportion analysis on the total data volume of the model training and a predefined constant, and then obtaining the model initial learning rate parameter operation component by exponential analysis; performing a square root analysis on the model initial training batch size, and then performing a proportion analysis on the YOLO algorithm training model scale correction factor to obtain the model initial learning rate training scale component; coupling the model learning rate adjustment coefficient and the model initial learning rate parameter operation component and the model initial learning rate training scale component in turn to obtain the model initial learning rate; performing a proportion analysis on the model data enhancement strength adjustment first coefficient and the total data volume of the model training and a predefined constant, and then performing a proportion analysis on the result of the square root analysis to obtain the model data enhancement strength data. The target average occlusion rate is corrected and analyzed by adjusting the second coefficient of the model data enhancement intensity to obtain the model data enhancement intensity complexity component; the total score of the two-dimensional entropy value is corrected and analyzed by adjusting the third coefficient of the model data enhancement intensity to obtain the model data enhancement intensity variability component; the model initial training batch size and the historical average value of the model initial training batch size are analyzed for proportion, and then the fourth coefficient is adjusted by the model data enhancement intensity after logarithmic analysis to obtain the model data enhancement intensity batch component; the model data enhancement intensity data volume component, the model data enhancement intensity complexity component and the model data enhancement intensity variability component are coupled and analyzed, and then the difference analysis is performed with the model data enhancement intensity batch component to obtain the model initial data enhancement intensity; the corresponding model parameters are configured according to the model initial learning rate and the model initial data enhancement intensity.

[0061] In this embodiment,

[0062] LRO represents the model's initial learning rate, a core hyperparameter that controls the parameter update step size during model training. Essentially, it determines how quickly the model "learns" from data and directly determines whether the model can converge to the optimal solution and the efficiency of convergence. This can be applied to training hyperparameter configuration files, such as default.yaml.

[0063] ZXS represents the total amount of data for model training, which is dimensionless. It can be the total number of pixels in all images in the training set, extracted from the YOLO algorithm model training database.

[0064] λ represents the model learning rate adjustment coefficient, which is dimensionless and is set here to the average learning rate of historical training data, that is, 0.01. In the following text, updates are triggered after actual model training.

[0065] Standardize the training data volume to a reasonable range. The training scale for classic datasets like COCO and VOC typically ranges from 10,000 to 100,000, with 10,000 being used as a reference. The computing power limitations of edge devices (such as the RK3576) require that the enhancement strength decay moderately as the data volume increases.

[0066] Controls the rate at which the enhancement strength decreases as the amount of data increases. A negative exponent indicates that the greater the amount of data, the lower the enhancement strength. This means that the enhancement strength is inversely proportional to the square root of the data volume, consistent with the statistical principle that the need for enhancement decreases with larger amounts of data. Experimental verification: In ablation experiments with YOLOv3 / v5, an exponent of -0.5 achieved the most significant improvement in validation set accuracy (approximately 2% to 3%).

[0067]

[0068] ZQD represents the model's initial data augmentation strength. Specifically, during the training of object detection models (such as YOLOv11), data augmentation strength refers to the severity of the random transformations applied to training images, regulating the model's adaptability to data diversity. Different augmentation strengths directly impact the model's generalization and robustness. This can be applied to training hyperparameter configuration files, such as default.yaml or hyp.yaml. When deployed on actual embedded edge computing devices, high augmentation strength during training increases model complexity and increases the computing resource consumption of the embedded edge computing device.

[0069] VZ represents the target average occlusion rate, which is dimensionless and ranges from 0 to 1. This is the initial value. The average target average occlusion rate in the historical dataset is taken as 0.3. The update will be triggered after the actual model training in the following text.

[0070] BATZ represents the historical average of the model's initial training batch size, which is extracted from the YOLO algorithm model training database.

[0071] Small datasets require high-intensity augmentation to compensate for data deficiencies and prevent overfitting. Large datasets, because their original distribution is already rich, require reduced augmentation to avoid introducing excessive noise. This means that as the total amount of data for model training increases, the augmentation intensity should be reduced.

[0072] When the average occlusion rate of the target is high, the intensity of the simulated occlusion needs to be increased to improve the model's anti-occlusion ability; however, excessive occlusion should be avoided to cause label noise, that is, to increase the intensity of model data enhancement.

[0073] The higher the two-dimensional entropy value of the image, the higher the variability of the training image. High-variability scenes require enhanced color perturbations and geometric transformations, that is, to increase the strength of model data enhancement.

[0074] When training with large batches, the gradient direction is more stable, so the augmentation intensity needs to be lowered to reduce noise interference. However, with small batches, due to severe gradient oscillation, high-intensity augmentation is required to improve generalization. In other words, the smaller the initial training batch size of the model, the higher the model data augmentation intensity needs to be.

[0075] η1 represents the first coefficient for adjusting the model data enhancement strength; η2 represents the second coefficient for adjusting the model data enhancement strength; η3 represents the third coefficient for adjusting the model data enhancement strength; and η4 represents the fourth coefficient for adjusting the model data enhancement strength. The complexity of enhancement operations (such as rotation and scaling) increases exponentially with strength. The two-dimensional entropy of the image needs to work in tandem with the enhancement strength, and both together amplify the computational load. Occlusion operations such as random erasure slow down their computing power requirements as the target average occlusion rate increases due to sparsity optimization. The total amount of model training data and the size of the initial model training batch will cause data loading and batch processing to occupy memory bandwidth, but the consumption of computing cores is relatively fixed. Based on the pre-run model, the ratios of the above four adjustment coefficients are adjusted to construct a mapping relationship between the actual operating chip computing power limit and the first, second, third, and fourth coefficients for adjusting the model data enhancement strength. The actual operating chip computing power limit is input to obtain the corresponding first, second, third, and fourth coefficients for adjusting the model data enhancement strength.

[0076] Furthermore, the initial parameters of the YOLO algorithm are configured and input into the training script for data training, specifically including: inputting the model with completed parameter configuration into the corresponding path through the predefined script to start model training; after the model training is completed, saving the model training process in the project directory configured by the predefined file, and testing the model results with the validation set; opening the predefined script file, configuring the model address and the picture to be tested, notifying relevant personnel to inspect and confirm the model processing effect, and obtaining the PC-side prediction and recognition model after confirmation.

[0077] In this embodiment, after completing the above configuration steps, you can start training the model. Use the train.py script and enter the data.yaml, default.yaml, and yolo11.yaml paths. The following code snippet example:

[0078]

[0079]

[0080] After training is complete, save the training process and the model that performs best on the validation set in the project directory specified in the default.yaml file. You can also perform model predictions to initially evaluate the model's effectiveness. Open the predict.py script and configure the model directory and the image to be tested.

[0081] Furthermore, the PC-side prediction and recognition model is converted into an embedded application environment model, specifically including: executing a predefined first conversion script on the PC side to convert the PC-side prediction and recognition model into an open neural network exchange format model, generating a standardized intermediate representation file, loading the precompiled model conversion tool image under the virtual machine terminal through a predefined software tool, mapping hardware device permissions and working directories, and building an isolated model conversion tool environment; mounting the local conversion script and the dataset directory to a specified path in the container, generating a quantization dataset list file containing the calibration image path, loading the open neural network exchange format model, configuring input standardization parameters and the target hardware platform, deploying the generated embedded-side prediction and recognition model to the target hardware platform, and loading and executing the embedded-side prediction and recognition model through a dedicated runtime library.

[0082] In this embodiment, if Figure 4A flow chart of the generation of the embedded prediction and recognition model provided in the embodiment of the present application. The following steps are used to achieve efficient adaptation of the embedded platform: first, input a predefined data set (such as VOC2007), and use a predefined script to automatically convert the data format to the YOLO standard format. At the same time, key parameters are extracted from the data set labels and images, including the total number of target categories, image texture complexity (two-dimensional entropy), hardware performance indicators, etc., to provide data support for subsequent model configuration; then, a multi-dimensional analysis is performed based on the extracted original parameters, combined with preset model scale correction factors, complexity constants and other standard values, to dynamically calculate the initial training parameters (such as image input size, batch size, learning rate), to solve the problems in traditional methods. The problem of scattered parameter configuration and reliance on manual experience is solved. Subsequently, the PC-side model training is started using uniformly configured parameters to generate a recognition model that has passed preliminary verification, and the model is converted into an open neural network exchange format (such as ONNX) through a script to achieve cross-platform compatibility. The standardized model is further deployed to embedded hardware (such as the RK3576 chip) to collect dynamic parameters (such as target occlusion rate, frame rate fluctuation, and computing power load) in the real operating environment to form real-time feedback. Finally, combined with the scheduling parameters collected by the embedded end, the initial training configuration is automatically updated, the model structure and inference parameters are re-optimized, and closed-loop tuning is completed. This method achieves deep adaptation of model training parameters to embedded hardware characteristics and scene changes through unified definition of data conversion and parameters, coupling of training and deployment processes, and dynamic feedback mechanism. It effectively solves the problems of parameter fragmentation, inability to perceive environmental changes, and lack of cross-platform flexibility in traditional processes, and significantly improves the scalability and adaptability of the YOLO algorithm in edge computing scenarios such as industrial detection and smart security.

[0083] The pt model is the PC-side prediction and recognition model mentioned above. The onnx model is the open neural network exchange format model.

[0084] Execute export.py on the PC to convert the pt model into onnx, as shown in the following code snippet:

[0085]

[0086]

[0087] The onnx model needs to be converted to the rknn model before it can run on EASY-EAI-Orin-nano, so you need to first build the rknn-toolkit model conversion tool environment.

[0088] The virtual machine is used here because it is convenient to configure the application conversion environment independently on the PC side.

[0089] Move the downloaded docker image to the rknn-toolkit2 directory of the ubuntu20.04 virtual machine; open the terminal of the ubuntu20.04 virtual machine in this directory and execute the following command to load the model conversion tool docker image. The sample code is as follows:

[0090] docker load--input rknn-toolkit2-v2.3.0-cp38-docker.tar.gz

[0091] Execute the following command to enter the mirror bash environment. The sample code is as follows:

[0092] docker run-ti --privileged-v / dev / bus / usb: / dev / bus / usb rknn-toolkit2:2.3.0-cp38 / bin / bash

[0093] Enter "python" to load Python-related libraries and the rknn library. At this point, the model conversion tool environment is complete.

[0094] EASY EAIOrin-nano supports the evaluation and operation of models with the .rknn suffix. Common tensorflow, tensroflowlite, caffe, darknet, onnx and Pytorch models can be converted to rknn models through the toolkit tool. For models trained with other frameworks, they can also be converted to onnx models first and then to rknn models. Figure 5 The figure shows a schematic diagram of the model conversion operation flow provided in an embodiment of the present application.

[0095] Unzip yolov11_model_convert.tar.bz2 and quant_dataset.zip to the virtual machine. Execute the following command to map the workspace into the Docker image, where / home / developer / rknn-toolkit2 / model_convert is the workspace, / test is mapped to the Docker image, and / dev / bus / usb: / dev / bus / usb is used to map the USB to the Docker image. The sample code is as follows: docker run -ti --privileged -v / dev / bus / usb: / dev / bus / usb -v / home / developer / rknn-toolkit2 / model_convert: / test rknn-toolkit2:2.3.0-cp38 / bin / bash

[0096] The model conversion test demo consists of yolov11_model_convert and quant_dataset. yolov11_model_convert stores the software script, and quant_dataset stores the data required for the quantitative model.

[0097] In the Docker environment, switch to the model conversion working directory and execute gen_list.py to generate a quantized image list. The rknn_convert.py script performs int8 quantization by default. Place the onnx model yolov11 s.onnx in the yolov11_model_convert directory. (Subsequently, if the user updates the model, regenerate a new onnx and replace the corresponding onnx.) Execute the rknn_convert.py script to convert the model. At this point, the model is generated and can be run in both the rknn environment and the EASY EAIOrin-nano environment. Note that EASY EAIOrin-nano is the corresponding runtime format for the embedded application platform.

[0098] Load and execute through the dedicated runtime library can connect to EASY-EAI-Orin-nano through the adb interface, such as Figure 6 , which is the hardware connection method between the PC and the embedded application platform provided in the embodiment of the present application.

[0099] Next, you need to transfer the source code to the board through adb. First, switch directories and execute the following instructions. Log in to the board, switch to the routine directory, and perform the compilation operation. After the compilation is successful, switch to the executable program directory. After the test network is completed, exit the board environment and output the test image. The test results are consistent with the PC results before the model conversion, indicating that it has been successfully applied to the embedded application platform.

[0100] Furthermore, the embedded-end prediction and recognition model is run on the embedded application platform and the original scheduling parameters are collected, specifically including: obtaining multiple target detection frames in each image through parameter detection and extraction during the operation of the embedded-end prediction and recognition model, performing pairwise intersection and comparison detection on them, and then summing and averaging them to obtain a new target average occlusion rate; obtaining the area sizes of multiple target detection frames in each image through parameter detection and extraction during the operation of the embedded-end prediction and recognition model; extracting the preset standard value of the number of target detection frames in the image, the average value of the area of ​​the target detection frames in the image, the preset constant of the model learning rate adjustment coefficient, the basic value of the confidence threshold of the embedded-end prediction and recognition model, the maximum value of the initial training batch size of the model, and the maximum value of the input size of the initial image of the model from the YOLO algorithm model training database; and using the embedded platform-specific library to obtain the embedded application platform main control CPU utilization, the embedded application platform main control memory occupancy and the average parameters of the embedded-end prediction and recognition model accuracy change when the embedded-end prediction and recognition model is running.

[0101] Furthermore, the initial parameters of the YOLO algorithm configuration are updated according to the YOLO algorithm configuration update model in combination with the original scheduling parameters, specifically including: substituting the new target average occlusion rate into and analyzing to obtain the new model initial data enhancement strength, recorded as the model update data enhancement strength; performing a ratio analysis on the average size of the target detection frame area and the average area of ​​the preset image target detection frame to obtain the model learning rate target detection frame area component; performing a ratio analysis on the standard value of the number of preset image target detection frames and the number of target detection frames to obtain the model learning rate target detection frame area component; coupling the model learning rate target detection frame area component and the model learning rate target detection frame area component, then performing a sum analysis, and then coupling analysis with the preset constant of the model learning rate adjustment coefficient to obtain a new model learning rate adjustment coefficient; substituting the new model learning rate adjustment coefficient into and analyzing to obtain the new model initial learning rate, recorded as the model update learning rate; reconfiguring the model update data enhancement strength and the model update learning rate into the corresponding model parameters.

[0102] In this embodiment, the target detection box in each image is obtained by statistically analyzing the recognition results after the YOLO algorithm training model. The sample code is as follows:

[0103]

[0104]

[0105]

[0106] Average object occlusion rate. For each bounding box in each image, pairwise IOU calculations are performed. This is based on the IOU distribution between bounding boxes, evaluating the occlusion ratio between objects in the same image. This parameter directly affects the image resolution and data augmentation strategy required for model training.

[0107] IOU (Intersection over Union) measures the degree of overlap between two detection boxes. The formula is: IOU = Area(A∩B) / Area(A∪B), where A and B are the intersection and union areas of the two detection boxes. A higher IOU value indicates more severe inter-object occlusion. This is a quantification method for occlusion rate.

[0108] For a target A, calculate the maximum IOU with all other detection boxes, for example: OccA = max(IOU(A, Bi)) (Bi ≠ A); B represents all detection boxes that intersect with target A, and i represents the object number of all detection boxes that intersect with target A. If OccA > θ, A is considered occluded. θ represents the confidence threshold for determining the occlusion rate of a single target, which is set by the relevant personnel.

[0109] Count the ratio of all detection boxes in the image that satisfy IOU>θ to all detection boxes in the image, reflecting the occlusion complexity of the entire image, and obtain the average occlusion rate of the target;

[0110]

[0111] ZQDN represents the model update data enhancement strength. Here, during the embedded application platform's embedded prediction and recognition model check, data is collected and calculated again to update the model's initial data enhancement strength.

[0112]

[0113] λ represents the new model learning rate adjustment coefficient, which is dimensionless. The average learning rate of the historical training data is 0.01, which is extracted from the corresponding average value of the historical training data in the YOLO algorithm model training database.

[0114] It represents the average size of the target detection box area in the X0th image, calculated based on the number of pixels; MJB represents the average area of ​​the target detection box in the preset image, which is extracted from the corresponding average value of the historical training data in the YOLO algorithm model training database.

[0115] Indicates the number of target detection boxes in the X0th image; KDB represents the standard value of the number of preset image target detection boxes, which is extracted from the YOLO algorithm model training database.

[0116] Furthermore, the embedded-end prediction and recognition model is reconfigured according to the updated initial parameters of the YOLO algorithm configuration, specifically including: performing a ratio analysis between the model's initial training batch size and the model's initial training batch size maximum value, and then coupling analysis with the embedded-end prediction and recognition model accuracy change average parameter to obtain a confidence threshold enhancement component; performing a ratio analysis between the model's initial image setting input size and the model's initial image setting input size maximum value, and then coupling analysis with the model update data enhancement strength to obtain a confidence threshold reduction component; performing a ratio analysis between the coupling analysis results of the embedded application platform main control CPU utilization and the embedded application platform main control memory occupancy and a preset fixed value to obtain an embedded application platform main control operation index, performing a difference analysis between the unit one and the embedded application platform main control operation index to obtain a confidence threshold ratio component; performing a ratio analysis between the confidence threshold enhancement component and the confidence threshold reduction component, and then coupling analysis with the confidence threshold base value and the confidence threshold ratio component of the embedded-end prediction and recognition model to obtain a confidence threshold; the real-time configuration of the confidence threshold is automatically set in a predefined configuration file through the command line in real time, and the embedded-end prediction and recognition model is updated.

[0117] In this embodiment, if Figure 7The figure shows a flow chart of the embedded-side prediction and recognition model update process provided by the embodiment of the present application. The embedded-side prediction and recognition model optimization process provided by the embodiment of the present application realizes the dynamic tuning of the model in the edge computing scenario through multi-stage collaboration: first, the trained PC-side prediction model is converted into a standardized ONNX format, and the computing power constraints of the embedded platform are adapted through quantization and compression technology; then it is deployed to the target hardware (such as the RK3576 chip), and key parameters such as the detection frame intersection-union ratio distribution, target number fluctuations, and CPU / memory resource occupancy are collected in real time during operation; based on the collected data, the characteristics of environmental changes are dynamically analyzed - the scene complexity is evaluated by calculating the average occlusion rate between targets, and the hardware resource load status is combined to automatically adjust Data augmentation is used to balance model robustness with computational overhead. The system leverages the difference between detection box area and a preset benchmark to drive a dynamic learning rate adjustment mechanism, ensuring efficient model convergence. Furthermore, real-time accuracy offset and hardware resource utilization metrics are integrated to construct an adaptive confidence threshold strategy. When detection accuracy decreases or resource constraints are exceeded, the threshold is raised to filter out low-quality detection boxes to reduce overhead; conversely, the threshold is relaxed to improve recall. Finally, a hot update mechanism is configured to instantly inject optimized parameters into running model instances, forming a closed-loop "deployment-monitoring-analysis-tuning" control loop, enabling adaptive operation of embedded models without manual intervention. This process, through four core steps: quantitative deployment, environmental awareness, parameter linkage, and dynamic loading, effectively addresses the mismatch between rigid model configurations and the dynamic nature of edge environments in traditional approaches. This significantly improves system stability and detection accuracy in real-time-critical scenarios such as industrial quality inspection and intelligent security. The confidence threshold (conf_threshold) ranges from 0.1 to 0.9. Configuration file: predict.py or inference script parameters. Configuration method: Passing command-line parameters or code.

[0118] The embedded-end prediction and recognition model operation monitoring time periods are numbered, MJ0 represents the number of the embedded-end prediction and recognition model operation monitoring time period, X represents the total number of the embedded-end prediction and recognition model operation monitoring time periods, MJ0 = 1, 2, ..., MJ.

[0119]

[0120] It represents the confidence threshold of the embedded-end prediction and recognition model during the MJ0th embedded-end prediction and recognition model operation monitoring period. CONFB represents the basic confidence threshold value of the embedded-end prediction and recognition model, which is generally 0.1. The confidence threshold value ranges from 0.1 to 0.9. Therefore, the constant 0.9 and the basic confidence threshold value 0.1 need to be given in the minimization objective function in the formula. It is extracted from the YOLO algorithm model training database.

[0121] This parameter represents the average change in the accuracy of the embedded-side predictive recognition model during the MJ0th embedded-side predictive recognition model monitoring period. This parameter is dimensionless and ranges from 0 to 1. The deviation of the current model's accuracy from the baseline model reflects the impact of environmental changes or hardware downtime on detection results. A pre-established accuracy baseline is obtained through trial training on a fixed dataset (such as the COCO validation set) and can be set by the average of historical data. By deploying a lightweight verification module (such as the MobileNet classifier) ​​on the embedded platform and verifying the accuracy with real-time sampling test results, the average change parameter of the embedded-side predictive recognition model's accuracy over the time period is obtained.

[0122] Indicates the initial training batch size of the model during the MJ0th embedded-end prediction and recognition model operation monitoring period; BATMA indicates the maximum value of the initial training batch size of the model that can be input on the current embedded platform, which is extracted from the YOLO algorithm model training database.

[0123] Indicates the model update data enhancement strength during the MJ0th embedded-end prediction and recognition model operation monitoring period.

[0124] MAG indicates the model initial image setting input size, and MAGMA indicates the maximum input size of the model initial image setting that can be input on the current embedded platform, which is extracted from the YOLO algorithm model training database.

[0125] Read data in real time from Linux files such as / proc / stat and / proc / meminfo. Use embedded platform-specific libraries (such as Jetson's nvpmodel) to obtain the corresponding data. To avoid performance jitter caused by frequent reading, a sampling period of 1 second or longer is recommended.

[0126] Indicates the CPU utilization of the embedded application platform during the monitoring period of the MJ0th embedded-end prediction and recognition model operation.

[0127] Indicates the main control memory usage of the embedded application platform during the MJ0th embedded-end prediction and recognition model running monitoring period.

[0128] The batch size directly affects the video memory usage and parallel computing efficiency. A larger batch size can improve throughput but increase the single-frame processing delay. It is necessary to balance the number of detections and real-time performance by adjusting the confidence threshold.

[0129] When the batch size increases, the confidence threshold needs to be increased to filter out low-quality detection boxes and reduce the computational load.

[0130] High resolution (e.g., 1280×1280) improves small object detection accuracy, but the computational complexity increases quadratically. Low resolution requires lowering the confidence threshold to compensate for detail loss. As the input image size decreases, the confidence threshold should be lowered to maintain recall.

[0131] Data augmentation strength (such as Mosaic and MixUp) improves model generalization, requiring a lower confidence threshold to capture more potential targets. For weak augmentation, raising the confidence threshold reduces false positives. Increasing data augmentation strength increases the model's robustness to noise, allowing for a lower confidence threshold.

[0132] like Figure 8 As shown, it is a structural diagram of the model training data processing system based on the improved YOLO algorithm provided in the embodiment of the present application. The model training data processing system based on the improved YOLO algorithm provided in the embodiment of the present application includes a data conversion and raw data configuration module, a PC-side prediction and recognition model generation module, an embedded-side prediction and recognition model generation module and an embedded-side prediction and recognition model update module: the data conversion and raw data configuration module is used to input a predefined data set, automatically convert it into a predefined format through a predefined script, and extract the predefined data set recognition statistics YOLO algorithm configuration raw parameters; the PC-side prediction and recognition model generation module is used to configure the raw parameters according to the YOLO algorithm The data is evaluated and analyzed to obtain the initial parameters of the YOLO algorithm configuration, and the training script is input according to the initial parameters of the YOLO algorithm configuration for data training to obtain the PC-side prediction and recognition model; the embedded-side prediction and recognition model generation module is used to convert the PC-side prediction and recognition model into an embedded application environment model to obtain the embedded-side prediction and recognition model, run the embedded-side prediction and recognition model on the embedded application platform and collect the original scheduling parameters; the embedded-side prediction and recognition model update module is used to update the YOLO algorithm configuration initial parameters according to the YOLO algorithm configuration update model in combination with the original scheduling parameters, and reconfigure the embedded-side prediction and recognition model according to the updated YOLO algorithm configuration initial parameters.

[0133] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0134] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A system that specifies the functions of a box or boxes.

[0135] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction system that is implemented in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0136] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0137] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0138] Obviously, those skilled in the art may make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if such modifications and variations fall within the scope of the claims and their equivalents, the present invention is intended to include such modifications and variations.

Claims

1. A model training data processing method based on an improved YOLO algorithm, characterized in that: The following steps are involved: Input a predefined data set, automatically convert it to a predefined format through a predefined script, and extract the predefined data set to identify statistics and configure the original parameters of the YOLO algorithm; According to the original parameters of the YOLO algorithm, an evaluation analysis is performed to obtain the initial parameters of the YOLO algorithm. According to the initial parameters of the YOLO algorithm, the training script is input to perform data training to obtain the PC-side prediction and recognition model; Convert the PC-side prediction and recognition model into an embedded application environment model to obtain an embedded-side prediction and recognition model, run the embedded-side prediction and recognition model on the embedded application platform and collect the original scheduling parameters; The YOLO algorithm configuration is updated based on the original scheduling parameters and the YOLO algorithm configuration initial parameters are updated. The embedded prediction and recognition model is reconfigured based on the updated YOLO algorithm configuration initial parameters.

2. The model training data processing method based on the improved YOLO algorithm as claimed in claim 1, characterized in that: The extraction of the predefined data set recognition statistics YOLO algorithm configuration original parameters specifically includes: Use predefined tool software to parse the label file of the predefined data set, count and extract the total number of different category identifiers; The predefined tool software analyzes the image grayscale of the predefined data set and calculates the two-dimensional entropy of the image; Extract the preset standard value of the total number of category identifiers, the preset standard value of the two-dimensional entropy value of the image, the PC CPU running memory capacity value, the YOLO algorithm training model scale correction factor, the model learning rate adjustment coefficient, and the historical average value of the model initial training batch size from the YOLO algorithm model training database; The number of convolutional layers, channels, kernel size, and total parameters in the YOLO algorithm preset model parameters are extracted through predefined software calls. The number of convolutional layers, channels, and kernel size are coupled and analyzed, and then coupled with the total parameter value to obtain the complexity constant of the YOLO algorithm training model.

3. The model training data processing method based on the improved YOLO algorithm as claimed in claim 2, characterized in that: The initial parameters of the YOLO algorithm configuration are obtained by performing evaluation and analysis based on the original parameters of the YOLO algorithm configuration, specifically including: Analyze the proportion of the total number of different category identifiers to the preset standard value of the total number of category identifiers to obtain the total number score of the category identifiers; The two-dimensional entropy values ​​of different images are compared with the preset standard values ​​of the two-dimensional entropy values ​​of the images, and then the sum and average are calculated to obtain the total score of the two-dimensional entropy values; Performing logarithmic processing on the coupling result of the total score of the category identifiers and the total score of the two-dimensional entropy values, and then coupling it with the preset image base size constant and the preset image setting input size constant respectively to obtain the model initial image setting input size; The maximum undetermined component of the initial batch is obtained by coupling the CPU running memory capacity value of the PC with the preset constant of the batch size; The YOLO algorithm training model complexity constant, the exponential analysis results of the model initial image input size setting, and the YOLO algorithm training model scale correction factor are coupled and analyzed to obtain the initial batch balance undetermined components; Analyze the proportion of the maximum undetermined component of the initial batch and the balanced undetermined component of the initial batch to obtain the initial training batch size of the model; Configure the corresponding model parameters according to the model's initial image input size and the model's initial training batch size.

4. The model training data processing method based on the improved YOLO algorithm as claimed in claim 3, characterized in that: The step of performing evaluation and analysis based on the original parameters of the YOLO algorithm configuration to obtain the initial parameters of the YOLO algorithm configuration further includes: The total amount of model training data is obtained by statistically analyzing the predefined data set, and the ratio of the total amount of model training data to the predefined constant is analyzed. Then, the initial learning rate parameter operation component of the model is obtained through exponential analysis. Perform square root analysis on the initial training batch size of the model, and then perform ratio analysis on it with the YOLO algorithm training model scale correction factor to obtain the model initial learning rate training scale component; The model learning rate adjustment coefficient, the model initial learning rate parameter operation component and the model initial learning rate training scale component are coupled and analyzed in sequence to obtain the model initial learning rate; The first coefficient of the model data enhancement intensity adjustment is analyzed with the total model training data volume and the predefined constant, and then the results of the square root analysis are analyzed to obtain the model data enhancement intensity data volume component; The target average occlusion rate is corrected and analyzed by adjusting the second coefficient of the model data enhancement strength to obtain the model data enhancement strength complexity component; The total score of the two-dimensional entropy value is modified and analyzed by adjusting the third coefficient of the model data enhancement intensity to obtain the model data enhancement intensity variability component; The model initial training batch size and the historical average of the model initial training batch size are analyzed, and then the fourth coefficient of the model data enhancement strength is adjusted through logarithmic analysis to obtain the model data enhancement strength batch component; The model data enhancement strength data volume component, model data enhancement strength complexity component and model data enhancement strength variability component are coupled and analyzed, and then the difference analysis is performed with the model data enhancement strength batch component to obtain the model initial data enhancement strength; Configure the corresponding model parameters according to the model's initial learning rate and the model's initial data enhancement strength.

5. The model training data processing method based on the improved YOLO algorithm as claimed in claim 1, characterized in that: The initial parameters of the YOLO algorithm are configured and input into the training script for data training, specifically including: Input the model with completed parameter configuration into the corresponding path through the predefined script to start model training; After the model training is completed, save the model training process in the project directory configured by the predefined file, and test the model results on the validation set; Open the predefined script file, configure the model address and the image to be tested, notify relevant personnel to inspect and confirm the model processing effect, and obtain the PC-side prediction and recognition model after confirmation.

6. The model training data processing method based on the improved YOLO algorithm as claimed in claim 1, characterized in that: The conversion of the PC-side prediction and recognition model into an embedded application environment model specifically includes: Execute the predefined first conversion script on the PC to convert the PC prediction and recognition model into an open neural network exchange format model, generate a standardized intermediate representation file, and use predefined software tools to load the precompiled model conversion tool image in the virtual machine terminal, map the hardware device permissions and working directory, and build an isolated model conversion tool environment; Mount the local conversion script and dataset directory to the specified path in the container, generate a quantization dataset list file containing the calibration image path, load the open neural network exchange format model, configure the input standardization parameters and target hardware platform, deploy the generated embedded-end prediction and recognition model to the target hardware platform, and load and execute the embedded-end prediction and recognition model through the dedicated runtime library.

7. The model training data processing method based on the improved YOLO algorithm as claimed in claim 4, characterized in that: The embedded-end prediction and recognition model is run on the embedded application platform and the original scheduling parameters are collected, specifically including: Through parameter detection and extraction during the runtime of the embedded predictive recognition model, multiple target detection frames in each image are obtained, and pairwise intersection and comparison detection are performed on them. The sum and average are then taken to obtain the new target average occlusion rate. By extracting parameters during the runtime of the embedded predictive recognition model, the area sizes of multiple target detection boxes in each image are obtained. Extract from the YOLO algorithm model training database the standard value of the number of preset image target detection boxes, the average area of ​​the preset image target detection boxes, the preset constant of the model learning rate adjustment coefficient, the basic value of the confidence threshold of the embedded end prediction recognition model, the maximum value of the model initial training batch size, and the maximum value of the model initial image input size; Use the embedded platform-specific library to obtain the embedded application platform main controller CPU utilization, embedded application platform main controller memory occupancy and the average parameters of the embedded end prediction and recognition model accuracy change when the embedded end prediction and recognition model is running.

8. The model training data processing method based on the improved YOLO algorithm as claimed in claim 7, characterized in that: The updating of the YOLO algorithm configuration initial parameters according to the YOLO algorithm configuration update model in combination with the original scheduling parameters specifically includes: Substitute the new target average occlusion rate and analyze to obtain the new model initial data enhancement strength, which is recorded as the model update data enhancement strength; Analyze the proportion of the average size of the target detection box area and the average size of the preset image target detection box area to obtain the target detection box area component of the model learning rate; The standard value of the number of target detection frames in the preset image is compared with the number of target detection frames to obtain the target detection frame area component of the model learning rate. After coupling analysis of the model learning rate target detection box area component and the model learning rate target detection box area component, a sum analysis is performed, and then a coupling analysis is performed with the preset constant of the model learning rate adjustment coefficient to obtain a new model learning rate adjustment coefficient; Substitute the new model learning rate adjustment coefficient and analyze to obtain the new model initial learning rate, which is recorded as the model update learning rate; Reconfigure the model update data enhancement strength and model update learning rate and substitute them into the corresponding model parameters.

9. The model training data processing method based on the improved YOLO algorithm as claimed in claim 8, characterized in that: The reconfiguration of the embedded-end prediction and recognition model according to the updated YOLO algorithm initial parameters specifically includes: The initial training batch size of the model and the maximum initial training batch size of the model are analyzed in proportion, and then coupled with the average parameter of the embedded end prediction and recognition model accuracy change to obtain the confidence threshold enhancement component; The proportion of the model's initial image setting input size and the model's initial image setting input size maximum value are analyzed, and then coupled with the model's updated data enhancement strength to obtain the confidence threshold attenuation component; The coupled analysis results of the embedded application platform main control CPU utilization and the embedded application platform main control memory occupancy are analyzed with the preset fixed value to obtain the embedded application platform main control operation index. The difference analysis between the unit one and the embedded application platform main control operation index is performed to obtain the confidence threshold ratio component; The confidence threshold enhancement component and the confidence threshold reduction component are analyzed for their proportion, and then coupled with the confidence threshold base value and confidence threshold ratio component of the embedded prediction and recognition model to obtain the confidence threshold; The confidence threshold is automatically configured in real time into a predefined configuration file through the command line, and the embedded prediction recognition model is updated.

10. Model training data processing system based on improved YOLO algorithm, characterized in that: It includes data conversion and raw data configuration module, PC-side prediction and recognition model generation module, embedded-side prediction and recognition model generation module and embedded-side prediction and recognition model update module: Data conversion and raw data configuration module: used to input predefined data sets, automatically convert them into predefined formats through predefined scripts, and extract the predefined data sets to identify statistics and configure the original parameters of the YOLO algorithm; PC-side prediction and recognition model generation module: used to configure the original parameters of the YOLO algorithm for evaluation and analysis to obtain the initial parameters of the YOLO algorithm configuration, input the training script according to the initial parameters of the YOLO algorithm configuration for data training, and obtain the PC-side prediction and recognition model; Embedded-end prediction and recognition model generation module: used to convert the PC-end prediction and recognition model into an embedded application environment model to obtain the embedded-end prediction and recognition model, run the embedded-end prediction and recognition model on the embedded application platform, and collect the original scheduling parameters; Embedded-end prediction and recognition model update module: used to update the initial parameters of the YOLO algorithm configuration according to the YOLO algorithm configuration update model in combination with the original scheduling parameters, and reconfigure the embedded-end prediction and recognition model according to the updated initial parameters of the YOLO algorithm configuration.

Citation Information

Patent Citations

  • A neural network acceleration method for AI chips based on FPGA

    CN113392973B

  • Construction method of Yolo improved network model based on experiment teaching and examination

    CN118313427A