Embedded stepped operation platform and method
By deploying AI frameworks and tools on the embedded platform, data preprocessing and model training is carried out, the problem of insufficient software support for embedded platforms in AI training operations is solved, and flexible and efficient AI model training is achieved, reducing operation difficulty and improving teaching value.
Patent Information
- Application Number
- CN202510067517.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing embedded experimental platform has insufficient software support when performing AI training operations, resulting in compatibility and performance problems, and the operation threshold is high, making it difficult to meet the needs of beginners of AI models to quickly get started.
It provides an embedded ladder operation platform and method. By deploying AI frameworks, installing development tools, data set management tools, user management tools, and performing data preprocessing, model training and testing, modular design and collaborative work are realized, and a complete AI model training process is supported.
It realizes flexibility, efficiency, modular design, collaborative work, and low-cost AI model training solutions, reduces development complexity, improves training efficiency, and has teaching and research value.
Smart Images

Figure CN120012882A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of embedded technology, and in particular to an embedded-based stepped operating platform and method. Background Art
[0002] The embedded experimental platform is a comprehensive environment that integrates hardware, software and development tools, and is specifically used for the development, testing and verification of embedded systems. It usually includes one or more embedded processors (such as ARM, DSP, etc.), various peripheral interfaces (such as GPIO, UART, I2C, SPI, etc.), memory (such as SRAM, DRAM, Flash, etc.) and corresponding development software and tool chains. These components work together to provide users with a complete embedded system development environment.
[0003] The main features of the embedded experiment platform include:
[0004] Highly integrated: Hardware and software components are highly integrated, which facilitates users to quickly develop and test.
[0005] Scalability: The platform usually supports a variety of processors and peripherals, and users can choose and expand according to their needs.
[0006] Real-time: Embedded systems usually require real-time response, and the experimental platform can provide real-time operating system and development tool support.
[0007] Low cost: Compared with complete embedded product development, the experimental platform has a low cost and is suitable for teaching and beginners.
[0008] However, although embedded experimental platforms have performed well in embedded system development, there is a clear lack of software support when performing AI training operations. Although some embedded platforms provide support for AI algorithms and frameworks, this support is usually limited and not mature enough. This leads to compatibility and performance issues when performing AI training on embedded platforms. Traditional model training on GPU services not only has a high operating threshold and is difficult to get started, but also has a complicated operation process, which greatly increases the difficulty for AI enthusiasts and practitioners to get started. Therefore, under the current wave of AI model training, for teaching platforms, or AI model training enthusiasts who want to experience AI model training operations, traditional embedded-based platforms cannot quickly and conveniently meet the needs of AI model beginners to quickly get started with AI model training.
[0009] Therefore, in order to solve the above problems, an embedded stepped operation platform and method are needed. Summary of the invention
[0010] The purpose of the present invention is to provide an embedded-based step-by-step operating platform and method. The present invention brings many beneficial effects such as flexibility, high efficiency, modular design, collaborative work, reusability, integrated environment, real-time support, low cost and high efficiency, and teaching and research value through the operating method, system module and platform design of AI model training on the embedded platform. These effects jointly promote the development and application of embedded AI technology.
[0011] The present invention is achieved in that:
[0012] The present invention provides an embedded step-by-step operation method which is specifically performed in the following steps:
[0013] S1: On the embedded experimental operation platform, deploy and install an embedded Linux system that supports the AI framework, and install and configure the AI framework, including TensorFlow Lite or PyTorch Mobile;
[0014] S2: Install development tools, including Python compiler, GCC compiler, and CMake build tool, for writing, compiling, and debugging AI model training code;
[0015] S3: Install and use dataset management tools, including Pandas and NumPy, for dataset import, preprocessing, annotation, and storage;
[0016] S4: Install and use Git control tools to perform version management on codes and datasets to ensure the traceability and repeatability of codes and datasets;
[0017] S5: Install and configure user management tools, including / etc / passwd, / etc / group files, and adduser, usermod, and deluser commands for creating, modifying, and deleting user accounts;
[0018] S6: Create a user account. Specifically, use the adduser command to create a new user account and set the user password and home directory information; for example: sudo adduser username, and then set the password and related information according to the prompts; use the usermod command to modify user permissions, including adding users to specific groups and modifying user home directories, for example: sudo usermod -aGgroupname username, to add users to the specified group;
[0019] Use the deluser command to delete a user account, and you can choose whether to delete the user's home directory; for example: sudodeluser --remove-home username, delete the user and delete its home directory.
[0020] S7: The user logs into the system. The user uses his own account and password to log into the embedded experimental platform. After logging in, the user enters his own home directory and accesses and operates the files and programs in the directory.
[0021] S8: The user uses the dataset management tool to import the dataset and perform preprocessing operations, including data cleaning and labeling. After the preprocessing is completed, the user saves the dataset to a specified location and performs model training later.
[0022] Please follow the steps below:
[0023] S8.1: When importing a dataset, first determine the source of the dataset. The dataset comes from the local computer, a remote server on the network, a database, or an API interface. According to the format and source of the dataset, use the import tools pandas and numpy libraries in Python, readr and data.table libraries in R, and SQL statements. Read the dataset using the selected import tool, according to the corresponding syntax or API call, and use different reading functions or methods, including read_csv, read_excel, and read_sql to read the dataset into the computing environment.
[0024] S8.2: Data cleaning, missing value processing, delete samples containing missing values; use mean, median or other statistics to fill missing values; use linear interpolation or model-based interpolation methods to estimate missing values;
[0025] Outlier processing: detect and identify outliers, and process outliers based on business context or statistical analysis, including deletion and correction;
[0026] Duplicate data processing, deleting or aggregating duplicate data rows;
[0027] The mean formula is as follows:
[0028]
[0029] The data set X is X = {x1, x2, ..., x n}, where Xi is the i-th non-missing value sample;
[0030] Use the median to fill in missing values, as follows:
[0031]
[0032] Arrange the non-missing value samples x1, x2, x3...xn in ascending order. If n is an odd number, the median M is the first value; if n is even, the median M is the average of the th value and the th value after sorting. In the calculation of linear interpolation, given two data points (x1, y1) and (x2, y2), to estimate the missing value y at x (x1 < x < x2), the linear interpolation formula is as follows:
[0033]
[0034] S8.3: Perform data integration, which involves merging data from different data sources into a consistent data store to support analysis and modeling; use database join operations to merge multiple data tables into one; merge data from different data sources based on a common unique identifier, such as a primary key;
[0035] S8.4: Data standardization, which scales numerical features to a similar range, specifically including standardizing the data to a standard normal distribution with a mean of 0 and a variance of 1;
[0036] Data normalization, which scales numerical features to a fixed range, including scaling the features to the interval [0, 1];
[0037] Logarithmic transformation, which performs natural logarithm or Log(x + 1) transformation on the data to reduce the skewness of the data;
[0038] Feature extraction, which extracts new features from the original data through feature selection, principal component analysis PCA, and polynomial feature construction methods to enhance the performance of the model;
[0039] Principal component analysis specifically maps high-dimensional data to a lower-dimensional space through a linear transformation, as shown in the following formula:
[0040]
[0041] where, is the mean vector, xi is the sample vector, and m is the number of samples;
[0042] And, new features are generated by combining the original features through polynomial feature construction to increase the complexity of the model and capture more non-linear relationships; as shown in the following formula:
[0043] P(x) = a n x n + a n-1 x n-1 + … + a1x + a0
[0044] where, P(x) is the polynomial feature, x is the original feature, an, an-1, …, a1, a0 are coefficients, and n is the degree of the polynomial.
[0045] Feature extraction, extracting new features from the original data through feature selection, principal component analysis PCA, and polynomial feature construction methods. Specifically, the number of features is reduced by removing features with variances below the threshold. Features with small variances have small changes in the data set. The formula is as follows:
[0046]
[0047] Among them, σ 2 is the variance, x i is the eigenvalue, μ is the mean of the feature, and N is the number of samples.
[0048] S8.5: Feature selection, selecting the most relevant features to reduce the dimensionality of the dataset; selection is based on statistical tests, feature importance scores, or feature selection algorithms used during model training;
[0049] Then perform data dimensionality reduction and use principal component analysis (PCA) to convert high-dimensional data into low-dimensional representation; reduce the storage and computing costs of the data set while maintaining important information of the data set.
[0050] S9: The user uses the AI framework to write the model training code and configure the training parameters. After writing, the user uses the development tool to compile and run the model training code.
[0051] Please follow the steps below:
[0052] S9.1: Write model training code. Users use AI frameworks, including TensorFlow and PyTorch, to write model training code; define the model structure, including the input layer, hidden layer, and output layer, and configure training parameters, including learning rate, batch size, and number of iterations;
[0053] S9.2: Users use development tools, including Jupyter Notebook and VS Code, to compile and run model training code;
[0054] During the training process, the model continuously updates weights through the gradient descent optimization algorithm to minimize the loss function;
[0055] S9.3: During the training process, the user monitors the loss function and accuracy indicators of the model; specifically, the commonly used performance evaluation indicators include mean square error (MSE), mean absolute error (MAE), and accuracy. According to the monitoring results, the user can adjust the training parameters and optimize the model in a timely manner.
[0056] During the training process, users monitor the model's loss function and accuracy indicators, and adjust the training parameters and optimize the model in a timely manner;
[0057] The specific loss function is used to measure the difference between the model prediction results and the actual results; as follows:
[0058]
[0059] Among them, yi is the true value, is the predicted value, N is the number of samples;
[0060] And use stochastic gradient descent SGD to update the model parameters to minimize the loss function; as shown below;
[0061]
[0062] Among them, θ j are model parameters, α is the learning rate, and J(θ) is the loss function.
[0063] S10: Perform model testing and deployment. After training is completed, the user uses the test data set to test the model to evaluate the performance and accuracy of the model. After the test passes, the user exports the model to a format recognizable by the embedded device and deploys it to the embedded experimental platform.
[0064] S11: The model training results of the above steps are uniformly transmitted to the backend server for unified management and scoring.
[0065] Furthermore, the present invention provides an embedded-based step-by-step operation platform, including a data preparation module for data set collection and upload, where the user needs to collect and prepare the data set required for training and upload it to the embedded experimental platform, including various types of data such as images, texts, and audios;
[0066] Data preprocessing: After uploading the data set, data preprocessing is performed, including data cleaning, data enhancement, image flipping, rotation, and data formatting operations;
[0067] Model training module, used for model selection and configuration. Users need to select the appropriate model type according to the application scenario and data characteristics, including convolutional neural network CNN for image classification and recurrent neural network RNN for sequence prediction. At the same time, the model hyperparameters are configured through the model training module, including learning rate, batch size, number of iterations, etc.
[0068] After configuring the model, users can use the embedded experimental platform to perform model training.
[0069] Model evaluation and optimization module, which specifically performs model evaluation. After training is completed, the model is evaluated to determine whether its performance meets the requirements. Specifically, the model is run on the validation set or test set and the performance indicators are calculated;
[0070] The platform support and service module specifically manages users and provides user registration, login, and permission management;
[0071] And carry out project management. Users create and manage multiple projects on the platform support and service module. Each project can contain multiple data sets, models and training tasks.
[0072] Furthermore, the present invention provides a computer-storable medium, wherein the computer-readable storage medium includes an embedded processing system and a stored program, and controls the above-mentioned embedded-based step-by-step operation method when the embedded system control program runs.
[0073] Compared with the prior art, the present invention has the following beneficial effects:
[0074] 1. The platform of the present invention allows users to flexibly select AI models and training algorithms according to actual needs, as well as adjust model parameters and training strategies. This flexibility enables users to optimize model performance for specific application scenarios.
[0075] Efficiency: Through reasonable operation methods, users can efficiently perform AI model training tasks on the embedded experimental platform. The hardware acceleration and parallel computing capabilities provided by the platform can significantly shorten the training time and improve the training efficiency.
[0076] 2. Modular design: The system module design of the embedded experimental platform makes each component relatively independent, which is convenient for users to carry out modular development and testing. This helps to reduce development complexity and improve development efficiency.
[0077] Collaboration: The collaborative work between system modules enables the embedded experimental platform to support the complete AI model training process. From data preparation, model training to model evaluation and optimization, each module works closely together to ensure the smooth progress of the training task.
[0078] 3. The platform of the present invention provides a low-cost and efficient AI model training solution, which enables more users to afford the development and application of AI technology.
[0079] Teaching and research value: The embedded experimental platform also has important teaching and research value. It can be used as a teaching and scientific research tool to help students and researchers understand the principles and applications of embedded systems and AI technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It is understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0081] Figure 1 is a flow chart of the method of the present invention;
[0082] Figure 2 It is a system structure diagram of the present invention. DETAILED DESCRIPTION
[0083] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work belong to the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the invention claimed for protection, but is only for selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work belong to the scope of protection of the present invention.
[0084] See also Figure 1-Figure 2 , the present invention provides an embedded-based stepped operating platform and method;
[0085] S1: On the embedded experimental operation platform, deploy and install an embedded Linux system that supports the AI framework, and install and configure the AI framework, including TensorFlow Lite or PyTorch Mobile;
[0086] S2: Install development tools, including Python compiler, GCC compiler, and CMake build tool, for writing, compiling, and debugging AI model training code;
[0087] S3: Install and use dataset management tools, including Pandas and NumPy, for dataset import, preprocessing, annotation, and storage;
[0088] S4: Install and use Git control tools to perform version management on codes and datasets to ensure the traceability and repeatability of codes and datasets;
[0089] S5: Install and configure user management tools, including / etc / passwd, / etc / group files, and adduser, usermod, and deluser commands for creating, modifying, and deleting user accounts;
[0090] S6: Create a user account. Specifically, use the adduser command to create a new user account and set the user password and home directory information; for example: sudo adduser username, and then set the password and related information according to the prompts; use the usermod command to modify user permissions, including adding users to specific groups and modifying user home directories, for example: sudo usermod -aGgroupname username, to add users to the specified group;
[0091] Use the deluser command to delete a user account, and you can choose whether to delete the user's home directory; for example: sudodeluser --remove-home username, delete the user and delete its home directory.
[0092] S7: The user logs into the system. The user uses his / her own account and password to log into the embedded experimental platform. After logging in, the user enters his / her home directory and accesses and operates the files and programs in the directory.
[0093] S8: The user uses the dataset management tool to import the dataset and perform preprocessing operations, including data cleaning and labeling. After the preprocessing is completed, the user saves the dataset to a specified location and performs model training later.
[0094] Please follow the steps below:
[0095] S8.1: When importing a dataset, first determine the source of the dataset. The dataset comes from the local computer, a remote server on the network, a database, or an API interface. According to the format and source of the dataset, use the import tools pandas and numpy libraries in Python, readr and data.table libraries in R, and SQL statements. Read the dataset using the selected import tool, according to the corresponding syntax or API call, and use different reading functions or methods, including read_csv, read_excel, and read_sql to read the dataset into the computing environment.
[0096] S8.2: Data cleaning, missing value processing, delete samples containing missing values; use mean, median or other statistics to fill missing values; use linear interpolation or model-based interpolation methods to estimate missing values;
[0097] Outlier handling, detecting and identifying outliers, and handling outliers according to the business background or statistical analysis, specifically including deletion and correction;
[0098] Duplicate data handling, deleting or aggregating duplicate data rows;
[0099] The mean formula is as follows:
[0100]
[0101] Among them, the data set X is X = {x1, x2,..., x n} , where Xi is the i-th non-missing value sample;
[0102] Using the median to fill in missing values, specifically as follows:
[0103]
[0104] Arrange the non-missing value samples x1, x2, x3... xn in ascending order. If n is odd, the median M is the th value after sorting; if n is even, the median M is the average of the th value and the th value after sorting. In the calculation of linear interpolation, given two data points (x1, y1) and (x2, y2), to estimate the missing value y at x (x1 < x < x2), the linear interpolation formula is as follows:
[0105]
[0106] S8.3: Perform data integration, which involves merging data from different data sources into a consistent data store to support analysis and modeling; use database join operations to merge multiple data tables into one; merge data from different data sources according to a common unique identifier, such as a primary key;
[0107] S8.4: Data standardization, scaling numerical features to a similar range, specifically including standardizing the data to a standard normal distribution with a mean of 0 and a variance of 1;
[0108] Data normalization, scaling numerical features to a fixed range, including scaling the features to the interval [0, 1];
[0109] Logarithmic transformation, performing natural logarithm or Log(x + 1) transformation on the data to reduce the skewness of the data;
[0110] Feature extraction, extracting new features from the original data through feature selection, principal component analysis PCA, and polynomial feature construction methods to enhance the performance of the model;
[0111] Principal component analysis specifically maps high-dimensional data to a lower-dimensional space through linear transformation, as shown below:
[0112]
[0113] in, is the mean vector, xi is the sample vector, and m is the number of samples;
[0114] In addition, the original features are combined polynomially through polynomial feature construction to generate new features, increase the complexity of the model, and capture more nonlinear relationships; as shown in the following formula:
[0115] P(x)=a n x n +a n-1 x n-1 +…+a1x+a0
[0116] Among them, P(x) is the polynomial feature, x is the original feature, an, an-1,…, a1, a0 are coefficients, and n is the degree of the polynomial.
[0117] Feature extraction, extracting new features from the original data through feature selection, principal component analysis PCA, and polynomial feature construction methods. Specifically, the number of features is reduced by removing features with variances below the threshold. Features with small variances have small changes in the data set. The formula is as follows:
[0118]
[0119] Among them, σ 2 is the variance, x i is the eigenvalue, μ is the mean of the feature, and N is the number of samples.
[0120] S8.5: Feature selection, selecting the most relevant features to reduce the dimensionality of the dataset; selection is based on statistical tests, feature importance scores, or feature selection algorithms used during model training;
[0121] Then perform data dimensionality reduction and use principal component analysis (PCA) to convert high-dimensional data into low-dimensional representation; reduce the storage and computing costs of the data set while maintaining important information of the data set.
[0122] S9: The user uses the AI framework to write the model training code and configure the training parameters. After writing, the user uses the development tool to compile and run the model training code.
[0123] In this embodiment, the following steps are specifically performed:
[0124] S9.1: Write model training code. Users use AI frameworks, including TensorFlow and PyTorch, to write model training code; define the model structure, including the input layer, hidden layer, and output layer, and configure training parameters, including learning rate, batch size, and number of iterations;
[0125] S9.2: Users use development tools, including Jupyter Notebook and VS Code, to compile and run model training code;
[0126] During the training process, the model continuously updates weights through the gradient descent optimization algorithm to minimize the loss function;
[0127] S9.3: During the training process, the user monitors the loss function and accuracy indicators of the model; specifically, the commonly used performance evaluation indicators include mean square error (MSE), mean absolute error (MAE), and accuracy. According to the monitoring results, the user can adjust the training parameters and optimize the model in a timely manner.
[0128] During the training process, users monitor the model's loss function and accuracy indicators, and adjust the training parameters and optimize the model in a timely manner;
[0129] The specific loss function is used to measure the difference between the model prediction results and the actual results; as follows:
[0130]
[0131] Among them, yi is the true value, is the predicted value, N is the number of samples;
[0132] And use stochastic gradient descent SGD to update the model parameters to minimize the loss function; as shown below;
[0133]
[0134] Among them, θ j are model parameters, α is the learning rate, and J(θ) is the loss function.
[0135] S10: Perform model testing and deployment. After training is completed, the user uses the test data set to test the model to evaluate the performance and accuracy of the model. After the test passes, the user exports the model to a format recognizable by the embedded device and deploys it to the embedded experimental platform.
[0136] S11: The model training results of the above steps are uniformly transmitted to the backend server for unified management and scoring.
[0137] In this embodiment, the present invention provides an embedded-based step-by-step operation platform, including a data preparation module for data set collection and upload. Users need to collect and prepare data sets required for training and upload them to the embedded experimental platform, including various types of data such as images, texts, and audios;
[0138] Data preprocessing: After uploading the data set, data preprocessing is performed, including data cleaning, data enhancement, image flipping, rotation, and data formatting operations;
[0139] Model training module, used for model selection and configuration. Users need to select the appropriate model type according to the application scenario and data characteristics, including convolutional neural network CNN for image classification and recurrent neural network RNN for sequence prediction. At the same time, the model hyperparameters are configured through the model training module, including learning rate, batch size, number of iterations, etc.
[0140] After configuring the model, users can use the embedded experimental platform to perform model training.
[0141] Model evaluation and optimization module, which specifically performs model evaluation. After training is completed, the model is evaluated to determine whether its performance meets the requirements. Specifically, the model is run on the validation set or test set and the performance indicators are calculated;
[0142] The platform support and service module specifically manages users and provides user registration, login, and permission management;
[0143] And carry out project management. Users create and manage multiple projects on the platform support and service module. Each project can contain multiple data sets, models and training tasks.
[0144] In this embodiment, the present invention provides a computer storable medium, the computer readable storage medium includes an embedded processing system and a stored program, and controls the above-mentioned embedded-based step-by-step operation method when the embedded system control program is running. In this embodiment, the hardware platform of the present invention is configured with a high-performance processor or an acceleration module, wherein the processor adopts a high-performance processor, such as ARM Cortex-A series or NXP i.MX series, etc., to meet the needs of AI model training.
[0145] In addition, in this embodiment, the platform of the present invention adds an AI acceleration module, such as a GPU and an NPU, to increase the speed of AI model training and reasoning.
[0146] In this embodiment, the platform of the present invention adds memory bars of the type DDR3 or DDR4 to the memory and storage expansion, and adds high-speed storage devices such as SSD or eMMC to store more data sets and model files.
[0147] On the interface module, add a network interface: ensure that the embedded experiment platform has a Gigabit Ethernet interface or a Wi-Fi module to facilitate remote data transmission and model updates.
[0148] In addition, USB, SD card and other interfaces are added to connect external storage devices to facilitate the import and export of data sets.
[0149] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention has various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A step-by-step operation method based on embedded system, characterized in that: Please follow the steps below: S1: On the embedded experimental operation platform, deploy and install an embedded Linux system that supports the AI framework, and install and configure the AI framework, including TensorFlow Lite or PyTorch Mobile; S2: Install development tools, including Python compiler, GCC compiler, and CMake build tool, for writing, compiling, and debugging AI model training code; S3: Install and use dataset management tools, including Pandas and NumPy, for dataset import, preprocessing, annotation, and storage; S4: Install and use Git control tools to perform version management on codes and datasets to ensure the traceability and repeatability of codes and datasets; S5: Install and configure user management tools, including / etc / passwd, / etc / group files, and adduser, usermod, and deluser commands for creating, modifying, and deleting user accounts; S6: Create a user account, specifically use the adduser command to create a new user account, and set the user password and home directory information; Use the usermod command to modify user permissions, including adding users to groups and modifying user home directories; Use the deluser command to delete a user account, and you can choose whether to delete the user's home directory; S7: User logs into the system. The user uses his / her own account and password to log into the embedded experimental platform; After logging in, the user enters his or her home directory and accesses and operates the files and programs in that directory; S8: The user uses the dataset management tool to import the dataset and perform preprocessing operations, including data cleaning and labeling; After preprocessing is completed, the user saves the dataset to a specified location and then conducts model training; S9: The user uses the AI framework to write the model training code and configure the training parameters. After writing, the user uses the development tool to compile and run the model training code. During the training process, users monitor the model's loss function and accuracy indicators, and adjust the training parameters and optimize the model in a timely manner; S10: Perform model testing and deployment. After training is completed, the user uses the test data set to test the model to evaluate the performance and accuracy of the model. After the test passes, the user exports the model to a format recognizable by the embedded device and deploys it to the embedded experimental platform. S11: The model training results of the above steps are uniformly transmitted to the backend server for unified management and scoring.
2. The embedded step-by-step operation method according to claim 1, characterized in that: In step S8, the following steps are specifically performed: S8.1: When importing a dataset, first determine the source of the dataset. The dataset comes from the local computer, a remote server on the network, a database, or an API interface. According to the format and source of the dataset, use the import tools pandas and numpy libraries in Python, readr and data.table libraries in R, and SQL statements. Read the dataset using the selected import tool, according to the corresponding syntax or API call, and use different reading functions or methods, including read_csv, read_excel, and read_sql to read the dataset into the computing environment. S8.2: Data cleaning, missing value processing, delete samples containing missing values; use mean, median or other statistics to fill missing values; use linear interpolation or model-based interpolation methods to estimate missing values; Outlier processing: detect and identify outliers, and process outliers based on business context or statistical analysis, including deletion and correction; Duplicate data processing, deleting or aggregating duplicate data rows; S8.3: Perform data integration, which involves merging data from different data sources into a consistent data store to support analysis and modeling; Use database join operations to combine multiple data tables into one; combine data from different data sources based on a common unique identifier; S8.4: Data standardization, scaling numerical features to similar ranges, specifically including standardizing the data to a standard normal distribution with a mean of 0 and a variance of 1; Data normalization, scaling numerical features to a fixed range, including scaling features to the interval [0,1]; Logarithmic transformation: perform natural logarithm or Log(x+1) transformation on the data to reduce the skewness of the data; Feature extraction: extract new features from raw data through feature selection, principal component analysis (PCA), and polynomial feature construction methods to enhance model performance; S8.5: Feature selection, selecting the most relevant features to reduce the dimensionality of the dataset; selection is based on statistical tests, feature importance scores, or feature selection algorithms used during model training; Then perform data dimensionality reduction and use principal component analysis (PCA) to convert high-dimensional data into low-dimensional representation; reduce the storage and computing costs of the data set while maintaining important information of the data set.
3. The embedded step-by-step operation method according to claim 1, characterized in that: In step S9, the following steps are specifically performed: S9.1: Write model training code. Users use AI frameworks, including TensorFlow and PyTorch, to write model training code; define the model structure, including the input layer, hidden layer, and output layer, and configure training parameters, including learning rate, batch size, and number of iterations; S9.2: Users use development tools, including Jupyter Notebook and VS Code, to compile and run model training code; During the training process, the model continuously updates weights through the gradient descent optimization algorithm to minimize the loss function; S9.3: During the training process, the user monitors the loss function and accuracy metrics of the model; specifically, the commonly used performance evaluation metrics include Mean Squared Error (MSE), Mean Absolute Error (MAE), and Accuracy. According to the monitoring results, the user adjusts the training parameters and optimizes the model in a timely manner.
4. The embedded step-by-step operation method according to claim 2, characterized in that: In step S8.2, the mean formula is as follows: The data set X is X = {x1, x2, ..., x n }, where Xi is the i-th non-missing value sample; Use the median to fill in the missing values, as follows: Arrange the non-missing value samples x1, x2, x3...xn in ascending order. If n is an odd number, the median M is the first values; if n is an even number, the median M is the first value after sorting. The value and The average of the values.
5. The embedded step-by-step operation method according to claim 2, characterized in that: In step S8.2, in the calculation of linear interpolation, given two data points (x1, y1) and (x2, y2), to estimate the missing value y at x (x1 < x < x2), the linear interpolation formula is as follows:
6. The embedded step-by-step operation method according to claim 2, characterized in that: In step S8.4, feature extraction is performed. New features are extracted from the original data through feature selection, Principal Component Analysis (PCA), and polynomial feature construction methods. Specifically, the number of features is reduced by removing features with variances lower than the threshold. Features with small variances change little in the dataset. The formula is as follows: Among them, σ 2 is the variance, x i is the eigenvalue, μ is the mean of the feature, and N is the number of samples.
7. The embedded step-by-step operation method according to claim 6, characterized in that: Principal Component Analysis specifically maps high-dimensional data to a lower-dimensional space through a linear transformation, as follows: in, is the mean vector, xi is the sample vector, and m is the number of samples; Moreover, new features are generated by combining the original features polynomially through polynomial feature construction to increase the complexity of the model and capture more non-linear relationships; as follows: P(x)=a n x n +a n-1 x n-1 +…+a1x+a0 Where P(x) is the polynomial feature, x is the original feature, an, an-1, …, a1, a0 are the coefficients, and n is the degree of the polynomial.
8. The embedded step-by-step operation method according to claim 1, characterized in that: In step S9, the specific loss function used measures the difference between the model's prediction result and the true result; as follows: Among them, yi is the true value, is the predicted value, N is the number of samples; And Stochastic Gradient Descent (SGD) is used to update the model parameters to minimize the loss function; as follows: Among them, θ j are model parameters, α is the learning rate, and J(θ) is the loss function.
9. An embedded stepped operating platform, characterized in that: It includes a data preparation module for dataset collection and upload. The user needs to collect and prepare the dataset required for training and upload it to the embedded experimental platform, including various types of data such as images, texts, and audios. Data preprocessing: After uploading the dataset, data preprocessing is performed, including data cleaning, data augmentation, image flipping, rotation, and data formatting operations. Model training module for model selection and configuration. The user needs to select a suitable model type according to the application scenario and data characteristics, including Convolutional Neural Network (CNN) for image classification and Recurrent Neural Network (RNN) for sequence prediction. At the same time, the hyperparameters of the model are also configured through the model training module, including the learning rate, batch size, number of iterations, etc. Perform model training. After configuring the model, the user uses the embedded experimental platform to perform model training. Model evaluation and optimization module specifically performs model evaluation. After training is completed, the model is evaluated to determine whether its performance meets the requirements. Specifically, the model is run on the validation set or test set, and the performance metrics are calculated. Platform support and service module specifically performs user management, providing user registration, login, and permission management. And project management is carried out. The user creates and manages multiple projects on the platform support and service module. Each project can contain multiple datasets, models, and training tasks.
10. A computer storable medium, characterized in that: The computer-readable storage medium includes an embedded processing system and a stored program, and when the embedded system control program is running, the embedded-based step-by-step operation method described in claims 1-8 is controlled.