Model Training Method, Device and Computer Equipment Adapted to a Chip
By using a pre-set feature extraction algorithm similar to the chip's, followed by hybrid and exclusive chip-based training, the method aligns model performance across server and chip environments, ensuring accuracy and robustness.
Patent Information
- Application Number
- CN202111020514.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-01
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-09-01
AI Technical Summary
In the prior art, after the server-side training model is transplanted to the end-side device, there is a problem of degradation in accuracy, mainly due to the differences in feature extraction algorithms of the server-side and end-side devices.
Through a progressive step by step, the sample data is extracted using a preset feature extraction algorithm close to the chip feature extraction algorithm to form the first training data, and the training data of the preset feature extraction algorithm and the chip feature extraction algorithm are combined in the same batch of training data. Finally, the chip feature extraction algorithm is used for final training to ensure the consistency of features during model training and application.
It ensures that the actual application effect of the final model on the end-side device is the same as that during server training, and improves the robustness and accuracy of the model.
Smart Images

Figure CN113919492B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of model training, and particularly to a model training method, device, and computer device adapted to a chip. Background Art
[0002] The end-side voice solution does not require an Internet connection and can achieve intelligent recognition through an AI chip. It does not rely on cloud services and can be used immediately after purchase, greatly reducing the user threshold and providing intelligent functions for homes. The existing end-side intelligent transplantation process is as follows:
[0003] 1. Train voice recognition, wake-up, and other models on the server;
[0004] 2. Deploy the model to the chip for inference, including feature extraction, model compression, and post-processing parts.
[0005] One problem with the above process is that the feature extraction algorithm used when training the model on the server often has a certain difference from the feature extraction algorithm implemented by the chip hardware. For the same input, it is impossible to achieve complete consistency, resulting in the actual performance of the model after being deployed to the end side being inferior to that during training. For example, the recognition accuracy of the wake-up word decreases. Summary of the Invention
[0006] The main purpose of this application is to provide a model training method, device, and computer device adapted to a chip, aiming to solve the drawback that the accuracy of the model trained on the existing server decreases after being transplanted to the end-side device.
[0007] To achieve the above purpose, this application provides a model training method adapted to a chip, including:
[0008] Obtain sample data, call a preset feature extraction algorithm to extract features from the sample data to form first training data, and call a chip feature extraction algorithm to extract features from the sample data to form second training data. The function implemented by the preset feature extraction algorithm is the same as the function implemented by the chip feature extraction algorithm, and the chip feature extraction algorithm is the feature extraction algorithm of the chip;
[0009] Use the first training data to train the neural network model until a first preset condition is met to obtain an initial model;
[0010] In each batch of training data, combine the first training data and the second training data of the same batch in a random ratio to form third training data, and use the third training data to train the initial model until a second preset condition is met to obtain a secondary model;
[0011] Use the second training data to train the quadratic model until a third preset condition is met, obtaining a final model.
[0012] This application also provides another model training method adapted to a chip, including:
[0013] Obtain sample data, and call a preset feature extraction algorithm to extract features from the sample data, obtaining first training data;
[0014] Use the first training data to train the neural network until a first preset condition is met, obtaining an initial model;
[0015] In each batch of training data, extract features from the sample data through the preset feature extraction algorithm according to a random ratio, and extract features from the sample data through the chip feature extraction algorithm according to the random ratio, and combine them to obtain third training data. The functions implemented by the preset feature extraction algorithm and the chip feature extraction algorithm are the same, and the chip feature extraction algorithm is the feature extraction algorithm of the chip;
[0016] Use the third training data to train the initial model until a second preset condition is met, obtaining a quadratic model;
[0017] Call the chip feature extraction algorithm to extract features from the sample data, forming second training data;
[0018] Use the second training data to train the quadratic model until a third preset condition is met, obtaining a final model.
[0019] This application also provides a model training device adapted to a chip, including:
[0020] A first extraction module, configured to obtain sample data, call a preset feature extraction algorithm to extract features from the sample data, form first training data, and call a chip feature extraction algorithm to extract features from the sample data, form second training data. The functions implemented by the preset feature extraction algorithm and the chip feature extraction algorithm are the same, and the chip feature extraction algorithm is the feature extraction algorithm of the chip;
[0021] A first training module, configured to use the first training data to train the neural network until a first preset condition is met, obtaining an initial model;
[0022] A second training module, configured to combine the first training data and the second training data in a random ratio in each batch of training data to form third training data, and use the third training data to perform model training on the initial model until a second preset condition is met, obtaining a secondary model;
[0023] A third training module, configured to perform model training on the secondary model using the second training data until a third preset condition is met, obtaining a final model.
[0024] This application also provides another model training device adapted to a chip, including:
[0025] A second extraction module, configured to obtain sample data and call a preset feature extraction algorithm to perform feature extraction on the sample data, obtaining first training data;
[0026] A fourth training module, configured to perform model training on a neural network using the first training data until a first preset condition is met, obtaining an initial model;
[0027] A combination module, configured to perform feature extraction on the sample data through the preset feature extraction algorithm in a random ratio in each batch of training data, and perform feature extraction on the sample data through a chip feature extraction algorithm in the random ratio, combining to obtain third training data, where the functions implemented by the preset feature extraction algorithm are the same as the functions implemented by the chip feature extraction algorithm, and the chip feature extraction algorithm is the feature extraction algorithm of the chip;
[0028] A fifth training module, configured to perform model training on the initial model using the third training data until a second preset condition is met, obtaining a secondary model;
[0029] A third extraction module, configured to call the chip feature extraction algorithm to perform feature extraction on the sample data, forming second training data;
[0030] A sixth training module, configured to perform model training on the secondary model using the second training data until a third preset condition is met, obtaining a final model.
[0031] This application also provides a computer device, including a memory and a processor, where a computer program is stored in the memory, and when the processor executes the computer program, the steps of the method described in any one of the above are implemented.
[0032] This application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.
[0033] A model training method, device, and computer device adapted to a chip provided in the present application first use a preset feature extraction algorithm close to the chip feature extraction algorithm to perform feature extraction on sample data to obtain first training data, and then perform model training on a neural network to obtain an initial model. Then, in the same batch of training data, the initial model is trained using third training data obtained by performing combined feature extraction on the sample data using the preset feature extraction algorithm and the chip feature extraction algorithm to obtain a secondary model. Finally, the secondary model is trained using second training samples obtained by extracting features from the sample data using the chip feature extraction algorithm to obtain a final model. The present application completes the process from the preset feature extraction algorithm to the chip feature extraction algorithm in a successive progressive manner, ensuring the consistency of features during model training and application. As a result, after the final model trained on the server side is deployed on the chip of the edge device, the actual processing effect of the final model in practical applications is the same as that during training on the server side, effectively ensuring the robustness and accuracy of the final model. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 FIG. is a schematic diagram of the steps of a model training method adapted to a chip according to an embodiment of the present application;
[0035] Figure 2 FIG. is a schematic diagram of the steps of another model training method adapted to a chip according to an embodiment of the present application;
[0036] Figure 3 FIG. is a block diagram of the overall structure of a model training device adapted to a chip according to an embodiment of the present application;
[0037] Figure 4 FIG. is a block diagram of the overall structure of another model training device adapted to a chip according to an embodiment of the present application;
[0038] Figure 5 FIG. is a schematic block diagram of the structure of a computer device according to an embodiment of the present application.
[0039] The implementation, functional features, and advantages of the objectives of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0041] Refer to Figure 1 , an embodiment of the present application provides a model training method adapted to a chip, including:
[0042] S1: Obtain sample data, call a preset feature extraction algorithm to extract features from the sample data to form first training data, and call a chip feature extraction algorithm to extract features from the sample data to form second training data. The function implemented by the preset feature extraction algorithm is the same as the function implemented by the chip feature extraction algorithm, and the chip feature extraction algorithm is the feature extraction algorithm of the chip;
[0043] S2: Use the first training data to train the neural network model until the first preset condition is met to obtain an initial model;
[0044] S3: In each batch of training data, combine the first training data and the second training data of the same batch according to a random ratio to form third training data, and use the third training data to train the initial model until the second preset condition is met to obtain a secondary model;
[0045] S4: Use the second training data to train the secondary model until the third preset condition is met to obtain a final model.
[0046] Preferably, after the step of using the second training data to train the secondary model until the third preset condition is met to obtain a final model, it includes:
[0047] S5: Deploy the final model in the embedded environment of the chip.
[0048] In this embodiment, the training system obtains sample data, and then separately calls a preset feature extraction algorithm and a chip feature extraction algorithm to extract features from the sample data. After the preset feature extraction algorithm extracts features from the sample data, first training data is obtained, and after the chip feature extraction algorithm extracts features from the sample data, second training data is formed. Among them, the functions implemented by the preset feature extraction algorithm and the chip feature extraction algorithm are the same, that is, the features extracted from the sample data by both belong to the same type; the chip feature extraction algorithm is the feature extraction algorithm possessed by the chip of the edge device itself. For general feature extraction algorithms, the general process is the same. For example, for the FBANK algorithm, the implementation process of the hardware (i.e., the chip of the edge device) and the software (i.e., the feature extraction algorithm used in the server-side training model) is basically similar, but there will be differences in details. Developers debug an algorithm close to the chip output with the help of the API provided by the chip, so as to obtain the preset feature extraction algorithm. Although the corresponding implementation source code of the chip feature extraction algorithm cannot be directly obtained (the feature extraction algorithm of the chip is usually not open source), during model training, the chip feature extraction algorithm can be called to extract features from the sample data by using the dynamic library or static library of the chip. After the first training data and the second training data are extracted, the training system caches the first training data and the second training data in the internal database for direct use during subsequent model training. The model training in this embodiment is divided into three stages, which are in turn: the first stage, the training system uses the first training data to train the neural network model until the first preset condition is met, such as training a preset number of times, or the loss function no longer decreases, or the accuracy rate no longer increases significantly, so as to obtain an initial model. In the first stage, by using the preset feature extraction algorithm close to the chip feature extraction algorithm to train the model, the convergence speed of the model can be accelerated, the training time of the model can be reduced, and the product development cycle can be accelerated; in addition, by using the preset feature extraction algorithm implemented by the developer himself, it is easier to obtain the corresponding pre-trained model, and it is easier to perform fine-tune (transfer learning) on it. The second stage, in each batch of training data, the training system combines the first training data and the second training data of the same batch according to a randomly generated random ratio to form third training data (the first training data and the second training data that make up the third training data both belong to the same batch to ensure the consistency of the training data). Then use the third training data to continue training the initial training model until the second preset condition is met to obtain a secondary model. Among them, the second preset condition can be: training a preset number of times, or the loss function no longer decreases, or the accuracy rate of the model no longer increases significantly; the second preset condition can be the same as or different from the first preset condition.In the second stage, the transition from the preset feature extraction algorithm to the chip feature extraction algorithm is achieved by combining and training the first training data and the second training data of the same batch, enabling the model training to converge consistently without divergence. Moreover, the method of simultaneously training the second training data extracted by the preset feature extraction algorithm and the first training data extracted by the chip feature extraction algorithm is equivalent to enhancing and perturbing the training data of the model, improving the robustness of the features, and making the trained model perform more robustly. In the third stage, the training system uses the second training data to train the secondary model until the third preset condition is met (the third preset condition can be: training a preset number of times, or the loss function no longer decreases, or the accuracy of the model no longer increases significantly), obtaining the final model. In the third stage, the second training data extracted entirely by the chip feature extraction algorithm is used to further train the secondary model, ensuring the consistency of data input during model training and model inference (i.e., the final trained model is deployed on the chip for application), and ensuring the accuracy of the final model obtained from training. After completing the model training, the final model is deployed in the embedded environment of the chip of the edge device, enabling the edge device to apply the final model.
[0049] In this embodiment, the training system completes the process from the preset feature extraction algorithm to the chip feature extraction algorithm by progressively training the model in sequence, ensuring the consistency of features during model training and application. Thus, after the final model trained on the server side is deployed on the chip of the edge device, the actual processing effect of the final model in application is the same as that during training on the server side, effectively ensuring the robustness and accuracy of the final model.
[0050] Further, the step of mixing the first training data and the second training data of the same batch in a random ratio in each batch of training data to form the third training data includes:
[0051] S301: Randomly generate the random ratio and obtain the total amount of data contained in a batch of training data;
[0052] S302: Calculate the first data amount and the second data amount respectively according to the total amount of data and the random ratio;
[0053] S303: Select the first training sub-data corresponding to the first data amount from the first training data of the same batch, and select the second training sub-data corresponding to the second data amount from the second training data of the same batch. The first training sub-data and the second training sub-data belong to the same batch of training data;
[0054] S304: Combine the first training sub-data and the second training sub-data to obtain the third training data, which is the training data required for one batch of model training.
[0055] In this embodiment, the training system randomly generates a random ratio, and then obtains the total amount of data included in the training data for one batch during model training. The training system calculates the first data amount and the second data amount respectively according to the total amount of data and the random ratio. The first data amount corresponds to the first training data, and the second data amount corresponds to the second training data. The training system selects the first training sub-data corresponding to the first data amount from the first training data of the same batch, and selects the second training sub-data corresponding to the second training data amount from the second training data of the same batch; wherein, the first training sub-data and the second training sub-data belong to the training data of the same batch. The training system combines the first training sub-data and the second training sub-data to obtain the third training data, which is the training data required for one batch during model training.
[0056] Further, the step of using the first training data to perform model training on the neural network until a first preset condition is met to obtain an initial model includes:
[0057] S201: Use the first training data to perform model training on the neural network and cumulatively count the training times in real time;
[0058] S202: When the number of training times reaches the number threshold, it is determined that the first preset condition is met, and the model training of the neural network is stopped to obtain the initial model.
[0059] In this embodiment, the training system uses the first training data to perform model training on the neural network and cumulatively counts the training times in real time. After each model training, the training system compares the cumulative number of training times with the preset number threshold. If the cumulative number of training times has not reached the number threshold (i.e., the number of training times is less than the number threshold), then continue to use the first training data to perform model training on the neural network. If the cumulative number of training times reaches the number threshold (i.e., the number of training times is equal to the number threshold), then the training system determines that the first preset condition is currently met, and stops continuing to perform model training on the neural network to obtain the initial model.
[0060] Further, before the steps of obtaining the sample data, calling a preset feature extraction algorithm to perform feature extraction on the sample data to form the first training data, and calling a chip feature extraction algorithm to perform feature extraction on the sample data to form the second training data, include:
[0061] S6: Obtain the output result type of the chip feature extraction algorithm;
[0062] S7: According to the pre-constructed mapping relationship table between the output result type and the algorithm number, match and obtain the algorithm number corresponding to the output result type;
[0063] S8: Retrieve the feature extraction algorithm corresponding to the algorithm number from the algorithm database as the preset feature extraction algorithm.
[0064] In this embodiment, the algorithm database of the training system stores feature extraction algorithms pre-written or input by developers. These feature extraction algorithms are correspondingly set with algorithm numbers according to the type of output result or the implemented function. The developers have pre-constructed a mapping relationship table between the output result type and the algorithm number, and this mapping relationship table between the output result type and the algorithm number includes multiple groups of one-to-one corresponding output result types and algorithm numbers. During model training, the training system obtains the output result type of the chip feature extraction algorithm (this output result type can be manually input by the developer or obtained by analyzing the features extracted by the training system through the chip feature extraction algorithm), and then according to the mapping relationship table between the output result and the algorithm number, matches and obtains the algorithm number corresponding to the current output result type. Finally, the training system retrieves the feature extraction algorithm corresponding to the algorithm number from the algorithm database as the preset feature extraction algorithm required for the current model training, realizing the rapid call of the preset feature extraction algorithm.
[0065] Refer to Figure 2 , another embodiment of the present application also provides a model training method adapted to a chip, including:
[0066] A1: Obtain sample data, and call the preset feature extraction algorithm to extract features from the sample data to obtain the first training data;
[0067] A2: Use the first training data to perform model training on the neural network until a first preset condition is satisfied to obtain an initial model;
[0068] A3: In each batch of training data, extract features from the sample data through the preset feature extraction algorithm according to a random ratio, and extract features from the sample data through the chip feature extraction algorithm according to the random ratio, and combine them to obtain the third training data. The function implemented by the preset feature extraction algorithm is the same as the function implemented by the chip feature extraction algorithm, and the chip feature extraction algorithm is the feature extraction algorithm of the chip;
[0069] A4: Use the third training data to perform model training on the initial model until a second preset condition is satisfied to obtain a secondary model;
[0070] A5: Call the chip feature extraction algorithm to extract features from the sample data to form the second training data;
[0071] A6: Use the second training data to train the quadratic model until a third preset condition is met, and obtain the final model.
[0072] This embodiment also provides another model training method adapted to the chip, which is also divided into three stages for model training. Specifically, in the first stage, the training system obtains sample data, and then calls a preset feature extraction algorithm close to the chip feature extraction algorithm to extract features from the sample data to obtain the first training data. Then, the training system uses the first training data to train the neural network model until a first preset condition is met, and an initial model is obtained. In the first stage, compared with calling the library function of the chip to extract features, that is, calling the chip feature extraction algorithm to extract features from the sample data, since the preset feature extraction algorithm is directly deployed on the model training server, when using the preset feature extraction algorithm to extract features from the sample data, matrix operation acceleration or GPU acceleration can be used to speed up the model training speed; while the chip feature extraction algorithm encapsulated in the chip cannot do this. The input parameters of the chip feature extraction algorithm are fixed, and the general input is a single audio, while the preset feature extraction algorithm can calculate in units of a batch. Therefore, in the first stage, by using the preset feature extraction algorithm close to the chip feature extraction algorithm to train the model, the convergence speed of the model can be accelerated, the training time of the model can be reduced, and the product development cycle can be speeded up. In the second stage, when the training system uses each batch of training data to train the model, it uses the preset feature algorithm and the chip feature algorithm to extract features from the sample data according to a randomly generated random ratio, and combines them to obtain the third training data (that is, the third training data is composed of partial feature data extracted from the sample data by the preset feature extraction algorithm and partial feature data extracted from the sample data by the chip feature extraction algorithm). Among them, the functions implemented by the preset feature extraction algorithm or the types of output results are the same as those implemented by the chip feature extraction algorithm, and the chip feature extraction algorithm is the feature extraction algorithm of the chip of the end-side device. The training system uses the third training data to train the initial model until a second preset condition is met, and a quadratic model is obtained. In the third stage, the training system calls the chip feature extraction algorithm to extract features from the sample data to form the second training data. Then, the second training data is used to train the quadratic model until a third preset condition is met, and the final model is obtained.
[0073] The model training method corresponding to steps A1 - A6 in this embodiment has the same basic idea as the model training method corresponding to steps S1 - S4 above. The difference is that the model training method corresponding to steps S1 - S4 pre - extracts and caches the first training data and the second training data, which can be directly called during subsequent model training, saving calculation time. However, the first training data and the second training data require additional memory space, and it takes a certain amount of reading time to retrieve the first training data and the second training data from the memory space. The model training method corresponding to steps A1 - A6 is to extract the sample data through the preset feature extraction algorithm and the chip feature extraction algorithm immediately when training data is needed during the model training in the corresponding stage, without occupying additional memory space, but it takes time for feature extraction in each stage. In practical applications, if the time taken for the preset feature extraction algorithm and the chip feature extraction algorithm to extract features from the sample data is less than the time taken to retrieve the first training data and the second training data from the memory space, then the model training method corresponding to steps A1 - A6 is preferred; if the time taken for the preset feature extraction algorithm and the chip feature extraction algorithm to extract features from the sample data is greater than the time taken to retrieve the first training data and the second training data from the memory space, then the model training method corresponding to steps S1 - S4 is preferred to improve the speed of model training.
[0074] Referring to Figure 3 , in an embodiment of the present application, a model training device adapted to a chip is further provided, including:
[0075] A first extraction module 1, configured to obtain sample data, call a preset feature extraction algorithm to extract features from the sample data to form first training data, and call a chip feature extraction algorithm to extract features from the sample data to form second training data. The functions implemented by the preset feature extraction algorithm are the same as those implemented by the chip feature extraction algorithm, and the chip feature extraction algorithm is the feature extraction algorithm of the chip;
[0076] A first training module 2, configured to use the first training data to perform model training on a neural network until a first preset condition is met to obtain an initial model;
[0077] A second training module 3, configured to combine the first training data and the second training data in a random ratio in each batch of training data to form third training data, and use the third training data to perform model training on the initial model until a second preset condition is met to obtain a secondary model;
[0078] A third training module 4, configured to use the second training data to perform model training on the secondary model until a third preset condition is met to obtain a final model.
[0079] Furthermore, the model training device further includes:
[0080] A deployment module 5, configured to deploy the final model in the embedded environment of the chip.
[0081] Furthermore, the second training model 3 includes:
[0082] A generation unit, configured to randomly generate the random ratio and obtain the total amount of data included in a batch of training data;
[0083] A calculation unit, configured to calculate a first data amount and a second data amount respectively according to the total amount of data and the random ratio;
[0084] A selection unit, configured to select first training sub-data corresponding to the first data amount from the first training data of the same batch, and select second training sub-data corresponding to the second data amount from the second training data of the same batch, where the first training sub-data and the second training sub-data belong to the training data of the same batch;
[0085] A combination unit, configured to combine the first training sub-data and the second training sub-data to obtain the third training data, where the third training data is the training data required for one batch of model training.
[0086] Furthermore, the first training module 3 includes:
[0087] An accumulation unit, configured to perform model training on the neural network using the first training data and accumulate the training times in real time;
[0088] A determination unit, configured to determine that the first preset condition is satisfied and stop performing model training on the neural network to obtain the initial model when the training times reach the times threshold.
[0089] Furthermore, the model training device further includes:
[0090] An acquisition module 6, configured to acquire the output result type of the chip feature extraction algorithm;
[0091] A matching module 7, configured to match the algorithm number corresponding to the output result type according to a pre-constructed mapping relation table between the output result type and the algorithm number;
[0092] An invocation module 8, configured to invoke the feature extraction algorithm corresponding to the algorithm number from the algorithm database as the preset feature extraction algorithm.
[0093] Refer to Figure 4, in an embodiment of the present application, another model training device adapted to the chip is further provided, including:
[0094] The second extraction module 9 is configured to obtain sample data and call a preset feature extraction algorithm to perform feature extraction on the sample data to obtain first training data;
[0095] The fourth training module 10 is configured to use the first training data to perform model training on the neural network until a first preset condition is met to obtain an initial model;
[0096] The combination module 11 is configured to, in each batch of training data, perform feature extraction on the sample data through the preset feature extraction algorithm according to a random ratio, and perform feature extraction on the sample data through the chip feature extraction algorithm according to the random ratio, and combine them to obtain third training data. The function implemented by the preset feature extraction algorithm is the same as the function implemented by the chip feature extraction algorithm, and the chip feature extraction algorithm is the feature extraction algorithm of the chip;
[0097] The fifth training module 12 is configured to use the third training data to perform model training on the initial model until a second preset condition is met to obtain a secondary model;
[0098] The third extraction module 13 is configured to call the chip feature extraction algorithm to perform feature extraction on the sample data to form second training data;
[0099] The sixth training module 14 is configured to use the second training data to perform model training on the secondary model until a third preset condition is met to obtain a final model.
[0100] In this embodiment, each module and unit in the model training method adapted to the chip is used to correspondingly execute each step in the above model training method adapted to the chip, and the specific implementation process is not described in detail here.
[0101] A model training device adapted to a chip provided in this embodiment first uses a preset feature extraction algorithm close to the chip feature extraction algorithm to perform feature extraction on sample data to obtain first training data for model training of a neural network, and an initial model is obtained. Then, in the same batch of training data, the initial model is trained using third training data obtained by performing combined feature extraction on the sample data using the preset feature extraction algorithm and the chip feature extraction algorithm to obtain a secondary model. Finally, the secondary model is trained using second training samples obtained by performing feature extraction on the sample data using the chip feature extraction algorithm to obtain a final model. Through a successive progressive manner, this application completes the process from the preset feature extraction algorithm to the chip feature extraction algorithm, ensuring the consistency of features during model training and application. Therefore, after the final model trained on the server side is deployed on the chip of the edge device, the processing effect of the actual application of the final model is the same as that during server-side training, effectively ensuring the robustness and accuracy of the final model.
[0102] Referring to Figure 5 , this embodiment of the present application also provides a computer device, which can be a server, and its internal structure can be as Figure 3 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data such as first training data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a model training method adapted to a chip.
[0103] The above processor executes the steps of the above model training method adapted to a chip:
[0104] S1: Obtain sample data, call a preset feature extraction algorithm to perform feature extraction on the sample data to form first training data, and call a chip feature extraction algorithm to perform feature extraction on the sample data to form second training data. The functions implemented by the preset feature extraction algorithm are the same as those implemented by the chip feature extraction algorithm, and the chip feature extraction algorithm is the feature extraction algorithm of the chip;
[0105] S2: Use the first training data to perform model training on a neural network until a first preset condition is met to obtain an initial model;
[0106] S3: In each batch of training data, combine the first training data and the second training data of the same batch according to a random ratio to form third training data, and use the third training data to perform model training on the initial model until a second preset condition is met to obtain a secondary model;
[0107] S4: Use the second training data to perform model training on the secondary model until a third preset condition is met to obtain a final model.
[0108] Preferably, after the step of using the second training data to perform model training on the secondary model until a third preset condition is met to obtain a final model, it includes:
[0109] S5: Deploy the final model in the embedded environment of the chip.
[0110] Further, the step of mixing the first training data and the second training data of the same batch according to a random ratio in each batch of training data to form third training data includes:
[0111] S301: Randomly generate the random ratio and obtain the total amount of data included in a batch of training data;
[0112] S302: Calculate a first data amount and a second data amount respectively according to the total amount of data and the random ratio;
[0113] S303: Select first training sub-data corresponding to the first data amount from the first training data of the same batch, and select second training sub-data corresponding to the second data amount from the second training data of the same batch. The first training sub-data and the second training sub-data belong to the same batch of training data;
[0114] S304: Combine the first training sub-data and the second training sub-data to obtain the third training data, and the third training data is the training data required for one batch of model training.
[0115] Further, the step of using the first training data to perform model training on the neural network until a first preset condition is met to obtain an initial model includes:
[0116] S201: Use the first training data to perform model training on the neural network and cumulatively count the number of training times in real time;
[0117] S202: When the number of training times reaches a threshold, it is determined that the first preset condition is met, and stop performing model training on the neural network to obtain the initial model.
[0118] Further, before the steps of obtaining sample data, calling a preset feature extraction algorithm to extract features from the sample data to form first training data, and calling a chip feature extraction algorithm to extract features from the sample data to form second training data, it includes:
[0119] S6: Obtain the output result type of the chip feature extraction algorithm;
[0120] S7: According to the pre-constructed mapping relationship table between the output result type and the algorithm number, match to obtain the algorithm number corresponding to the output result type;
[0121] S8: Retrieve the feature extraction algorithm corresponding to the algorithm number from the algorithm database as the preset feature extraction algorithm.
[0122] The above processor also executes the steps of the above model training method adapted to the chip:
[0123] A1: Obtain sample data, and call a preset feature extraction algorithm to extract features from the sample data to obtain first training data;
[0124] A2: Use the first training data to train the neural network model until a first preset condition is met to obtain an initial model;
[0125] A3: In each batch of training data, extract features from the sample data through the preset feature extraction algorithm according to a random ratio, and extract features from the sample data through the chip feature extraction algorithm according to the random ratio, and combine to obtain third training data. The functions implemented by the preset feature extraction algorithm and the chip feature extraction algorithm are the same, and the chip feature extraction algorithm is the feature extraction algorithm of the chip;
[0126] A4: Use the third training data to train the initial model until a second preset condition is met to obtain a secondary model;
[0127] A5: Call the chip feature extraction algorithm to extract features from the sample data to form second training data;
[0128] A6: Use the second training data to train the secondary model until a third preset condition is met to obtain a final model.
[0129] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a model training method adapted to a chip. The model training method adapted to the chip is specifically:
[0130] S1: Obtain sample data, call a preset feature extraction algorithm to extract features from the sample data to form first training data, and call a chip feature extraction algorithm to extract features from the sample data to form second training data. The function implemented by the preset feature extraction algorithm is the same as the function implemented by the chip feature extraction algorithm, and the chip feature extraction algorithm is the feature extraction algorithm of the chip;
[0131] S2: Use the first training data to train the neural network model until the first preset condition is met to obtain an initial model;
[0132] S3: In each batch of training data, combine the first training data and the second training data of the same batch according to a random ratio to form third training data, and use the third training data to train the initial model until the second preset condition is met to obtain a secondary model;
[0133] S4: Use the second training data to train the secondary model until the third preset condition is met to obtain a final model.
[0134] Preferably, after the step of using the second training data to train the secondary model until the third preset condition is met to obtain a final model, it includes:
[0135] S5: Deploy the final model in the embedded environment of the chip.
[0136] Further, the step of combining the first training data and the second training data of the same batch according to a random ratio in each batch of training data to form third training data includes:
[0137] S301: Randomly generate the random ratio and obtain the total amount of data included in a batch of training data;
[0138] S302: Calculate the first data amount and the second data amount respectively according to the total amount of data and the random ratio;
[0139] S303: Select first training sub-data corresponding to the first data amount from the first training data of the same batch, and select second training sub-data corresponding to the second data amount from the second training data of the same batch. The first training sub-data and the second training sub-data belong to the same batch of training data;
[0140] S304: Combine the first training sub-data and the second training sub-data to obtain the third training data, and the third training data is the training data required for one batch of model training.
[0141] Further, the step of training the neural network using the first training data until a first preset condition is met to obtain an initial model includes:
[0142] S201: Train the neural network using the first training data and cumulatively count the training times in real time;
[0143] S202: When the training times reach a threshold, it is determined that the first preset condition is met, and the training of the neural network is stopped to obtain the initial model.
[0144] Further, before the steps of obtaining sample data, calling a preset feature extraction algorithm to extract features from the sample data to form first training data, and calling a chip feature extraction algorithm to extract features from the sample data to form second training data, include:
[0145] S6: Obtain the output result type of the chip feature extraction algorithm;
[0146] S7: According to a pre-constructed mapping relationship table between the output result type and the algorithm number, match to obtain the algorithm number corresponding to the output result type;
[0147] S8: Retrieve the feature extraction algorithm corresponding to the algorithm number from the algorithm database as the preset feature extraction algorithm.
[0148] Another implementation of the model training method adapted to the chip when the computer program is executed by the processor:
[0149] A1: Obtain sample data and call a preset feature extraction algorithm to extract features from the sample data to obtain first training data;
[0150] A2: Train the neural network using the first training data until a first preset condition is met to obtain an initial model;
[0151] A3: In each batch of training data, extract features from the sample data through the preset feature extraction algorithm according to a random ratio, and extract features from the sample data through the chip feature extraction algorithm according to the random ratio, and combine to obtain third training data. The functions implemented by the preset feature extraction algorithm and the chip feature extraction algorithm are the same, and the chip feature extraction algorithm is the feature extraction algorithm of the chip;
[0152] A4: Train the initial model using the third training data until a second preset condition is met to obtain a secondary model;
[0153] A5: Invoke the chip feature extraction algorithm to extract features from the sample data to form the second training data;
[0154] A6: Use the second training data to train the secondary model until the third preset condition is met to obtain the final model.
[0155] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided in this application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0156] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, device, first object, or method including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, device, first object, or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, device, first object, or method including that element.
[0157] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A model training method adapted to a chip, characterized in that, Including: Obtain sample data, call a preset feature extraction algorithm to perform feature extraction on the sample data to form first training data, and call a chip feature extraction algorithm to perform feature extraction on the sample data to form second training data. The function implemented by the preset feature extraction algorithm is the same as the function implemented by the chip feature extraction algorithm, and the chip feature extraction algorithm is the feature extraction algorithm of the chip. Use the first training data to train the neural network model until a first preset condition is met to obtain an initial model. In each batch of training data, combine the first training data and the second training data of the same batch according to a random ratio to form third training data, and use the third training data to train the initial model until a second preset condition is met to obtain a secondary model. Use the second training data to train the secondary model until a third preset condition is met to obtain a final model. Before the step of obtaining sample data, calling a preset feature extraction algorithm to perform feature extraction on the sample data to form first training data, and calling a chip feature extraction algorithm to perform feature extraction on the sample data to form second training data, it includes: Obtain the output result type of the chip feature extraction algorithm. According to a pre-constructed mapping relationship table between the output result type and the algorithm number, match and obtain the algorithm number corresponding to the output result type. Retrieve from the algorithm database the feature extraction algorithm corresponding to the algorithm number as the preset feature extraction algorithm.
2. The model training method adapted to the chip according to claim 1, wherein After the step of using the second training data to train the secondary model until a third preset condition is met to obtain a final model, it includes: Deploy the final model in the embedded environment of the chip.
3. The model training method adapted to a chip according to claim 1, wherein The step of combining the first training data and the second training data according to a random ratio in each batch of training data to form third training data includes: Randomly generate the random ratio and obtain the total amount of data included in a batch of training data. According to the total amount of data and the random ratio, calculate the first data amount and the second data amount respectively. Select first training sub-data corresponding to the first data amount from the first training data of the same batch, and select second training sub-data corresponding to the second data amount from the second training data of the same batch. The first training sub-data and the second training sub-data belong to the same batch of training data. Combine the first training sub-data and the second training sub-data to obtain the third training data, and the third training data is the training data required for one batch of model training.
4. The model training method adapted to the chip according to claim 1, characterized in that The step of using the first training data to train the neural network model until a first preset condition is met to obtain an initial model includes: Use the first training data to train the neural network model and accumulate the training times in real time. When the training times reach the times threshold, it is determined that the first preset condition is met, and stop training the neural network model to obtain the initial model.
5. A model training method adapted to a chip, characterized in that, Including: Obtain sample data, and call a preset feature extraction algorithm to extract features from the sample data to obtain first training data; Use the first training data to train the neural network model until a first preset condition is met to obtain an initial model; In each batch of training data, extract features from the sample data through the preset feature extraction algorithm according to a random ratio, and extract features from the sample data through the chip feature extraction algorithm according to the random ratio, and combine them to obtain third training data. The functions implemented by the preset feature extraction algorithm and the chip feature extraction algorithm are the same. The chip feature extraction algorithm is the feature extraction algorithm of the chip. Among them, by obtaining the output result type of the chip feature extraction algorithm; according to the pre-constructed mapping relationship table between the output result type and the algorithm number, match to obtain the algorithm number corresponding to the output result type; retrieve the feature extraction algorithm corresponding to the algorithm number from the algorithm database to obtain the preset feature extraction algorithm; Use the third training data to train the initial model until a second preset condition is met to obtain a secondary model; Call the chip feature extraction algorithm to extract features from the sample data to form second training data; Use the second training data to train the secondary model until a third preset condition is met to obtain a final model.
6. A model training device adapted to a chip, for performing the model training method adapted to a chip according to any one of claims 1-4, characterized in that, Including: A first extraction module, configured to obtain sample data, call a preset feature extraction algorithm to extract features from the sample data to form first training data, and call a chip feature extraction algorithm to extract features from the sample data to form second training data. The functions implemented by the preset feature extraction algorithm and the chip feature extraction algorithm are the same. The chip feature extraction algorithm is the feature extraction algorithm of the chip; A first training module, configured to use the first training data to train the neural network model until a first preset condition is met to obtain an initial model; A second training module, configured to, in each batch of training data, combine the first training data and the second training data according to a random ratio to form third training data, and use the third training data to train the initial model until a second preset condition is met to obtain a secondary model; A third training module, configured to use the second training data to train the secondary model until a third preset condition is met to obtain a final model.
7. A model training device adapted to a chip, for performing the model training method adapted to a chip described in claim 5, characterized in that, Including: A second extraction module, configured to obtain sample data, and call a preset feature extraction algorithm to extract features from the sample data to obtain first training data; A fourth training module, configured to use the first training data to train the neural network model until a first preset condition is met to obtain an initial model; A combined module is used to extract features from the sample data according to a random ratio through the preset feature extraction algorithm and extract features from the sample data according to the random ratio through the chip feature extraction algorithm in each batch of training data, and combine them to obtain third training data. The function implemented by the preset feature extraction algorithm is the same as the function implemented by the chip feature extraction algorithm, and the chip feature extraction algorithm is the feature extraction algorithm of the chip; A fifth training module is used to train the initial model with the third training data until a second preset condition is met to obtain a secondary model; A third extraction module is used to call the chip feature extraction algorithm to extract features from the sample data to form second training data; A sixth training module is used to train the secondary model with the second training data until a third preset condition is met to obtain a final model.
8. A computer device, comprising a memory and a processor, wherein a computer program is stored in the memory, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Method and Apparatus of Training Acoustic Feature Extracting Model, Device and Computer Storage Medium
US20180336888A1
Chip adaptation determination method and related product
WO2020041960A1