Neural network training method and system based on model parameter sharing
The parameter sharing of the DFSMN acoustic model is solved through the Seq-Skip parameter sharing strategy, which solves the problem of efficient and accurate speech recognition on low-power chips, improves the accuracy and training efficiency of speech recognition, and adapts to different task requirements.
Patent Information
- Application Number
- CN202510428462.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-18
AI Technical Summary
It is difficult for the existing technology to meet the requirements of efficient and accurate recognition while under the limitations of memory and computing power.
The Seq-Skip parameter sharing strategy is adopted to share parameters for the DFSMN acoustic model. By selecting some parameters to build a sharing layer, using the remaining computing power to generate an N-layer model, and conducting training and performance testing to ensure that the memory usage remains unchanged.
It improves speech recognition accuracy and model training efficiency, adapts to different languages and tasks, and improves the efficiency and performance of natural language processing tasks.
Smart Images

Figure CN120338024A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of deep learning technology, and particularly to a neural network training method and system based on model parameter sharing. Background Art
[0002] With the rapid development of deep learning and neural network technologies in the field of speech recognition, how to achieve efficient and accurate speech recognition on low-power chips has become the focus of attention in the industry. On low-power chips, due to limitations in memory and computing power, the number of parameters and computational complexity of neural network models often need to be appropriately controlled to meet the requirements of actual application scenarios.
[0003] Currently, the existing technologies mainly include two solutions: one is the SEQUENCE parameter sharing strategy based on the Transformer multi-layer encoder-decoder structure, that is, by setting the parameters of consecutive hidden layers to be the same to reduce the number of model parameters; the other is the acoustic model compression technology, which reduces the model size through methods such as pruning and quantization to meet the limitations of memory and computing power on low-power chips. Both of these solutions have achieved compression of model parameters and computing resources to varying degrees.
[0004] However, in practical applications, a single parameter sharing or model compression strategy is difficult to simultaneously meet the requirements of efficient computing and accurate recognition on low-power speech chips. Specifically, although the first technology effectively reduces the number of parameters through parameter sharing, due to the large scale of the weight matrix after sharing in each layer, the computational complexity and processing time increase significantly. And the other technology may cause the model to be unable to fully capture complex speech features due to excessive compression, thereby affecting the recognition accuracy. To sum up, in a resource-constrained low-power environment, how to achieve efficient training of neural network models and accurate speech recognition while ensuring memory and computing power requirements has become the main technical problem to be solved urgently. Summary of the Invention
[0005] This application provides a neural network training method and system based on model parameter sharing, which can, on the premise of ensuring the same memory occupancy in a resource-constrained low-power environment, through the Seq-Skip parameter sharing method, utilize the remaining extra computing power to achieve efficient training of neural network models and accurate speech recognition. This application provides the following technical solutions:
[0006] In a first aspect, this application provides a neural network training method based on model parameter sharing, and the method includes:
[0007] For the M-layer DFSMN acoustic model already deployed on the edge side, calculate its computing power and the number of parameters, appropriately increase the total computing power requirement based on the same number of parameters of this model, and determine whether there is remaining computing power on the chip;
[0008] If there is remaining computing power in the chip, select some parameters of each layer of the DFSMN acoustic model to construct a shared layer;
[0009] Organize and allocate the constructed shared layer according to the Seq-Skip parameter sharing strategy;
[0010] Generate an N-layer DFSMN acoustic model under the selected Seq-Skip parameter sharing strategy, train and perform performance tests on the DFSMN acoustic model after parameter sharing, and determine whether the indicators of the model meet the expected requirements.
[0011] In a specific implementable solution, for the existing M-layer DFSMN acoustic model deployed on the edge side, calculate its computing power and the number of parameters, appropriately increase the total computing power requirement based on the same number of parameters of this model, and determining whether there is remaining chip computing power includes:
[0012] Starting from each layer of the DFSMN model, for various types of parameters included in each layer, analyze its network structure to count the number of parameters;
[0013] After counting the number of parameters in each layer, estimate the amount of computation required for each layer based on the number of multiplication-addition operations participated by each parameter in the forward propagation, and add up the computing requirements of all layers to obtain the total computing power requirement of the entire DFSMN model.
[0014] In a specific implementable solution, for the existing M-layer DFSMN acoustic model deployed on the edge side, calculate its computing power and the number of parameters, appropriately increase the total computing power requirement based on the same number of parameters of this model, and determining whether there is remaining chip computing power further includes:
[0015] Evaluate the remaining computing power of the chip in the target application scenario, and compare the actual available computing power of the chip with the total computing power requirement of the DFSMN model;
[0016] If the chip has no additional remaining computing power, that is, the actual available computing power is less than or equal to the model requirement, then directly end.
[0017] In a specific implementable solution, the step of selecting some parameters of each layer of the DFSMN acoustic model to construct a shared layer when there is remaining chip computing power includes:
[0018] For each layer in the original M-layer DFSMN acoustic model, divide the parameters into three parts: hidden layer, linear mapping layer, and feed-forward sequence memory layer;
[0019] Separate and prepare independent parameters for each layer of the M layers. According to the calculation results of the model parameters and the remaining computing power of the chip in the early stage, selectively share some parameters in each layer of the DFSMN model, construct the corresponding parameter sharing layer, and allocate the selected shared parameters to the N layers;
[0020] N represents the total number of layers of the finally constructed DFSMN model, M represents the number of layers of the initial DFSMN model, and N is greater than or equal to M.
[0021] In a specific implementable solution, for each layer of the original M-layer DFSMN acoustic model, the parameters are divided into three parts: the hidden layer, the linear mapping layer, and the feedforward sequence memory layer, including:
[0022] The hidden layer part adopts a fully connected structure, and the activation function is ReLU;
[0023] The linear mapping layer adopts a low-rank linear mapping matrix;
[0024] The feedforward sequence memory layer models the historical and future temporal information and integrates it into a fixed-dimensional encoding.
[0025] In a specific implementable solution, the organization and allocation of the constructed sharing layer according to the Seq-Skip parameter sharing strategy include:
[0026] For a certain type of parameters selected in the DFSMN acoustic model, the Seq-Skip parameter sharing strategy first sets the same parameters of adjacent consecutive layers to be the same, and then increases the sharing layer by using the remaining extra computing power while keeping the overall model parameter quantity unchanged;
[0027] Next, the parameter layers to be shared are divided into several independent groups. The parameters of each layer within each group remain independent, but these groups also use the remaining extra computing power in the network in a jumping (skip) manner and are repeatedly used to enhance the expression ability of the model.
[0028] In a specific implementable solution, under the selected Seq-Skip parameter sharing strategy, an N-layer DFSMN acoustic model is generated, and the DFSMN acoustic model after parameter sharing is trained and performance tested to determine whether all indicators of the model meet the expected requirements:
[0029] Start the training process according to the constructed model structure, continuously record and monitor the key performance indicators during the training process. If all indicators meet the design requirements after sufficient training and testing, the entire process ends;
[0030] If the test results show that the model fails to meet the requirements on certain key indicators, it is necessary to enter the feedback adjustment stage, that is, reselect some parameters of each layer in the DFSMN acoustic model to reconstruct the shared layer, and execute the Seq-Skip parameter sharing strategy again for parameter allocation; retrain and perform performance testing on the model with the new parameter sharing configuration until all performance indicators of the model reach the predetermined standards.
[0031] In a second aspect, the present application provides a neural network training system based on model parameter sharing, adopting the following technical solution:
[0032] A neural network training system based on model parameter sharing includes:
[0033] A computing power calculation module, configured to calculate the computing power and the number of parameters of an M-layer DFSMN acoustic model already deployed on the edge side, appropriately increase the total computing power requirement based on the same number of parameters of the model, and determine whether there is remaining chip computing power;
[0034] A parameter selection module, configured to select some parameters of each layer of the DFSMN acoustic model to construct a shared layer if there is remaining chip computing power;
[0035] A parameter allocation module, configured to organize and allocate the constructed shared layer according to the Seq-Skip parameter sharing strategy;
[0036] A model training module, configured to generate an N-layer DFSMN acoustic model under the selected Seq-Skip parameter sharing strategy, train and perform performance testing on the DFSMN acoustic model after parameter sharing, and determine whether each index of the model meets the expected requirements.
[0037] In a third aspect, the present application provides an electronic device, the device includes a processor and a memory; a program is stored in the memory, and the program is loaded and executed by the processor to implement a neural network training method based on model parameter sharing as described in the first aspect.
[0038] In a fourth aspect, the present application provides a computer-readable storage medium, a program is stored in the storage medium, and the program is used to implement a neural network training method based on model parameter sharing as described in the first aspect when executed by a processor.
[0039] In summary, the beneficial effects of the present application at least include:
[0040] (1) Aiming at the limitations of memory and computing power of low-power voice chips, a new parameter sharing scheme is proposed while keeping the number of parameters of the original M-layer DFSMN acoustic model unchanged. Exploration is carried out on three different types of parameters: hidden, projection, and fsmn. By calculating the remaining computing power and selecting specific partial parameters of each layer to construct the corresponding shared layer, using the proposed Seq-Skip parameter sharing strategy, first share parameters between adjacent consecutive layers, and use the remaining extra computing power to add shared layers, keeping the overall number of model parameters unchanged so that the model is more suitable for running on devices with limited memory. Then, divide the parameter layers to be shared into several independent groups. The parameters within each group are independent, but these groups also use the remaining extra computing power in a jumping (skip) manner repeatedly in the network to enhance the model's expressive ability, allowing the model to learn different feature representations between different layer groups, which helps the model to better capture long-term dependencies when processing sequence data. By adding (N - M) shared layers (1 ≤ M ≤ N) and using the extra computing power, compared with acoustic models of the same size, the speech recognition accuracy can be effectively improved.
[0041] (2) The parameter sharing scheme can accelerate the model training process by reducing the size of the weight matrix of each layer. Especially in the case of resource constraints, the proposed Seq-Skip parameter sharing strategy can make the model more flexible, which helps the model's adaptability when processing different language pairs and tasks. It can also improve efficiency and performance in fields such as machine translation, providing a method to improve efficiency and performance for applications in natural language processing (NLP) tasks.
[0042] Aiming at the limitations of memory and computing power of low-power voice chips, the Seq-Skip parameter sharing scheme proposed in this application is based on exploratory experiments of the existing acoustic model of M-layer DFSMN deployed on the edge side. It can, without increasing the total original parameters, calculate the total computing power requirement after increasing the acoustic model DFSMN and the remaining computing power of the chip. At the same time, in order to avoid large dimensions of each layer of the model, that is, large weight matrices, which increase the computational complexity and time, only select partial parameters of each layer of DFSMN to construct the corresponding shared layer. Through the Seq-Skip parameter sharing strategy, add (N - M) parameter shared layers, and while keeping the memory occupancy unchanged, use the extra computing power to optimize the model structure to improve the speech recognition performance.
[0043] The above description is only an overview of the technical solution of this application. In order to be able to understand the technical means of this application more clearly and implement it according to the content of the description, the following takes the preferred embodiments of this application and combines with the attached drawings to elaborate in detail as follows. Brief Description of the Drawings
[0044] Figure 1It is a schematic diagram of the overall process of the neural network training method based on model parameter sharing in the embodiments of the present application.
[0045] Figure 2 It is a schematic diagram of use cases of the Seq-Skip sharing strategy under different parameters in the embodiments of the present application.
[0046] Figure 3 It is a block diagram of the structure of the neural network training system based on model parameter sharing in the embodiments of the present application.
[0047] Figure 4 It is a block diagram of an electronic device for neural network training based on model parameter sharing in the embodiments of the present application. Detailed implementation manners
[0048] The following combines the accompanying drawings and embodiments to further describe in detail the specific implementation manners of the present application. The following embodiments are used to illustrate the present application, but are not used to limit the scope of the present application.
[0049] Optionally, the present application takes the neural network training method based on model parameter sharing provided in each embodiment and used in an electronic device as an example for illustration. The electronic device is a terminal or a server. The terminal can be a mobile phone, a computer, a tablet computer, etc. The type of the electronic device is not limited in this embodiment.
[0050] Refer to Figure 1 , which is a schematic diagram of the process of the neural network training method based on model parameter sharing provided in an embodiment of the present application. The method at least includes the following steps:
[0051] Step S101: For the existing M-layer DFSMN acoustic model deployed on the edge side, calculate its computing power and the number of parameters, appropriately increase the total computing power requirement based on the same number of parameters of the model, and determine whether there is remaining chip computing power.
[0052] In step S101, the present application first calculates the computing power and the number of parameters of the existing M-layer DFSMN (Deep Feedforward Sequential Memory Network) acoustic model deployed on the edge side, and evaluates the remaining computing power of the chip to ensure that the design of the new parameter sharing model can be successfully implemented on a low-power voice chip.
[0053] Specifically, starting from each layer of the DFSMN model, for various types of parameters contained in each layer, the number of parameters is counted by analyzing its network structure. After counting the number of parameters in each layer, based on the number of multiplication and addition operations participated by each parameter in the forward propagation, the computational requirements of each layer are estimated, and the total computational power requirements of the entire DFSMN model are obtained by adding up the computational requirements of all layers. Next, according to the technical specifications of the low-power voice chip, starting from the theoretical peak of the chip and the computational power that can be allocated to the speech recognition task during actual operation, the remaining computational power of the chip in the target application scenario is evaluated. The key to judging whether there is remaining computational power in the chip lies in comparing the actual available computational power of the chip with the total computational power requirements of the DFSMN model: if there is no additional remaining computational power in the chip, that is, the actual available computational power is less than or equal to the model requirements, it is proved that the hardware resources are not sufficient to support the introduction of additional parameter sharing layers without increasing the total amount of model parameters, and the process ends directly. Otherwise, it is necessary to consider adjusting the configuration of the sharing layer or further optimizing the model in the subsequent design to meet the increased total computational power requirements. Through this detailed calculation and comparison process, not only the computational power consumption of the DFSMN model is clarified, but also a quantitative basis is provided for the hardware feasibility of the new parameter sharing strategy, thus ensuring that the entire technical solution can operate efficiently in a low-power environment.
[0054] Step S102: If there is remaining computational power in the chip, select some parameters of each layer of the DFSMN acoustic model to construct a sharing layer.
[0055] In step S102, after the evaluation result shows that the remaining computing power of the chip is sufficient to support the new model design, it enters the stage of selecting shared parameters. N represents the total number of layers of the finally constructed DFSMN model, and M represents the number of layers of the initial DFSMN model, where N is greater than or equal to M. First, for each of the original M layers of the DFSMN acoustic model, the parameters are divided according to actual needs, specifically into three parts: the hidden layer, the projection layer, and the feed-forward sequence memory layer (fsmn). In this process, independent parameters are prepared for each of the M layers, and these parameters will serve as the basis for constructing the entire N-layer model, so that the parameter allocation can be adjusted according to the actual needs of each layer, rather than simply using a uniform allocation method. Subsequently, according to the calculation results of the model parameters and the remaining computing power of the chip in the early stage, some of the parameters in each layer of the DFSMN model are selectively shared, that is, the corresponding parameter sharing layer is constructed, and the selected shared parameters are allocated to the N layers. Specifically, the hidden layer part uses a fully connected structure with the ReLU activation function; the projection layer uses a low-rank linear mapping matrix, and its working principle is similar to that of a fully connected layer without an activation function; while the feed-forward sequence memory layer focuses on modeling historical and future temporal information and integrates this information into a fixed-dimensional encoding. Based on the limitations of the memory and computing power of the low-power voice chip, the proposed Seq-Skip parameter sharing scheme is developed based on the experimental exploration of the M-layer DFSMN acoustic model. By only selecting some of the parameters of each layer to construct the sharing layer, it not only avoids the problems of excessive weight matrices, increased computational complexity and time caused by the excessive dimension of each layer, but also effectively improves the overall model performance using the remaining computing power of the chip.
[0056] It should be noted that the memory module (block) in this application includes the above three layers of parameters. After the outputs of these three parts are added together, they are linearly mapped and input into the next hidden layer, and then the skip connection is introduced, so that the deeper network structure can be better utilized to improve the model recognition effect. In addition, the introduction of skip-frame context temporal modeling can effectively remove the redundancy between adjacent frames, which has more obvious advantages than BLSTM.
[0057] In implementation, when constructing the shared layer, M represents the number of independent parameters in each layer of the original DFSMN acoustic model. That is to say, the independent parameters form the basis for constructing the shared layer. By selecting some parameters of each layer to construct the shared layer, it is actually extracting the key parameters from these M layers and reusing these parameters throughout the model. In other words, when constructing an N-layer model (where N is greater than or equal to M), only the independent parameters of M layers need to be prepared, and the remaining (N - M) layers are formed by sharing these parameters. That is to say, M determines how many unique parameter sets there are, and the shared layer uses these parameters as "modules" and repeatedly distributes them throughout the network, achieving continuous sharing and skip sharing through the Seq-Skip parameter sharing strategy, so as to expand to N layers without increasing the total number of parameters. Therefore, M not only represents the number of layers of the original independent parameters, but also limits the types of parameters available when constructing the shared layer, ultimately enabling the entire model to maintain both depth (N layers) and control the total number of parameters.
[0058] Step S103: Organize and distribute the constructed shared layer according to the Seq-Skip parameter sharing strategy.
[0059] In step S103, after constructing the specific shared layer, organize and distribute the shared layer according to the Seq-Skip parameter sharing strategy. Specifically, the Seq-Skip parameter sharing strategy first targets a certain type of parameter selected in the DFSMN acoustic model. By setting the same parameters of consecutive layers to be the same, the number of independent parameters is relatively reduced while keeping the total number of parameters of the overall model unchanged. Then, the parameter layers to be shared are divided into several independent groups. The parameters of each layer within each group remain independent, but these groups also repeatedly use the remaining extra computing power in the network in a skip manner to enhance the expression ability of the model. This method allows only the independent parameters of M layers (where 1 ≤ M ≤ N) to be prepared when constructing an N-layer model, achieving efficient sharing of parameters without increasing the total original number of parameters.
[0060] Take Figure 2 (a) the hidden parameter sharing as an example. When setting M = 4 and N = 8, according to the Seq-Skip sharing strategy, the hidden parameters of the first, second, and fourth layers share the same set of parameters, while the fifth, sixth, and eighth layers share another set of parameters. Every four layers form a parameter sharing block, so the entire 8-layer DFSMN model can be regarded as composed of two sharing blocks. Through this way of combining continuous sharing and skip sharing, it not only reduces the memory occupation of the model while ensuring the model depth and the complexity of the memory module, but also maintains the diversity of parameters between different layer groups, which helps the model better capture long-term dependencies and complex data features, thereby improving the performance and expression ability of the overall model.
[0061] Step S104: Train and perform performance testing on the DFSMN acoustic model under the selected Seq-Skip parameter sharing strategy, and determine whether the indicators of the model meet the expected requirements.
[0062] In step S104, after selecting a suitable parameter sharing strategy, comprehensively train the constructed DFSMN model and simultaneously conduct strict performance testing to determine whether the indicators of the model meet the expected requirements. Specifically, at this stage, first start the training process according to the constructed model structure, continuously record and monitor key performance indicators during the training process, including the memory occupancy of the model, computational efficiency, and speech recognition accuracy, etc. If all indicators meet the design requirements after sufficient training and testing, the entire process ends and the model can be directly applied to the actual scenario; however, if the test results show that the model fails to meet the requirements in some key indicators, it is necessary to enter the feedback adjustment stage, that is, reselect some parameters of each layer in the DFSMN acoustic model to reconstruct the shared layer, and execute the Seq-Skip parameter sharing strategy again for parameter allocation. Subsequently, retrain and perform performance testing on the model using the new parameter sharing configuration, and this process will continue to cycle until the performance indicators of the model reach the predetermined standard. Through this closed-loop mechanism of repeated debugging and optimization, it is ensured that the finally constructed DFSMN model can not only make full use of the remaining computing power of the low-power voice chip, but also achieve the optimal balance of speech recognition performance and system efficiency without increasing the total amount of original parameters.
[0063] In summary, aiming at the limitations of the memory and computing power of the low-power voice chip, the Seq-Skip parameter sharing scheme proposed in this application is based on the exploration experiment of the acoustic model of the M-layer DFSMN already deployed on the edge side. Without increasing the total amount of original parameters, it calculates the total computing power requirement after increasing the acoustic model DFSMN and the remaining computing power of the chip. At the same time, in order to avoid the large dimension of each layer of the model, that is, the large weight matrix, which increases the computational complexity and time, only select some parameters of each layer of DFSMN to construct the corresponding shared layer. Through the Seq-Skip parameter sharing strategy, add (N - M) layers of parameter sharing layers (1 ≤ M ≤ N), and use the additional computing power to optimize the model structure to improve the speech recognition performance.
[0064] Compared with existing technologies such as SEQUENCE parameter sharing and model quantization, a Seq-Skip parameter sharing strategy for DFSMN acoustic models under different parameters is proposed. First, by setting the parameters of adjacent consecutive layers to be the same, the remaining extra computing power is used to increase the shared layers. Second, by grouping the parameter layers to be shared, the parameters within each group are independent, and then the remaining extra computing power is used to reuse these groups in a jumping manner to improve the expression ability of the model. This strategy allows us to prepare only the parameters of M layers (1 ≤ M ≤ N) when constructing an N-layer model. In addition, this flexibility enables different parameters to be shared between different layers, rather than simply sharing the parameters of a single layer between all layers. By reducing the size of the weight matrix of each layer, the computational complexity is significantly reduced, thereby reducing the inference time and improving the efficiency. Without changing the total number of parameters, the remaining computing power of low-power chips can be utilized to improve the ASR performance through the proposed Seq-Skip parameter sharing scheme.
[0065] Figure 3 FIG. is a structural block diagram of a neural network training system based on model parameter sharing provided by an embodiment of the present application. The device at least includes the following modules:
[0066] A computing power calculation module, configured to calculate the computing power and the number of parameters of an M-layer DFSMN acoustic model that has been deployed on the edge side, appropriately increase the total computing power requirement based on the same number of parameters of the model, and determine whether there is remaining chip computing power;
[0067] A parameter selection module, configured to select partial parameters of each layer of the DFSMN acoustic model to construct a shared layer if there is remaining chip computing power;
[0068] A parameter allocation module, configured to organize and allocate the constructed shared layer according to the Seq-Skip parameter sharing strategy;
[0069] A model training module, configured to generate an N-layer DFSMN acoustic model under the selected Seq-Skip parameter sharing strategy, train and perform performance tests on the DFSMN acoustic model after parameter sharing, and determine whether the indicators of the model meet the expected requirements.
[0070] For relevant details, refer to the above method embodiment.
[0071] Figure 4 FIG. is a block diagram of an electronic device provided by an embodiment of the present application. The device at least includes a processor 401 and a memory 402.
[0072] The processor 401 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 401 may be implemented in at least one of the following hardware forms: DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 401 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 401 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 401 may further include an AI (Artificial Intelligence) processor, which is used to process computational operations related to machine learning.
[0073] The memory 402 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 402 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 402 is used to store at least one instruction, and the at least one instruction is to be executed by the processor 401 to implement the neural network training method based on model parameter sharing provided in the method embodiments of the present application.
[0074] In some embodiments, the electronic device may further optionally include: a peripheral device interface and at least one peripheral device. The processor 401, the memory 402, and the peripheral device interface may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface through a bus, signal lines, or a circuit board. Schematically, the peripheral devices include, but are not limited to: a radio frequency circuit, a touch display screen, an audio circuit, and a power supply, etc.
[0075] Of course, the electronic device may also include fewer or more components, and this embodiment does not limit this.
[0076] Optionally, the present application further provides a computer-readable storage medium, and a program is stored in the computer-readable storage medium, and the program is loaded and executed by the processor to implement the neural network training method based on model parameter sharing in the above method embodiments.
[0077] Optionally, the present application further provides a computer product, which includes a computer-readable storage medium storing a program that is loaded and executed by a processor to implement the neural network training method based on model parameter sharing in the above method embodiments.
[0078] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered that the scope described in this specification.
[0079] The above embodiments only represent several implementation manners of the present application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A neural network training method based on model parameter sharing, characterized in that The method includes: For the M-layer DFSMN acoustic model already deployed on the edge side, calculate its computing power and the number of parameters, appropriately increase the total computing power requirement based on the same number of parameters of this model, and determine whether there is remaining chip computing power; If there is remaining chip computing power, select partial parameters of each layer of the DFSMN acoustic model to construct a shared layer; Organize and allocate the constructed shared layer according to the Seq-Skip parameter sharing strategy; Generate an N-layer DFSMN acoustic model under the selected Seq-Skip parameter sharing strategy, train and perform performance testing on the DFSMN acoustic model after parameter sharing, and determine whether each index of the model meets the expected requirements.
2. The neural network training method based on model parameter sharing according to claim 1, wherein The step of "For the M-layer DFSMN acoustic model already deployed on the edge side, calculate its computing power and the number of parameters, appropriately increase the total computing power requirement based on the same number of parameters of this model, and determine whether there is remaining chip computing power" includes: Starting from each layer of the DFSMN model, for various types of parameters included in each layer, analyze its network structure to count the number of parameters; After counting the number of parameters of each layer, estimate the amount of computation required for each layer based on the number of multiply-accumulate operations participated by each parameter in the forward propagation, and add up the computational requirements of all layers to obtain the total computing power requirement of the entire DFSMN model.
3. The neural network training method based on model parameter sharing according to claim 2, wherein The step of "For the M-layer DFSMN acoustic model already deployed on the edge side, calculate its computing power and the number of parameters, appropriately increase the total computing power requirement based on the same number of parameters of this model, and determine whether there is remaining chip computing power" also includes: Evaluate the remaining extra computing power of the chip in the target application scenario, and compare the actual available computing power of the chip with the total computing power requirement of the DFSMN model; If the chip has no extra remaining computing power, that is, the actual available computing power is less than or equal to the model requirement, directly end.
4. The neural network training method based on model parameter sharing according to claim 1, wherein The step of "If there is remaining chip computing power, select partial parameters of each layer of the DFSMN acoustic model to construct a shared layer" includes: For each layer in the original M-layer DFSMN acoustic model, divide the parameters into three parts: hidden layer, linear mapping layer, and feed-forward sequence memory layer; Prepare independent parameters for each of the M layers. According to the calculation results of the model parameters and the remaining chip computing power in the early stage, selectively share partial parameters in each layer of the DFSMN model to construct the corresponding parameter sharing layer, and allocate the selected shared parameters to the N layers; N represents the total number of layers of the finally constructed DFSMN model, M represents the number of layers of the initial DFSMN model, and N is greater than or equal to M.
5. The neural network training method based on model parameter sharing according to claim 4, wherein The step of "For each layer in the original M-layer DFSMN acoustic model, divide the parameters into three parts: hidden layer, linear mapping layer, and feed-forward sequence memory layer" includes: The hidden layer part adopts a fully connected structure, and the activation function is ReLU; The linear mapping layer adopts a low-rank linear mapping matrix; The feed-forward sequence memory layer models historical and future temporal information and integrates it into a fixed-dimensional encoding.
6. The neural network training method based on model parameter sharing according to claim 1, wherein, The step of "Organize and allocate the constructed shared layer according to the Seq-Skip parameter sharing strategy" includes: For a certain type of parameters selected in the DFSMN acoustic model, the Seq-Skip parameter sharing strategy first sets the same parameters in adjacent consecutive layers to be consistent, and uses the remaining extra computing power to add shared layers while keeping the overall model parameter quantity unchanged; Next, the parameter layers to be shared are divided into several independent groups. The parameters within each group are kept independent, but these groups are also repeatedly used in a skip manner in the network using the remaining extra computing power to enhance the model's expressive ability.
7. The neural network training method based on model parameter sharing according to claim 1, characterized in that, Under the selected Seq-Skip parameter sharing strategy, an N-layer DFSMN acoustic model is generated, and the DFSMN acoustic model after parameter sharing is trained and performance-tested to determine whether the various model indicators meet the expected requirements: Start the training process according to the constructed model structure, continuously record and monitor the key performance indicators during the training process. If all indicators meet the design requirements after sufficient training and testing, the entire process ends; If the test results show that the model fails to meet the requirements on some key indicators, it is necessary to enter the feedback adjustment stage, that is, re-select some parameters of each layer in the DFSMN acoustic model to reconstruct the shared layer, and execute the Seq-Skip parameter sharing strategy again for parameter allocation; re-train and performance-test the model with the new parameter sharing configuration until the various performance indicators of the model reach the predetermined standard.
8. A neural network training system based on model parameter sharing, characterized in that Including: A computing power calculation module, for an M-layer DFSMN acoustic model already deployed on the edge side, calculates its computing power and parameter quantity, appropriately increases the total computing power requirement based on the same parameter quantity of this model, and determines whether there is remaining chip computing power; A parameter selection module, used to select some parameters of each layer of the DFSMN acoustic model to construct a shared layer if there is remaining chip computing power; A parameter allocation module, used to organize and allocate the constructed shared layer according to the Seq-Skip parameter sharing strategy; A model training module, used to generate an N-layer DFSMN acoustic model under the selected Seq-Skip parameter sharing strategy, train and performance-test the DFSMN acoustic model after parameter sharing, and determine whether the various model indicators meet the expected requirements.
9. An electronic device, characterized in that, The device includes a processor and a memory; a program is stored in the memory, and the program is loaded and executed by the processor to implement a neural network training method based on model parameter sharing according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A program is stored in the storage medium, and when the program is executed by a processor, it is used to implement a neural network training method based on model parameter sharing according to any one of claims 1 to 7.