A method, system, device and storage medium for keyword recognition

Through hypernetwork training and hardware constraint optimization, automatic search for the best neural network structure and accelerator design solves the problems of high cost and long cycle of keyword recognition neural network design in the prior art, and achieves lower cost and more efficient keyword recognition.

CN114724553BActive Publication Date: 2025-06-10SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210295370.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-24
Publication Date
2025-06-10
Estimated Expiration
2042-03-24

AI Technical Summary

Technical Problem

The existing keyword recognition neural network framework has high design costs and long cycles, making it difficult to obtain global optimal results on the software and hardware side while ensuring low design costs.

Method used

By obtaining audio signals and hardware constraints for preprocessing, the candidate neural network architecture is obtained using hypernetwork training, and the hardware architecture parameters are determined in combination with hardware constraints. Through evaluation functions, the neural network accelerator framework is finally built to identify keywords.

Benefits of technology

It realizes the accuracy of keyword recognition while reducing development costs, and solves the problem of high and long cost of manual keyword recognition neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114724553B_ABST
    Figure CN114724553B_ABST
Patent Text Reader

Abstract

A method, system, device and medium for identifying keywords provided by the present invention mainly include the following steps: obtaining an audio signal and hardware constraints of a storage array, and preprocessing the audio signal to obtain a speech feature map; performing hypernetwork training according to the speech feature map, training to obtain a candidate neural network architecture, and determining network architecture parameters and weight parameters of the candidate neural network architecture; determining hardware architecture parameters according to the hardware constraints; determining a framework search space according to the hardware architecture parameters; updating the network architecture parameters and the hardware architecture parameters in the framework search space through an evaluation function; constructing a neural network accelerator framework according to the updated network architecture parameters, the updated hardware architecture parameters and the weight parameters, and outputting keywords in a target audio through the neural network accelerator framework, which can be widely applied to the technical field of machine learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and in particular to a method, system, device and medium for identifying keywords. Background Art

[0002] Keyword recognition refers to the process of identifying target keywords from an audio stream. Currently, popular keyword recognition systems generally rely on a neural network framework.

[0003] As Figure 1 shown, in the related art framework, the audio signal first passes through pre-processing of speech recognition such as a feature extraction module, and is converted into a spectrogram composed of feature vectors, and then is put into a neural network for continuous training and iteration of weight parameters, and the accuracy is obtained by means of inference on a validation set. Under the limitation of stopping conditions such as the number of iterations or the loss value, the best accuracy under a certain weight is obtained, and finally the weight parameters can be put into a neural network accelerator for verification. Among them, the pre-processing of speech recognition includes feature extraction algorithms such as Mel Frequency Cepstral Coefficients and Linear Predictive Cepstral Coefficients; common neural network structures include CNN, DNN, RNN, LSTM, etc. Since the training process of the neural network is realized by continuous backpropagation, the coordinated update between layers can make the system easier to optimize globally, and the recognition rate also has obvious advantages compared with the traditional Hidden Markov Model; after obtaining the best accuracy, the best weight parameters and network structure can be transplanted into hardware such as a neural network accelerator and an embedded system for actual speech recognition applications.

[0004] However, the above process is highly dependent on rich prior knowledge and choices of neural networks, and independent designs need to be carried out for different data sets and target tasks to obtain the most suitable network structure. The manually designed keyword recognition neural network and the corresponding accelerator require a large amount of manpower and material resources and a long design cycle.

[0005] At the same time, many of the manually designed neural networks only consider the standard of network accuracy, and do not consider problems such as power consumption, real-time performance and resource overhead faced by the hardware side in the actual scenario. This has led to difficulties and challenges for researchers to apply deep learning to specific tasks and platforms.

[0006] Therefore, if the traditional keyword recognition neural network framework is adopted, it is difficult to obtain the globally optimal results for both software and hardware while ensuring low design costs. Summary of the Invention

[0007] In view of this, in order to at least partially solve one of the above technical problems, an object of the embodiments of the present invention is to provide a method for identifying keywords with lower cost and better effect, as well as a corresponding system, device and storage medium capable of implementing the method.

[0008] On the one hand, the technical solution of the present application provides a method for identifying keywords, including the following steps:

[0009] Obtain an audio signal and hardware constraints of a storage array, and preprocess the audio signal to obtain a speech feature map;

[0010] Perform hypernetwork training based on the speech feature map, train to obtain a candidate neural network architecture, and determine network architecture parameters and weight parameters of the candidate neural network architecture;

[0011] Determine hardware architecture parameters according to the hardware constraints;

[0012] Determine a framework search space according to the hardware architecture parameters;

[0013] Update the network architecture parameters and the hardware architecture parameters in the framework search space through an evaluation function;

[0014] Construct a neural network accelerator framework according to the updated network architecture parameters, the updated hardware architecture parameters, and the weight parameters, and output keywords in the target audio through the neural network accelerator framework.

[0015] In a feasible embodiment of the solution of the present application, the step of obtaining an audio signal and hardware constraints of a storage array, and preprocessing the audio signal to obtain a speech feature map includes:

[0016] Perform pre-emphasis on the audio signal, and then frame the pre-emphasized audio signal to obtain a long-term speech signal;

[0017] Perform windowing on the long-term speech signal to obtain a short-term speech signal;

[0018] Perform a fast Fourier transform on the short-term speech signal to obtain frequency components;

[0019] Perform conversion according to the frequency components to obtain a Mel frequency signal, and perform cepstrum operation on the Mel frequency to obtain the speech feature map.

[0020] In a feasible embodiment of the solution of the present application, the step of performing hypernetwork training based on the speech feature map, training to obtain a candidate neural network architecture, and determining network architecture parameters and weight values of the candidate neural network architecture includes:

[0021] Generate binary architecture parameters, and determine an activation path according to the binary architecture parameters;

[0022] Construct a training set based on the speech feature map, and train the weight parameters of the activation path according to the training set by the stochastic gradient descent method;

[0023] Fix the weight parameters, construct a validation set according to the speech feature map, and determine the network architecture parameters through the validation set;

[0024] Generate a candidate neural network architecture by sampling binary gates according to the network architecture parameters.

[0025] In a feasible embodiment of the solution of the present application, the determining the framework search space according to the hardware architecture parameters includes at least one of the following steps:

[0026] Determine the framework search space according to the parallelism of the computing array;

[0027] Determine the framework search space according to the reuse method of the input feature map;

[0028] Determine the framework search space according to the operator type of the neural network.

[0029] In a feasible embodiment of the solution of the present application, the step of updating the network architecture parameters and the hardware architecture parameters through an evaluation function in the framework search space includes:

[0030] Obtain the accuracy of the candidate neural network architecture;

[0031] Determine the hardware latency according to the hardware architecture parameters;

[0032] Calculate the evaluation function score according to the accuracy and the hardware latency;

[0033] Iteratively update the network architecture parameters and the hardware architecture parameters according to the evaluation function score.

[0034] In a feasible embodiment of the solution of the present application, the step of determining the hardware latency according to the hardware architecture parameters includes:

[0035] Determine the hardware latency according to the sum of the computational latency and the off-chip memory access latency;

[0036] The computational latency is obtained by dividing the sum of the single-layer computational amounts by the effective computational parallelism; the off-chip memory access latency is obtained by multiplying the total memory access data volume by the data bit width and dividing by the bandwidth.

[0037] In a feasible embodiment of the solution of the present application, the step of constructing a neural network accelerator framework based on the updated network architecture parameters, the updated hardware architecture parameters, and the weight parameters, and outputting keywords in the target audio through the neural network accelerator framework includes:

[0038] Construct a target neural network according to the updated network architecture parameters;

[0039] Determine the data reuse mode and the computing parallelism according to the updated hardware architecture parameters;

[0040] Input the weight parameters and the target feature map data into the target neural network;

[0041] Control the target neural network to perform calculations according to the data reuse mode and the computing parallelism.

[0042] On the other hand, the technical solution of the present application also provides a keyword recognition system, and the system includes:

[0043] An information acquisition unit; used to acquire audio signals and the hardware constraints of the storage array;

[0044] A neural network accelerator software and hardware co-optimization unit; used to preprocess the audio signal to obtain a speech feature map; determine the hardware architecture parameters according to the hardware constraints; and perform hypernetwork training according to the speech feature map, train to obtain a candidate neural network architecture, and determine the network architecture parameters and weight parameters of the candidate neural network architecture; determine the framework search space according to the hardware architecture parameters; update the network architecture parameters and the hardware architecture parameters in the framework search space through an evaluation function;

[0045] A reconfigurable accelerator design unit; used to construct a neural network accelerator framework according to the updated network architecture parameters, the updated hardware architecture parameters, and the weight parameters, and output keywords in the target audio through the neural network accelerator framework.

[0046] On the other hand, the technical solution of the present application also provides a keyword recognition device, and the device includes:

[0047] At least one processor;

[0048] At least one memory, used to store at least one program;

[0049] When the at least one program is executed by the at least one processor, the at least one processor runs a keyword recognition method as described in any item of the first aspect.

[0050] On the other hand, the technical solution of the present application also provides a storage medium storing a program executable by a processor, and the program executable by the processor is used to execute a method for identifying a keyword as described in any one of the first aspects when executed by the processor.

[0051] The advantages and beneficial effects of the present invention will be partially given in the following description, and the other parts can be understood through the specific implementation manners of the present invention:

[0052] The method for identifying a keyword proposed by the technical solution of the present application reflects the automation characteristics of the design through the neural network architecture search method and the reconfigurable accelerator design respectively, and considers both the accuracy of the neural network and the hardware delay of the hardware architecture and other software and hardware indicators in the design of the algorithm end. Users can use this method to automatically and comprehensively obtain the optimal neural network structure required for the target task and the corresponding accelerator design; the solution considers the neural network accuracy at the algorithm end and the hardware resource overhead at the accelerator end, comprehensively considers multiple performance indicators of software and hardware in practical applications, reduces the development cost while improving the accuracy of keyword recognition, and can effectively solve the problems of high cost and long cycle in manually designing a keyword recognition neural network. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0054] Figure 1 It is a block diagram of a keyword recognition neural network provided in the related art;

[0055] Figure 2 It is a flowchart of the steps of a method for identifying a keyword according to the technical solution of the present application;

[0056] Figure 3 It is a flowchart of the steps of preprocessing for speech recognition in the technical solution of the present application;

[0057] Figure 4 It is a flowchart of the steps of the super network training process in the technical solution of the present application;

[0058] Figure 5 It is a flowchart of the steps of neural network architecture search in the technical solution of the present application;

[0059] Figure 6 It is a flowchart of the steps of performance evaluation of software and hardware co-search in the technical solution of the present application;

[0060] Figure 7 This is a schematic structural diagram of a reconfigurable accelerator design unit in the technical solution of this application;

[0061] Figure 8 This is a system structure diagram of a keyword recognition method provided in the technical solution of this application. Specific implementation manners

[0062] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, in which the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention. For the step numbers in the following embodiments, they are only set for the convenience of description and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0063] Based on what is pointed out in the background art, in the related art, there are generally disadvantages of high cost and insufficient consideration of software and hardware in the design process of keyword recognition neural network frameworks. This application proposes a software and hardware collaborative automatic optimization scheme for keyword recognition neural networks. In the design of the algorithm end (software end) and the accelerator end (hardware end) of the scheme, the automation characteristics of the design are respectively reflected through the neural network architecture search method and the reconfigurable accelerator design. And in the design of the algorithm end, software and hardware indicators such as the accuracy of the neural network and the hardware delay of the hardware architecture are considered at the same time. Users can automatically and comprehensively obtain the best neural network structure required for the target task and the corresponding accelerator design by adopting the embodiments provided by this scheme.

[0064] On the one hand, the technical solution of this application provides a method for recognizing keywords; this method mainly includes steps S100 - S600:

[0065] S100. Obtain an audio signal and the hardware constraints of the storage array, and preprocess the audio signal to obtain a speech feature map;

[0066] Specifically in the embodiment, first, the audio signal needs to be transformed into a speech feature map through pre - processing before speech recognition and directly input into the neural network - accelerator software and hardware collaborative optimization framework.

[0067] S200. Perform hyper - network training according to the speech feature map, train to obtain a candidate neural network architecture, and determine the network architecture parameters and weight parameters of the candidate neural network architecture;

[0068] Specifically in the embodiment, based on the speech feature map obtained in step S100, the neural network architecture search module inside the framework provided by the embodiment receives the data of the speech input feature map and starts training the super network. The neural network architecture search adopted in this embodiment is different from traditional neural networks. It needs to alternately train the weight parameters and architecture parameters. The weight parameters represent the weights and biases in the convolutional calculation, and the architecture parameters represent the selection rate of each operator in the super network. Through repeated training and inference, this framework can automatically search for the optimal network structure.

[0069] S300. Determine the hardware architecture parameters according to the hardware constraints; that is, through a necessary user interaction process, obtain the hardware architecture parameters of the embodiment architecture.

[0070] Specifically in the embodiment, since hardware metrics need to be considered during the training phase, the embodiment establishes a delay model for the accelerator parallelism parameter and the reuse method, takes the accelerator delay as the optimization target at the hardware end, and puts it into the evaluation function together with the inference accuracy obtained from each training for evaluation.

[0071] S400. Determine the framework search space according to the hardware architecture parameters.

[0072] Specifically in the embodiment, for the proposed hardware requirements, the embodiment can select the required design space, including the parallelism of the computing array, the reuse method of the input feature map, and the operator type of the neural network.

[0073] S500. Update the network architecture parameters and the hardware architecture parameters in the framework search space through the evaluation function;

[0074] Specifically in the embodiment, the accelerator delay can be used as the optimization target at the hardware end and put into the evaluation function together with the inference accuracy obtained from each training for evaluation. After the evaluation, it is judged whether to search again according to the quality of each evaluation result. After repeatedly iterating the above process, the optimal network structure considering both the neural network accuracy and the hardware delay and the optimal hardware design scheme can be obtained. It should be noted that the inference progress in the embodiment is the accuracy value obtained by inferring and predicting the current (candidate) neural network architecture through the test set.

[0075] S600. Construct a neural network accelerator framework based on the updated network architecture parameters, the updated hardware architecture parameters, and the weight parameters, and output the keywords in the target audio through the neural network accelerator framework;

[0076] Specifically, in the embodiment, the network structure parameters and the hardware design scheme are input into the reconfigurable accelerator design module, which automatically modifies the accelerator configuration according to the hardware design scheme to obtain the corresponding accelerator RTL code. At the same time, the module trains the best network obtained by searching and quantizes the optimal weights after training. The quantized weights and the corresponding accelerator templates can be used for specific keyword language recognition applications.

[0077] As Figure 2 shown, for the embodiment of the keyword recognition method provided in this application, the complete implementation process is described as follows:

[0078] First, an audio signal needs to be input, usually selected as a.wav format file in the speech dataset. Then, the user can set the target hardware constraints at the algorithm end, such as the number of DSPs at the accelerator end, the on-chip cache size, etc. At the same time, the user needs to set the initial parameters of the super network in the neural network architecture search, including the required number of super network layers, the number of convolutional calculation channels, and the initial learning rate, etc. Finally, for the hardware requirements, the user can select the required design space, including the parallelism of the computing array, the reuse method of the input feature map, and the operator type of the neural network. After the above initial settings are completed, the optimization framework will automatically perform the neural search work to obtain the best network model that takes into account both software and hardware parameters and the optimal hardware design scheme, and input them into the reconfigurable accelerator module to generate the corresponding accelerator template and weight parameters.

[0079] It should be noted that the technical solution of this application can modify the speech recognition preprocessing to other processing units and apply this optimization method to different fields. Moreover, there are many neural network architecture search methods and evaluation function schemes in the optimization framework, and different combinations can be used to achieve similar effects.

[0080] In some alternative embodiments, as Figure 3 shown, in the embodiment method, the Mel Frequency Cepstral Coefficient (MFCC) algorithm is used for speech recognition preprocessing. It converts the audio signal into a low-dimensional frequency domain signal and can well extract the feature information in the audio. The traditional MFCC feature extraction mainly includes: pre-emphasis, framing, windowing, Fast Fourier Transform (FFT), Mel filtering, taking the logarithm, and Discrete Cosine Transform (DCT). Furthermore, for the step S100 of obtaining the audio signal and the hardware constraints of the storage array and preprocessing the audio signal to obtain the speech feature map in the method, it may include steps S110 - S140:

[0081] S110. Pre-emphasize the audio signal, and then frame the pre-emphasized audio signal to obtain long-term speech signals. Specifically, in the embodiment, pre-emphasis adds a high-pass filter to the audio, making the spectrum look flatter. Framing cuts the speech sequence into a finite number of equal-length short-term speech sequences.

[0082] S120. Window the long-term speech signal to obtain short-term speech signals. Among them, windowing can make the edges of the short-term speech signals after framing smooth and ensure the continuity of the endpoints.

[0083] S130. Perform a fast Fourier transform on the short-term speech signal to obtain frequency components. The fast Fourier transform (FFT) greatly reduces the computational complexity of the Fourier transform, and its purpose is to convert the time-domain signal into the distribution of frequency components on the time axis. It should be noted that in the embodiment, the short-time Fourier transform (STFT) as shown in Figure 3 can also be used to obtain frequency components.

[0084] S140. Convert according to the frequency components to obtain Mel frequency signals, and perform cepstrum operations on the Mel frequencies to obtain the speech feature map. In the embodiment, in order to adapt to the human ear's hearing, it is necessary to convert the actual frequencies of the speech sequence into Mel frequencies that simulate the human ear's hearing. Since the characteristics of the sound signal are mainly distributed in the low-frequency part of the energy spectrum, in order to directly obtain low-frequency information, it is necessary to take the logarithm of the Mel energy and perform a DCT cepstrum operation, and finally obtain a speech feature map containing keyword feature parameters.

[0085] In some alternative embodiments, for step S200 of training a super network according to the speech feature map, training to obtain a candidate neural network architecture, and determining the network architecture parameters and weight values of the candidate neural network architecture, it may include steps S210 - S240:

[0086] S210. Generate binary architecture parameters, and determine the activation path according to the binary architecture parameters.

[0087] S220. Construct a training set according to the speech feature map, and train the weight parameters of the activation path through the random gradient descent method according to the training set.

[0088] S230. Fix the weight parameters, construct a validation set according to the speech feature map, and determine the network architecture parameters through the validation set.

[0089] S240. Generate a candidate neural network architecture according to the network architecture parameters by sampling binary gates.

[0090] Specifically, in the embodiment, as Figure 4As shown, when training the example supernetwork, a method of repeatedly iterating the weight parameters and the architecture parameters is adopted. The complete training process is as follows:

[0091] First, take the input feature map data for keyword recognition, activate one path according to the randomly generated binary architecture parameters, and freeze other paths at the same time. After fixing the architecture parameters, use the stochastic gradient descent (SGD) method to train the weight parameters of this path on the training set. After fixing the trained weight parameters, use the validation set to update the network architecture parameters, and obtain the current architecture by sampling the binary gates. And keep looping through the steps of training the weight parameters and training the network architecture parameters.

[0092] In addition, as Figure 5 shown, the optimization algorithm of the technical solution of this application mainly makes changes on the basis of neural network architecture search (NAS). Therefore, in step S400 of determining the framework search space according to the network architecture parameters and the hardware architecture parameters, the example method may further include steps S410 - S430:

[0093] S410. Determine the framework search space according to the parallelism of the computing array;

[0094] S420. Determine the framework search space according to the reuse method of the input feature map;

[0095] S430. Determine the framework search space according to the operator type of the neural network.

[0096] Specifically in the example, when using the neural network architecture search framework, it is necessary to adopt a search strategy to search for the network structure in the search space. The result will be put into the performance evaluator to obtain the accuracy of the model. Subsequently, the returned accuracy will guide the search strategy to converge to a better structure in the search space, and the optimal network structure will be obtained after repeated iterations. In order to consider the hardware architecture parameters at the same time, for example: accelerator parameters, the example also needs to expand the framework search space, that is, add restrictions such as accelerator parallelism parameters, the reuse method of feature maps, the operator type of the neural network, on-chip cache size, and on-chip cache bandwidth to the framework search space. Through this method, the neural network search can automatically generate the neural network structure while considering the software and hardware architecture parameters.

[0097] In some selectable examples, in step S500 of updating the network architecture parameters and the hardware architecture parameters by the evaluation function in the framework search space of the example method, steps S510 - S540 may be included:

[0098] S510. Obtain the accuracy of the candidate neural network architecture;

[0099] S520. Determine the hardware latency according to the hardware architecture parameters;

[0100] S530. Calculate the evaluation function score based on the accuracy and the hardware latency;

[0101] S540. Iteratively update the network architecture parameters and the hardware architecture parameters according to the evaluation function score.

[0102] Specifically, as Figure 6 shown, in the process of hypernetwork training, in order to evaluate the accuracy (precision) of the neural network on the software side and the latency on the hardware side simultaneously, the present invention uses an evaluation function based on reinforcement learning;

[0103] Put the hardware latency of the current architecture and the neural network accuracy corresponding to this architecture into the evaluation function of reinforcement learning, so that when updating the architecture weights, the accuracy of the validation set inference and the latency obtained from the hardware architecture search can be considered simultaneously:

[0104]

[0105] Among them, Reward is the evaluation function score, acc is the neural network accuracy of the current architecture, ref_latency and latency are the reference hardware latency and ratio that need to be set in advance, and latency is the hardware latency of the current architecture.

[0106] Furthermore, in step S530 of determining the hardware latency according to the hardware architecture parameters, it determines the hardware latency according to the sum of the computational latency and the off-chip memory access latency; among them, the computational latency is obtained by dividing the sum of the single-layer computational amounts by the effective computational parallelism; the off-chip memory access latency is obtained by multiplying the total memory access data volume by the data bit width and dividing by the bandwidth.

[0107] Specifically, in the hardware latency model, the total latency latency is:

[0108] latency = latency PP + latency MA

[0109] The latency latency of a single-layer calculation OP is equal to the sum of the single-layer computational amounts TotalOperations divided by the effective computational parallelism, that is, it satisfies:

[0110]

[0111] Among them, PE is the parallelism, O represents the output feature map of a single layer, I represents the input feature map of a single layer, K represents the convolution kernel, and P represents the parallelism in each direction of the accelerator.

[0112] Ow , O h respectively represent the width and height of the output feature map; I c , O c respectively represent the number of channels of the input feature map and the output feature map; K w , K h respectively represent the width and height of the convolutional kernel; P ox , P oy , P ic , P oc respectively represent the computational parallelism in the xy direction of the output feature map, the output channel direction, and the input channel direction. PE_ represents the effective utilization rate of the PE array. That is, PE × PE_Utility represents the effective computational parallelism.

[0113] In addition, the off-chip memory access latency is equal to the total amount of data accessed Total Amount of Data multiplied by the data bit width DataWidth divided by the bandwidth Bandwidth, that is, it satisfies:

[0114]

[0115] In some alternative embodiments, the method of the embodiment constructs a neural network accelerator framework according to the updated network architecture parameters, the updated hardware architecture parameters, and the weight parameters. The step S600 of obtaining the keywords in the target audio through the neural network accelerator framework may include steps S610 - S640:

[0116] S610. Construct a target neural network according to the updated network architecture parameters;

[0117] S620. Determine the data reuse method and computational parallelism according to the updated hardware architecture parameters;

[0118] S630. Input the weight parameters and the target feature map data into the target neural network;

[0119] S640. Control the target neural network to perform calculations according to the data reuse method and the computational parallelism;

[0120] Specifically in the embodiment, as Figure 7 shown, the network architecture parameters and the hardware scheme (hardware architecture parameters) will be passed into the processing system in an encoded form, and then the programmable logic module of the accelerator will be initialized and configured by the central processor. At the same time, the weight parameters will be quantized in the processing system in the INT8 format to be more conveniently passed into the accelerator for verification.

[0121] The configuration register in the programmable logic module receives the initialization configuration signal transmitted from the CPU through the AXI-lite bus. According to the parameters in the hardware solution, it automatically modifies the data multiplexing method and the computing parallelism in the top-level file, and guides the global controller to start controlling the computing process of the neural network. The optimal weight parameters transmitted from the software and hardware optimization module to the processing system will be transmitted to the weight cache in the programmable logic module through the AXI-4 bus after quantization. Under the control of the storage controller, they are transmitted to the computing unit array together with the input feature map data in data cache A for the convolution calculation of the neural network. The final calculation result will be transmitted out through data cache B and external storage to obtain the required keyword recognition probability.

[0122] On the other hand, as Figure 8 shown, the technical solution of this application also provides a keyword recognition system, which includes:

[0123] An information acquisition unit; used to acquire the audio signal and the hardware constraints of the storage array;

[0124] A neural network accelerator software and hardware co-optimization unit; used to preprocess the audio signal to obtain a speech feature map; determine the hardware architecture parameters according to the hardware constraints; and perform hypernetwork training according to the speech feature map, train to obtain a candidate neural network architecture, and determine the network architecture parameters and weight parameters of the candidate neural network architecture; determine the framework search space according to the hardware architecture parameters; update the network architecture parameters and the hardware architecture parameters through an evaluation function in the framework search space;

[0125] A reconfigurable accelerator design unit; used to construct a neural network accelerator framework according to the updated network architecture parameters, the updated hardware architecture parameters, and the weight parameters, and output the keywords in the target audio through the neural network accelerator framework.

[0126] In the third aspect, the technical solution of this application also provides a keyword recognition device, which includes at least one processor; at least one memory, and this memory is used to store at least one program; when at least one program is executed by at least one processor, it causes at least one processor to run a keyword recognition method as in the first aspect.

[0127] The embodiment of the present invention also provides a storage medium in which a program is stored, and the program is executed by a processor to implement any one of the keyword recognition methods in the first aspect.

[0128] From the above specific implementation process, it can be summarized that the technical solution provided by the present invention has the following advantages or advantages compared with the prior art:

[0129] 1. The technical solution of this application proposes a software and hardware co - optimization method for keyword recognition neural networks, which is used to solve the problem of high cost and long cycle in manually designing keyword recognition neural networks.

[0130] 2. When searching for the neural network architecture in the technical solution of this application, the neural network accuracy on the algorithm side and the hardware resource overhead on the accelerator side are considered, comprehensively considering multiple performance indicators of software and hardware in practical applications.

[0131] In some alternative embodiments, the functions / operations mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the functions / operations involved, two consecutively shown blocks can actually be executed substantially simultaneously or the blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed method is not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated where the order of various operations is changed and where sub - operations described as part of a larger operation are executed independently.

[0132] In addition, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It can also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. Rather, considering the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the module will be understood within the ordinary skills of an engineer. Therefore, those skilled in the art can implement the present invention as set forth in the claims without undue experimentation. It can also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.

[0133] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definitional sequence of executable instructions for implementing logical functions, which can be specifically implemented in any computer - readable medium for use by or in connection with an instruction execution system, apparatus, or device (such as a computer - based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device).

[0134] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0135] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.

[0136] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without violating the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A method for identifying keywords, characterized in that, it includes the following steps: Obtain an audio signal and the hardware constraints of the storage array, and preprocess the audio signal to obtain a speech feature map; Perform hypernetwork training based on the speech feature map, train to obtain a candidate neural network architecture, and determine the network architecture parameters and weight parameters of the candidate neural network architecture; Determine the hardware architecture parameters according to the hardware constraints; Determine the framework search space according to the hardware architecture parameters; Update the network architecture parameters and the hardware architecture parameters in the framework search space through an evaluation function; Construct a neural network accelerator framework according to the updated network architecture parameters, the updated hardware architecture parameters, and the weight parameters, and output the keywords in the target audio through the neural network accelerator framework.

2. The method for identifying keywords according to claim 1, characterized in that, the step of obtaining the audio signal and the hardware constraints of the storage array, and preprocessing the audio signal to obtain a speech feature map includes: Perform pre-emphasis on the audio signal, and then frame the pre-emphasized audio signal to obtain a long-term speech signal; Perform windowing on the long-term speech signal to obtain a short-term speech signal; Perform fast Fourier transform on the short-term speech signal to obtain frequency components; Perform conversion according to the frequency components to obtain a Mel frequency signal, and perform cepstrum operation on the Mel frequency to obtain the speech feature map.

3. The method for identifying keywords according to claim 1, characterized in that, the step of performing hypernetwork training based on the speech feature map, training to obtain a candidate neural network architecture, and determining the network architecture parameters and weight values of the candidate neural network architecture includes: Generate binary architecture parameters, and determine the activation path according to the binary architecture parameters; Construct a training set according to the speech feature map, and train the weight parameters of the activation path through the training set by the stochastic gradient descent method; Fix the weight parameters, construct a validation set according to the speech feature map, and determine the network architecture parameters through the validation set; Generate a candidate neural network architecture according to the network architecture parameters by sampling binary gates.

4. The method for identifying keywords according to claim 1, characterized in that, determining the framework search space according to the hardware architecture parameters includes at least one of the following steps: Determine the framework search space according to the parallelism of the computing array; Determine the framework search space according to the reuse method of the input feature map; Determine the framework search space according to the operator type of the neural network.

5. The method for identifying keywords according to claim 1, characterized in that, the step of updating the network architecture parameters and the hardware architecture parameters in the framework search space through an evaluation function includes: Obtain the accuracy of the candidate neural network architecture; Determine the hardware latency according to the hardware architecture parameters; Calculate the evaluation function score according to the accuracy and the hardware latency; Iteratively update the network architecture parameters and the hardware architecture parameters according to the scores of the evaluation function.

6. A method for identifying keywords according to claim 5, wherein, the step of determining the hardware latency according to the hardware architecture parameters includes: determining the hardware latency according to the sum of the computational latency and the off-chip memory access latency; the computational latency is obtained by dividing the sum of the single-layer computational amounts by the effective computational parallelism; the off-chip memory access latency is obtained by multiplying the total memory access data volume by the data bit width and dividing by the bandwidth.

7. A method for identifying keywords according to any one of claims 1-6, wherein, the step of constructing a neural network accelerator framework according to the updated network architecture parameters, the updated hardware architecture parameters, and the weight parameters, and outputting the keywords in the target audio through the neural network accelerator framework includes: constructing a target neural network according to the updated network architecture parameters; determining the data reuse mode and the computational parallelism according to the updated hardware architecture parameters; inputting the weight parameters and the target feature map data into the target neural network; controlling the target neural network to perform calculations according to the data reuse mode and the computational parallelism.

8. A keyword identification system, wherein, the system includes: an information acquisition unit; configured to acquire an audio signal and the hardware constraints of the storage array; a neural network accelerator software and hardware co-optimization unit; configured to preprocess the audio signal to obtain a speech feature map; determine the hardware architecture parameters according to the hardware constraints; and perform hypernetwork training according to the speech feature map, train to obtain a candidate neural network architecture, and determine the network architecture parameters and weight parameters of the candidate neural network architecture; determine the framework search space according to the hardware architecture parameters; update the network architecture parameters and the hardware architecture parameters in the framework search space through an evaluation function; a reconfigurable accelerator design unit; configured to construct a neural network accelerator framework according to the updated network architecture parameters, the updated hardware architecture parameters, and the weight parameters, and output the keywords in the target audio through the neural network accelerator framework.

9. A keyword identification device, wherein, it includes: at least one processor; at least one memory for storing at least one program; when the at least one program is executed by the at least one processor, the at least one processor runs a method for identifying keywords according to any one of claims 1-7.

10. A storage medium storing a program executable by a processor, wherein, the program executable by the processor is used to run a method for identifying keywords according to any one of claims 1-7 when executed by the processor.

Citation Information

Patent Citations

  • Lightweight neural network voice keyword recognition method based on hierarchical quantification

    CN112786021A

  • Multi-stream target-speech detection and channel fusion

    US20200184985A1