A method, apparatus, device, and storage medium for building a neural network model

By setting up interfaces on the neural network layer, collecting and calculating weight information in real time, the problem of low efficiency in building neural network models in the existing technology is solved, and a more efficient model construction process is achieved.

CN114358260BActive Publication Date: 2025-06-03PENG CHENG LAB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111507466.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-10
Publication Date
2025-06-03
Estimated Expiration
2041-12-10

AI Technical Summary

Technical Problem

The existing technology is less efficient in building neural network models, and it is necessary to manually calculate the weight information of the neural network layer.

Method used

By setting up interfaces on the neural network layer, weight information is collected and calculated, and changing information is returned in real time during the construction process, parallel calculation of weight information and model construction are realized.

Benefits of technology

The speed of neural network model construction has been improved, and the overall efficiency has been improved by calculating weight information in parallel and building models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114358260B_ABST
    Figure CN114358260B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of neural networks, and specifically relates to a method, device, equipment and storage medium for building a neural network model. An interface is provided on the neural network layer in the present invention. This interface can be used to collect weight information during the process of building a neural network model based on the neural network layer. This interface can also calculate the change information of the weights during the process of building the neural network model and return the change information of the weights to the neural network model being built. Therefore, the setting of the interface in the present invention enables the calculation of weight information and the building of the neural network model to be parallel and synchronous, thereby improving the speed of building the neural network model in the present invention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of neural networks, and specifically relates to a method, device, equipment and storage medium for building a neural network model. Background Art

[0002] Neural network models can be used for face recognition, speech recognition, etc. When using a neural network to perform the above tasks, it is first necessary to build an independent and discrete neural network layer into a neural network model according to the task to be performed. Most of the neural network model building processes use two base classes to construct, namely the neural network abstract layer (Layer) and the model abstract layer (Model). Among them, the neural network abstract layer is used to collect the parameters corresponding to the neural network layer and define the connection relationship (network building information) of the neural network layer during the training of the neural network layer, and the model abstract layer is used to implement the connection of the neural network layer according to the defined connection relationship of the neural network layer. The functions of these two base classes are to implement the neural network reference layer through the Layer class, and then use the Model class to connect these Layers to build a deep learning model. In the above neural network model building process, it is necessary to manually calculate the weight information corresponding to the neural network layer, thereby reducing the efficiency of model building.

[0003] In summary, the efficiency of building a neural network model in the prior art is low.

[0004] Therefore, the prior art still needs to be improved and enhanced. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a method, device, equipment and storage medium for building a neural network model, which solves the problem of low efficiency in building a neural network model in the prior art.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] In the first aspect, the present invention provides a method for building a neural network model, which includes:

[0008] Obtain a target neural network model to be built, where the target neural network model is used to complete a target task;

[0009] According to the target neural network model, obtain each neural network layer corresponding to the target neural network model;

[0010] According to the output information in adjacent neural network layers, obtain the weight information between adjacent neural network layers;

[0011] Complete the building of the target neural network model according to the weight information.

[0012] In one implementation, obtaining the weight information between adjacent neural network layers based on the output information in the adjacent neural network layers includes:

[0013] Based on adjacent neural network layers, obtain the upper neural network layer and the lower neural network layer included in the adjacent neural network layers, and the output of the upper neural network layer is the input of the lower neural network layer;

[0014] Based on the upper neural network layer, obtain the upper output information corresponding to the upper neural network layer in the output information;

[0015] Based on the lower neural network layer, obtain the lower output information corresponding to the lower neural network layer in the output information;

[0016] Based on the upper output information and the lower output information, obtain the weight information between adjacent neural network layers.

[0017] In one implementation, obtaining the weight information between adjacent neural network layers based on the upper output information and the lower output information includes:

[0018] Based on the upper output information, obtain the number of neurons in the upper layer in the upper output information;

[0019] Based on the lower output information, obtain the number of neurons in the lower layer in the lower output information;

[0020] Based on the number of neurons in the upper layer and the number of neurons in the lower layer, obtain the weight data size in the weight information.

[0021] In one implementation, obtaining the weight data size in the weight information based on the number of neurons in the upper layer and the number of neurons in the lower layer includes:

[0022] Based on the number of neurons in the upper layer and the number of neurons in the lower layer, obtain the product result corresponding to the product of the number of neurons in the upper layer and the number of neurons in the lower layer,

[0023] Based on the product result, obtain the weight data size.

[0024] In one implementation, obtaining the weight information between adjacent neural network layers based on the output information in the adjacent neural network layers includes:

[0025] Obtain whether there is a corresponding initial weight between each pair of adjacent neural network layers;

[0026] When there is no corresponding initial weight between adjacent neural network layers, input the data information corresponding to the target task into the neural network layer serving as the input layer to obtain the output information in the adjacent neural network layers;

[0027] Obtain the weight information between adjacent neural network layers based on the output information in the adjacent neural network layers.

[0028] In one implementation, obtaining whether there is a corresponding initial weight between each pair of adjacent neural network layers includes:

[0029] Input the data information corresponding to the target task into the neural network layer serving as the input layer to obtain the output result information of each neural network layer;

[0030] Obtain whether there is a corresponding initial weight between each pair of adjacent neural network layers based on the output result information.

[0031] In one implementation, obtaining the neural network layers corresponding to the target neural network model based on the target neural network model includes:

[0032] Obtain the target component information corresponding to the target neural network model based on the target neural network model;

[0033] Obtain the network layer connection information corresponding to the target component information based on the target component information;

[0034] Connect the target components according to the network layer connection information to obtain the neural network layers corresponding to the target neural network model.

[0035] In one implementation, completing the construction of the target neural network model based on the weight information includes:

[0036] Obtain the weight data size in the weight information based on the weight information;

[0037] Initialize the weight parameters according to the weight data size;

[0038] Complete the construction of the target neural network model based on the weight after the initialization.

[0039] In one implementation, completing the construction of the target neural network model based on the weight parameters after the initialization includes:

[0040] Train the neural network layer according to the weight parameters after the initialization, and collect the weight parameters being trained during the training process;

[0041] Send the weight parameters to be trained to an optimizer to obtain the optimized weight parameters, and complete the construction of the target neural network model.

[0042] In one implementation, the step of sending the weight parameters to be trained to an optimizer to obtain the optimized weight parameters and complete the construction of the target neural network model includes:

[0043] During the process of training the neural network layer, collect the connection relationships corresponding to the neural network layer, where the connection relationships correspond to the weight parameters to be trained;

[0044] Send the connection relationships and the weight parameters to be trained to the optimizer, and the optimizer outputs the connection relationships and the optimized weight parameters corresponding to the connection relationships;

[0045] Complete the construction of the target neural network model according to the connection relationships and the optimized weight parameters corresponding to the connection relationships output by the optimizer.

[0046] In a second aspect, an embodiment of the present invention further provides an apparatus for a neural network model construction method, where the apparatus includes the following components:

[0047] A target neural network model acquisition module, configured to acquire a target neural network model to be constructed, where the target neural network model is used to complete a target task;

[0048] A neural network layer generation module, configured to obtain each neural network layer corresponding to the target neural network model according to the target neural network model;

[0049] A weight calculation module, configured to obtain weight information between adjacent neural network layers according to output information in adjacent neural network layers;

[0050] A model construction module, configured to complete the construction of the target neural network model according to the weight information.

[0051] In a third aspect, an embodiment of the present invention further provides a terminal device, where the terminal device includes a memory, a processor, and a neural network model construction program stored in the memory and executable on the processor. When the processor executes the neural network model construction program, the steps of the above-mentioned neural network model construction method are implemented.

[0052] Fourthly, an embodiment of the present invention further provides a computer-readable storage medium, on which a neural network model building program is stored. When the neural network model building program is executed by a processor, the steps of the above-mentioned neural network model building method are implemented.

[0053] Beneficial effects: An interface is provided on the neural network layer in the present invention. This interface can be used to collect weight information during the process of building a neural network model based on the neural network layer. This interface can also calculate the change information of the weights during the process of building the neural network model and return the change information of the weights to the neural network model being built. Therefore, the setting of the interface in the present invention enables the calculation of weight information and the building of the neural network model to be parallel and synchronous, thereby improving the speed of building the neural network model in the present invention. Description of the Drawings

[0054] Figure 1 is the overall flowchart of the present invention;

[0055] Figure 2 is the schematic hardware structure diagram of the present invention;

[0056] Figure 3 is the flowchart for checking the initialization of weight parameters of the present invention;

[0057] Figure 4 is the flowchart for calculating the value of weight parameters of the present invention;

[0058] Figure 5 is the internal structure principle block diagram of the terminal device provided by the embodiment of the present invention. Detailed Embodiments

[0059] The following describes the technical solutions in the present invention clearly and completely in conjunction with the embodiments and the drawings of the specification. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0060] It has been found through research that neural network models can be used for face recognition, speech recognition, etc. When using a neural network to perform the above tasks, it is first necessary to build an independent and discrete neural network layer into a neural network model according to the task to be performed. The process of building a neural network model mostly uses two base classes to construct, namely the neural network abstraction layer (Layer) and the model abstraction layer (Model). Among them, the neural network abstraction layer is used to collect the parameters corresponding to the neural network layer and define the connection relationship (networking information) of the neural network layer during the training of the neural network layer, and the model abstraction layer is used to implement the connection of the neural network layer according to the defined connection relationship of the neural network layer. The functions of these two base classes are respectively to implement the neural network reference layer through the Layer class, and then use the Model class to connect these Layers to build a deep learning model. During the above process of building a neural network model, it is necessary to manually calculate the weight information corresponding to the neural network layer, thereby reducing the efficiency of model building.

[0061] To solve the above technical problems, the present invention provides a method, device, equipment and storage medium for building a neural network model, which solves the problem of low efficiency in building a neural network model in the prior art. Specifically in implementation, first, according to the target neural network model to be built, the corresponding neural network layer is obtained, and then the weight information between the neural network layers is constructed. Since the neural network layer of the present invention is provided with an interface, this interface can be used to collect the weight information during the process of building a neural network model based on the neural network layer, and this interface can also calculate the change information of the weight during the process of building a neural network model and return the change information of the weight to the neural network model being built. Therefore, the setting of the interface of the present invention enables the calculation of weight information and the building of a neural network model to be parallel and synchronous, thereby improving the speed of building a neural network model of the present invention.

[0062] For example, the target neural network model to be built is a face recognition neural network model, which is a neural network model capable of face recognition. It requires a neural network layer A as the input layer, a neural network layer B as the middle layer, and a neural network layer C as the output layer. In this embodiment, interfaces are set on the neural network layer A, the neural network layer B, and the neural network layer C, and a program for calculating weight information is set in the interfaces. The weight information is the weight information between the neural network layer A and the neural network layer B, and the weight information between the neural network layer B and the neural network layer C. Building a neural network model means inputting data into the neural network layer A as the input layer. After the data is processed by the neural network layer B and the neural network layer C, the neural network layer C outputs data, and the weight information is adjusted according to the output data. In this embodiment, during the process of the neural network layer A, the neural network layer B, and the neural network layer C processing the input data, the weight information is iteratively calculated through the interfaces at the same time. Therefore, this embodiment enables the processing of input data and the calculation of weight information to be carried out in parallel, thereby improving the efficiency of building the neural network model in this embodiment.

[0063] Exemplary method

[0064] The method for building a neural network model in this embodiment can be applied to a terminal device, and the terminal device can be a terminal product with computing functions, such as a computer, etc. In this embodiment, as Figure 1 shown in

[0065] S100, obtain the target neural network model to be built, and the target neural network model is used to complete the target task.

[0066] The target neural network model that needs to be built can be a neural network model for performing face recognition tasks, or a neural network model for performing speech recognition tasks.

[0067] S200, according to the target neural network model, obtain each neural network layer corresponding to the target neural network model.

[0068] In step S100, it is known what kind of task the neural network model to be built can perform, and for what kind of task, there need to be corresponding types of neural network layers.

[0069] Step S200 includes the following steps S201, S202, and S203:

[0070] S201, according to the target neural network model, obtain the target component information corresponding to the target neural network model.

[0071] S202, according to the target component information, obtain the network layer connection information corresponding to the target component information.

[0072] S203. Connect the target component according to the network layer connection information to obtain the neural network layer corresponding to the target neural network model.

[0073] In this embodiment, according to the target neural network model to be built, it is possible to know how many neural network layers are required to build the target neural network model and the attributes of each neural network layer (target component information). For example, if it is known that an input layer, two convolutional layers, and an output layer are required, these layers are the target component information. After knowing the component information, search for the network layer connection information corresponding to these component information, that is, how the input layer is connected to the two convolutional layers, and how the convolutional layer is connected to the output layer. These network layer connection information are stored in the already constructed base class Moudle. The base class Moudle of this embodiment provides neural network layer component parameter collection (for collecting the weight information generated during the training process of the connected neural network layer later) and neural network layer collection function (network layer connection information).

[0074] S300. Obtain the weight information between adjacent neural network layers according to the output information in the adjacent neural network layers. An interface is provided on the neural network layer, and the interface is used to collect the output information and calculate the weight information according to the output information.

[0075] Step S300 includes three parts: The first part first initializes (initialized in init) the neural network model composed of each neural network layer, that is, assigns initial values to each weight parameter in the neural network model; The second part is to check whether each weight parameter has an initial value after initialization; The third part is to re-initialize the weight parameters that do not have initial values. The following introduces these three parts separately:

[0076] The first part includes initializing in init to obtain the initial values of the weight parameters corresponding to each neural network layer Layers corresponding to the target component information.

[0077] The second part includes inputting the data information corresponding to the target task into the neural network layer serving as the input layer to obtain the output result information of each neural network layer; According to the output result information, obtain whether there are corresponding initial weights (initial values of weight parameters) between each adjacent neural network layer. That is, perform the first forward propagation in the neural network model composed of each neural network layer to check whether each connected neural network layer has the corresponding weight parameter initial value.

[0078] The third part includes the following steps S301, S302, S303, S304, S305:

[0079] S301, based on the adjacent neural network layer, obtain the upper neural network layer and the lower neural network layer included in the adjacent neural network layer. The output of the upper neural network layer is the input of the lower neural network layer.

[0080] In a neural network model, if the output of one layer is used as the input of another layer, these two layers are adjacent neural network layers.

[0081] S302, based on the upper neural network layer, obtain the upper-layer output information corresponding to the upper neural network layer in the output information.

[0082] In this embodiment, the upper-layer output information is the number of neurons included in the upper neural network layer carried in the result output by the upper neural network layer. In this embodiment, the number of neurons included in the upper neural network layer is collected by setting an interface in each neural network layer (the interface layer provides a high-level API interface for users, and the API interface has the function of extracting the number of neurons, and the number of neurons is the shape size of the output of the upper neural network layer), and the upper-layer output information (the number of neurons) is saved in params.

[0083] S303, based on the lower neural network layer, obtain the lower-layer output information corresponding to the lower neural network layer in the output information.

[0084] The process of step S303 is similar to that of step S302.

[0085] S304, based on the number of upper-layer neurons and the number of lower-layer neurons, obtain the product result corresponding to the product of the number of upper-layer neurons and the number of lower-layer neurons.

[0086] S305, based on the product result, obtain the weight data size.

[0087] For example, if the product result is 9, that is, the weight data size is 9. For example, a matrix with 9 elements. At this time, the matrix is still an empty matrix, and subsequent steps are required to add values to this empty matrix to obtain the weight value corresponding to the weight parameter.

[0088] S400, based on the weight information, complete the construction of the target neural network model.

[0089] Step S400 includes the following steps S401, S402, S403, S404, S405:

[0090] S401, based on the weight data size, initialize the weight parameters.

[0091] In this embodiment, the weight parameters are initialized randomly.

[0092] S402. Train the neural network layer according to the weight parameters after the initialization, and collect the trained weight parameters during the training process.

[0093] In this embodiment, according to whether the result output by the neural network model where the neural network layer is located meets the requirements, the weight parameters are continuously adjusted. In this embodiment, the neural network layer component parameter collection program provided by the base class Moudle through the interface collects the changing weight parameters in real time during the training process, adjusts the weight parameters during the process of the neural network model outputting the result, and returns the adjusted weight parameters to the neural network model again for the next training. In this embodiment, the adjustment of the weight parameters and the output result of the neural network model are carried out in parallel and synchronously, which can improve the efficiency of building the neural network model.

[0094] S403. During the process of training the neural network layer, collect the connection relationship corresponding to the neural network layer, and the connection relationship corresponds to the trained weight parameters.

[0095] For example, there is a connection relationship S1 between neural network layer A and neural network layer B, and there is a connection relationship S2 between neural network layer C and neural network layer D. S1 and S2 are different. The weight parameters corresponding to neural network layer A and neural network layer B are E1, and the weight parameters corresponding to neural network layer C and neural network layer D are E2. E1 is corresponded with S1, and E2 is corresponded with S2. Then after optimizing E1 and E2 and returning the optimized E1 and E2 to the neural network layer, it is known where they should be placed between which two layers respectively.

[0096] S404. Send the connection relationship and the trained weight parameters to the optimizer, and the optimizer outputs the connection relationship and the optimized weight parameters corresponding to the connection relationship.

[0097] S405. Complete the construction of the target neural network model according to the connection relationship and the optimized weight parameters corresponding to the connection relationship output by the optimizer.

[0098] The overall process of building the neural network model of the present invention is as follows:

[0099] 1. Construct the base class Moudle, and the base class Moudle provides the collection of neural network layer component parameters and the collection function of the neural network layer.

[0100] 2. Obtain the target component information corresponding to the target neural network model to be constructed.

[0101] 3. Initialize each neural network layer Layers corresponding to the target component information in init according to the target component information.

[0102] 4. Define the forward propagation process in forward to connect each neural network layer to form a network.

[0103] 5. When the first forward propagation is executed, check whether each neural network layer component after connection has the corresponding initial value of the weight parameter.

[0104] 6. If there is no corresponding weight, initialize the weight layer by layer according to the forward propagation process.

[0105] 7. Form a network structure with neural network layers and training parameters with weights, thus completing the construction of the target neural network model.

[0106] The method for building a neural network model of the present invention relies on a hardware structure. The hardware structure is as Figure 2 shown. The hardware structure includes a kernel layer, an interface layer, and an algorithm layer. Hardware, backend, and language are all existing tools and devices. The hardware layer is the computing core for software operation, including CPUs, GPUs, TPUs, etc.; the backend layer is the tool required for software writing, including software such as TensorFlow and OpenMPI; the kernel layer is the core of the entire software. In this implementation, a Module class is designed as the base class. Based on this class, developers or users can develop high-level APIs and build deep learning models; the interface layer provides high-level APIs for users to build deep learning models. The APIs of the interface layer do not require the input of the shape and size of the upper layer; the language layer is the language for software writing; the algorithm layer provides deep learning models that can be directly called or pre-trained by users. The algorithm is a deep learning algorithm built based on the interface layer and the kernel.

[0107] Among them, the design of the Module class mainly focuses on adding training parameters and nesting Modules. The __setattr__ method is used to check all parameters in the neural network model. If they belong to weight parameters, the weight parameters will be added to _params (saved in dictionary format) through insert_param_to_layer. At the same time, to solve the problem of network layer nesting, Modules are also checked. If it is a Layer module, the Layer will be added to _layers (saved in dictionary format). The weights operation within the Module mainly consists of three parts. During model training, trainable_weights are used for gradient calculation to collect the parameters that need to be trained. nontrainable_weights returns the parameters that are not trained, and all_weights will return all parameters, which are obtained by traversing the Layers using the layers_and_name method. forward is used for users to customize the forward propagation after inheriting Module; build is used to implement the initialization of parameters and operator operations. After the model is built, set_train is used to train the model, and set_eval is used to make predictions on the model.

[0108] In this embodiment, when calculating the size of the weight data, it is necessary to infer the output size (the number of output neurons) of the previous neural network layer. The process of inferring the output size of the previous neural network layer is as follows:

[0109] 1. As Figure 3 shown, initialize the Layer, including the construction status (False) of the Layer, the status of the Layer, the parameter status of the Layer, and the training status of the Layer. This step is mainly to check whether each weight parameter has an initial value.

[0110] 2. Judge the construction status of the Layer. If it is not constructed, perform a forward propagation calculation on the previous layer to obtain the output size of the previous layer, and mark the construction status of the Layer as True.

[0111] 3. Use the output size as a parameter to be passed to the weight construction function to construct the training weights.

[0112] 4. As Figure 4 shown, collect the Layer layer and save the parameters.

[0113] In summary, the present invention provides an interface on the neural network layer. This interface can be used to collect weight information during the process of building a neural network model based on the neural network layer. It can also calculate the change information of the weights during the process of building the neural network model and return the change information of the weights to the neural network model being built. Therefore, the setting of the interface in the present invention enables the parallel and synchronous calculation of weight information and the building of the neural network model, thereby improving the speed of building the neural network model in the present invention.

[0114] Exemplary device

[0115] This embodiment also provides a device for a neural network model building method. The device includes the following components:

[0116] A target neural network model acquisition module, configured to acquire a target neural network model to be built, where the target neural network model is used to complete a target task;

[0117] A neural network layer generation module, configured to obtain each neural network layer corresponding to the target neural network model according to the target neural network model;

[0118] A weight calculation module, configured to obtain weight information between adjacent neural network layers according to output information in adjacent neural network layers;

[0119] A model building module, configured to complete the building of the target neural network model according to the weight information

[0120] Based on the above embodiments, the present invention also provides a terminal device, and its principle block diagram can be as Figure 5 shown. The terminal device includes a processor, a memory, a network interface, a display screen, and a temperature sensor connected through a system bus. Among them, the processor of the terminal device is used to provide computing and control capabilities. The memory of the terminal device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the terminal device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements a neural network model building method. The display screen of the terminal device can be a liquid crystal display screen or an electronic ink display screen, and the temperature sensor of the terminal device is pre-set inside the terminal device to detect the operating temperature of the internal device.

[0121] Those skilled in the art can understand, Figure 5The principle block diagram shown only shows the block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the terminal device to which the solution of the present invention is applied. The specific terminal device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0122] In one embodiment, a terminal device is provided. The terminal device includes a memory, a processor, and a neural network model building program stored in the memory and executable on the processor. When the processor executes the neural network model building program, the following operation instructions are implemented:

[0123] Obtain a target neural network model to be built, where the target neural network model is used to complete a target task;

[0124] According to the target neural network model, obtain each neural network layer corresponding to the target neural network model;

[0125] According to the output information in adjacent neural network layers, obtain the weight information between adjacent neural network layers. An interface is provided on the neural network layer, and the interface is used to collect the output information and calculate the weight information according to the output information;

[0126] Complete the building of the target neural network model according to the weight information.

[0127] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to the memory, storage, database, or other media used in the embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0128] In summary, the present invention discloses a method, apparatus, device, and storage medium for building a neural network model. The method includes: obtaining each neural network layer corresponding to the target neural network model according to the target neural network model; obtaining the weight information between adjacent neural network layers according to the output information in the adjacent neural network layers. An interface is provided on the neural network layer, and the interface is used to collect the output information and calculate the weight information according to the output information, thereby completing the building of the target neural network model. The setting of the interface in the present invention enables the calculation of the weight information and the building of the neural network model to be parallel and synchronous, thereby improving the speed of building the neural network model in the present invention.

[0129] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for building a neural network model, characterized in that, it includes: Obtain a target neural network model to be built, where the target neural network model is used to complete a target task, the target neural network model is used to perform a face recognition task, or the target neural network model is used to perform a speech recognition task; According to the target neural network model, obtain each neural network layer corresponding to the target neural network model; According to the output information in adjacent neural network layers, obtain the weight information between adjacent neural network layers. An interface is provided on the neural network layer, and the interface is used to collect the output information and calculate the weight information according to the output information. The interface calculates the change information of the weight during the process of building the neural network model and returns the weight change information to the neural network model being built; Complete the building of the target neural network model according to the weight information; The step of completing the building of the target neural network model according to the weight information includes: According to the weight information, obtain the weight data size in the weight information; Initialize the weight parameters according to the weight data size; Train the neural network layer according to the weight parameters after initialization, and collect the trained weight parameters through the interface during the training process; Send the trained weight parameters to an optimizer to obtain the optimized weight parameters and complete the building of the target neural network model.

2. The method for building a neural network model according to claim 1, characterized in that, The step of obtaining the weight information between adjacent neural network layers according to the output information in adjacent neural network layers, where an interface is provided on the neural network layer, and the interface is used to collect the output information and calculate the weight information according to the output information, includes: According to adjacent neural network layers, obtain the upper neural network layer and the lower neural network layer included in the adjacent neural network layers. The output of the upper neural network layer is the input of the lower neural network layer; According to the upper neural network layer, obtain the upper output information corresponding to the upper neural network layer in the output information; According to the lower neural network layer, obtain the lower output information corresponding to the lower neural network layer in the output information; Obtain the weight information between adjacent neural network layers according to the upper output information and the lower output information.

3. The method for building a neural network model according to claim 2, characterized in that, The step of obtaining the weight information between adjacent neural network layers according to the upper output information and the lower output information includes: According to the upper output information, obtain the number of upper neurons in the upper output information; According to the lower output information, obtain the number of lower neurons in the lower output information; Obtain the weight data size in the weight information according to the number of upper neurons and the number of lower neurons.

4. The method for building a neural network model according to claim 3, wherein, obtaining the weight data size in the weight information according to the number of neurons in the upper layer and the number of neurons in the lower layer includes: obtaining the product result corresponding to the product of the number of neurons in the upper layer and the number of neurons in the lower layer according to the number of neurons in the upper layer and the number of neurons in the lower layer; obtaining the weight data size according to the product result.

5. The method for building a neural network model according to claim 1, wherein, obtaining the weight information between adjacent neural network layers according to the output information in adjacent neural network layers, an interface is provided on the neural network layer, and the interface is used to collect the output information and calculate the weight information according to the output information, including: judging whether there are corresponding initial weights between each adjacent neural network layer; when there is no corresponding initial weight between adjacent neural network layers, inputting the data information corresponding to the target task into the neural network layer serving as the input layer to obtain the output information in adjacent neural network layers; obtaining the weight information between adjacent neural network layers according to the output information in adjacent neural network layers.

6. The method for building a neural network model according to claim 5, wherein, judging whether there are corresponding initial weights between each adjacent neural network layer includes: inputting the data information corresponding to the target task into the neural network layer serving as the input layer to obtain the output result information of each neural network layer; obtaining the judgment result of whether there are corresponding initial weights between each adjacent neural network layer according to the output result information.

7. The method for building a neural network model according to claim 1, wherein, obtaining the neural network layer corresponding to the target neural network model according to the target neural network model includes: obtaining the target component information corresponding to the target neural network model according to the target neural network model; obtaining the network layer connection information corresponding to the target component information according to the target component information; connecting the target components according to the network layer connection information to obtain the neural network layer corresponding to the target neural network model.

8. The method for building a neural network model according to claim 1, wherein, sending the weight parameters to be trained to an optimizer to obtain the optimized weight parameters, and completing the building of the target neural network model, including: during the process of training the neural network layer, collecting the connection relationship corresponding to the weight parameters to be trained of the neural network layer; sending the connection relationship and the weight parameters to be trained to the optimizer, and the optimizer outputs the connection relationship and the optimized weight parameters corresponding to the connection relationship; completing the building of the target neural network model according to the connection relationship output by the optimizer and the optimized weight parameters corresponding to the connection relationship.

9. An apparatus for a method of building a neural network model, characterized in that, the apparatus comprises the following components: A target neural network model acquisition module, configured to acquire a target neural network model to be built, the target neural network model being used to complete a target task, the target neural network model being used to perform a face recognition task, or the target neural network model being used to perform a speech recognition task; A neural network layer generation module, configured to obtain each neural network layer corresponding to the target neural network model according to the target neural network model; A weight calculation module, configured to obtain weight information between adjacent neural network layers according to output information in adjacent neural network layers, the interface calculates change information of weights during the process of building a neural network model, and returns the weight change information to the neural network model being built; A model building module, configured to complete the building of the target neural network model according to the weight information; The completing the building of the target neural network model according to the weight information includes: Obtaining the weight data size in the weight information according to the weight information; Initializing the weight parameters according to the weight data size; Training the neural network layer according to the weight parameters after the initialization, and collecting the trained weight parameters through the interface during the training process; Sending the trained weight parameters to an optimizer to obtain the optimized weight parameters, and completing the building of the target neural network model.

10. A terminal device, characterized in that, the terminal device includes a memory, a processor, and a neural network model building program stored in the memory and executable on the processor. When the processor executes the neural network model building program, the steps of the neural network model building method according to any one of claims 1-8 are implemented.

11. A computer-readable storage medium, characterized in that, a neural network model building program is stored on the computer-readable storage medium. When the neural network model building program is executed by a processor, the steps of the neural network model building method according to any one of claims 1-8 are implemented.