XML-based neural network model generation method, device and storage medium

Through the XML-based neural network model generation method, the deep learning framework attributes in XML instance text are parsed and the target neural network model is generated, which solves the problem of poor compatibility between deep learning frameworks in the existing technology, and realizes efficient model generation and migration.

CN114492321BActive Publication Date: 2025-05-06INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111673232.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-05-06
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

There is poor compatibility between existing deep learning frameworks, which leads to low generation efficiency and difficulty in migration of neural network models, which in turn increases the workload of model generation.

Method used

Using the XML-based neural network model generation method, the target neural network model is generated by obtaining and parsing the deep learning framework attributes and attribute values ​​in the XML instance text. This method includes initializing the neural network model, parsing child elements in XML text, updating the model until all child elements are parsed.

Benefits of technology

Reduces the generation workload of neural network models, improves model compatibility and migration efficiency, and eliminates the need to design and debug unique models for each deep learning framework.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114492321B_ABST
    Figure CN114492321B_ABST
Patent Text Reader

Abstract

The present invention provides a method, device and storage medium for generating a neural network model based on XML, wherein the method comprises: obtaining an XML instance text to be parsed, wherein the XML instance text defines deep learning framework attributes and their deep learning framework attribute values; parsing the XML instance text to generate a target neural network model corresponding to the deep learning framework attribute values. The method, device and storage medium provided by the present invention only need to modify the deep learning framework attribute values ​​for different deep learning frameworks. The XML instance text is compatible with multiple deep learning frameworks, and the user does not need to design, implement and debug his own unique model for each deep learning framework, thereby greatly reducing the workload of generating the neural network model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an XML-based neural network model generation method, device and storage medium. Background Art

[0002] With the rapid development of deep learning technology, deep learning frameworks such as TensorFlow, Pytorch, and PaddlePaddle continue to emerge and develop. However, the design patterns, features, usage, and basic support libraries of these deep learning frameworks are mostly different.

[0003] The differences between different deep learning frameworks require researchers and engineers to invest a lot of time in learning the various details of different deep learning frameworks in order to generate neural network models corresponding to various deep learning frameworks, resulting in low model generation efficiency; it also brings about the problem of difficulty in model migration. Specifically, the models and algorithms developed by different deep learning frameworks are not compatible and migrated with each other, which seriously restricts the migration efficiency and sharing capabilities of models and algorithms; moreover, due to the huge differences between different deep learning frameworks, the performance of the same type of models and algorithms is inconsistent in different deep learning frameworks, which seriously affects the fair evaluation in scientific monitoring, industrial production and other fields.

[0004] In summary, the neural network models generated based on the current mainstream deep learning frameworks have poor compatibility with each other, which makes model migration difficult. Users are required to design, implement and debug their own unique models for each deep learning framework, which greatly increases the workload of generating neural network models. Summary of the invention

[0005] The present invention provides an XML-based neural network model generation method, device and storage medium, which are used to solve the defect of heavy workload in the generation of neural network models in the prior art.

[0006] The present invention provides a method for generating a neural network model based on XML, comprising:

[0007] Obtain an XML instance text to be parsed, wherein the XML instance text defines a deep learning framework attribute and its deep learning framework attribute value;

[0008] The XML instance text is parsed to generate a target neural network model corresponding to the deep learning framework attribute value.

[0009] According to an XML-based neural network model generation method provided by the present invention, the XML instance text is parsed to generate a target neural network model corresponding to the deep learning framework attribute value, including:

[0010] Initializing based on the attributes of the root element in the XML instance text to obtain an initial neural network model;

[0011] Determine a first target sub-element to be parsed in the root element, and call a corresponding parsing algorithm based on the type of the first target sub-element to parse the first target sub-element to update the initial neural network model;

[0012] Return to the step of determining the first target sub-element to be parsed in the root element until all sub-elements in the root element are parsed, and the last updated initial neural network model is determined as the target neural network model.

[0013] According to an XML-based neural network model generation method provided by the present invention, the root element includes a hyperparameter sub-element, and the execution steps of the parsing algorithm corresponding to the hyperparameter sub-element are as follows:

[0014] The attribute value of the hyperparameter attribute in the hyperparameter sub-element is read, and the initial neural network model is updated based on the read attribute value of the hyperparameter attribute.

[0015] According to an XML-based neural network model generation method provided by the present invention, the root element includes a network layer sub-element, and the execution steps of the parsing algorithm corresponding to the network layer sub-element are as follows:

[0016] Reading the attribute value of the network layer attribute in the network layer sub-element, and constructing the target neural network layer to be added based on the attribute value of the read network layer attribute;

[0017] Adding the target neural network layer to the initial neural network model;

[0018] The root element also includes an element-by-element operation layer sub-element, and the execution steps of the parsing algorithm corresponding to the element-by-element operation layer sub-element are as follows:

[0019] Read the attribute value of the element-by-element operation layer attribute in the element-by-element operation layer sub-element, and construct the target element-by-element operation layer to be added based on the read attribute value of the element-by-element operation layer attribute;

[0020] The target element-wise operation layer is added to the initial neural network model.

[0021] According to an XML-based neural network model generation method provided by the present invention, the root element includes a block sub-element, and the execution steps of the parsing algorithm corresponding to the block sub-element are as follows:

[0022] Based on the attributes of the block sub-element, initialization is performed to obtain an initial block;

[0023] Determine a second target sub-element to be parsed currently in the block sub-element, and call a corresponding parsing algorithm based on the type of the second target sub-element to parse the second target sub-element, so as to update the initial block;

[0024] Return to the step of determining the second target sub-element currently to be parsed in the block sub-element until all sub-elements in the block sub-element are parsed, and add the last updated initial block to the initial neural network model.

[0025] According to an XML-based neural network model generation method provided by the present invention, the deep learning framework attribute is defined in the root element;

[0026] The initialization based on the attributes of the root element in the XML instance text to obtain an initial neural network model includes:

[0027] Based on the deep learning framework attribute value of the root element, initialization is performed to obtain an initial neural network model corresponding to the deep learning framework attribute value.

[0028] According to an XML-based neural network model generation method provided by the present invention, the XML instance text includes a root element, the root element includes a hyperparameter sub-element, a network layer sub-element, an element-by-element operation layer sub-element, a block sub-element, and a data set sub-element, and the block sub-element includes a network layer sub-element and an element-by-element operation layer sub-element.

[0029] According to an XML-based neural network model generation method provided by the present invention, the step of obtaining an XML instance text to be parsed includes:

[0030] Obtaining an input XML text, and verifying the XML text based on a preset standardization definition specification;

[0031] If the XML text complies with the standardization definition specification, determining the XML text as the XML instance text;

[0032] If the XML text does not conform to the standard definition specification, an error prompt message is issued, and the error prompt message is used to prompt the XML text to be modified;

[0033] Return to the step of obtaining the input XML text until the modified XML text complies with the standardization definition specification.

[0034] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of any of the above-mentioned XML-based neural network model generation methods are implemented.

[0035] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the XML-based neural network model generation methods described above.

[0036] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the steps of any of the above-mentioned XML-based neural network model generation methods are implemented.

[0037] The XML-based neural network model generation method, device and storage medium provided by the present invention obtain an XML instance text to be parsed, wherein the XML instance text defines deep learning framework attributes and their deep learning framework attribute values; the XML instance text is parsed to generate a target neural network model corresponding to the deep learning framework attribute values. In the above manner, the deep learning framework attributes and their deep learning framework attribute values ​​are defined in the XML instance text, so that when the XML instance text is subsequently parsed, a neural network model corresponding to the defined deep learning framework can be generated. Based on this, for different deep learning frameworks, it is only necessary to modify the deep learning framework attribute values. The XML instance text is compatible with multiple deep learning frameworks, and there is no need for users to design, implement and debug their own unique models for each deep learning framework, thereby greatly reducing the workload of generating the neural network model. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0039] Figure 1 One of the flow charts of the XML-based neural network model generation method provided by the present invention;

[0040] Figure 2 The second flowchart of the XML-based neural network model generation method provided by the present invention;

[0041] Figure 3 This is a schematic structural diagram of an electronic device provided by the present invention. DETAILED DESCRIPTION

[0042] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0043] With the rapid development of deep learning technology, deep learning frameworks such as TensorFlow, Pytorch, and PaddlePaddle continue to emerge and develop. However, the design patterns, features, usage, and basic support libraries of these deep learning frameworks are mostly different. For example, early versions of TensorFlow built neural networks through static graphs. This static computational graph mode cannot be modified at runtime. It usually has a faster running speed, but the program debugging is more complicated and less user-friendly. In Pytorch, users can define and run a dynamic computational graph at the same time, which is more user-friendly in programming and debugging, but its dynamic characteristics also lead to a slight disadvantage in the running speed of the model.

[0044] The differences between different deep learning frameworks require researchers and engineers to invest a lot of time in learning the various details of different deep learning frameworks in order to generate neural network models corresponding to various deep learning frameworks, resulting in low model generation efficiency; it also brings about the problem of difficulty in model migration. Specifically, the models and algorithms developed by different deep learning frameworks are not compatible and migrated with each other, which seriously restricts the migration efficiency and sharing capabilities of models and algorithms; moreover, due to the huge differences between different deep learning frameworks, the performance of the same type of models and algorithms is inconsistent in different deep learning frameworks, which seriously affects the fair evaluation in scientific monitoring, industrial production and other fields.

[0045] In summary, the neural network models generated based on the current mainstream deep learning frameworks have poor compatibility with each other, which makes model migration difficult. Users are required to design, implement and debug their own unique models for each deep learning framework, which greatly increases the workload of generating neural network models and reduces the generation efficiency of neural network models.

[0046] Based on the above problems, the present invention provides a method for generating a neural network model based on XML. Figure 1 One of the flow charts of the XML-based neural network model generation method provided by the present invention is as follows: Figure 1 As shown, the method includes:

[0047] Step 110, obtaining an XML instance text to be parsed, wherein the XML instance text defines deep learning framework attributes and their deep learning framework attribute values.

[0048] Here, the XML instance text is the text to be parsed, that is, the XML instance text is the text to be decoded. Specifically, the XML instance text is parsed and the XML instance text can be converted into a neural network model.

[0049] The XML instance text can be written by the user, and further, the user can write the corresponding XML text according to the standardization definition specification and his / her own needs. The standardization definition specification is used to assist and constrain the user to write the corresponding XML text.

[0050] Here, the deep learning framework attribute is used to specify which deep learning framework to embed. For example, if the deep learning framework attribute is platform, and platform = "Pytorch", that is, the deep learning framework attribute value is Pytorch, it is specified to embed in the Pytorch deep learning framework.

[0051] In a specific embodiment, the deep learning framework attributes and their deep learning framework attribute values ​​can be defined in the root element of the XML instance text. For example, the root element of the XML instance text is <model>, the deep learning framework attribute is platform, and the deep learning framework attribute value is Pytorch. At this time, the root element is<model platform="Pytorch”> .

[0052] Specifically, the XML instance text includes information such as network structure and network parameters. More specifically, the neural network components supported by the XML instance text include one or more of the following: convolutional neural network, pooling layer, padding layer, batch normalization layer, dropout layer, upsampling layer, residual connection, various element-wise operations, various tensor operations (such as flatten, unsqueeze, reshape, etc.) and various mathematical calculation operations, etc., as well as attention module, activation function module, loss function module, parameter initialization module, optimizer module, evaluation index module, etc.

[0053] Among them, the convolutional neural network can be a 2-dimensional convolutional neural network, a 3-dimensional convolutional neural network, or a convolutional neural network of other dimensions, which is not limited here.

[0054] Attention module, the attention mechanisms provided include but are not limited to: additive attention mechanism, point product attention mechanism and fully connected attention mechanism.

[0055] The activation function module includes all activation function layers supported by various deep learning frameworks, for example, all activation function layers supported by TensorFlow or Pytorch, and the activation function layers may be ReLU activation function layers, Sigmoid activation function layers, Softmax activation function layers, etc.

[0056] The parameter initialization module supports all parameter initialization methods built into various deep learning frameworks, for example, it supports all parameter initialization methods built into TensorFlow or Pytorch, including random uniform initialization, random normal initialization, orthogonal initialization, Xavier initialization, etc. In addition, based on this parameter initialization module, fine-grained initialization can be provided. For example, users can specify the mean and variance of random normal initialization based on this module.

[0057] The optimizer module provides all optimizers supported by various deep learning frameworks, for example, all optimizers supported by TensorFlow or Pytorch, including the stochastic gradient descent (SGD) optimizer, the momentum optimizer, the Adam optimizer, etc. In addition, based on the optimizer module, fine-grained control of the learning rate can be provided. For example, users can specify the initial learning rate, the minimum learning rate, the learning rate decay rate, etc. based on this module.

[0058] The evaluation index module supports accuracy, precision, recall, F1 value and other indicators for classification tasks; mAP evaluation indicators for object detection tasks; pixel accuracy, IoU and other evaluation indicators for image segmentation tasks. Of course, there can be more or fewer evaluation indicators for other tasks, which are not limited here.

[0059] In a specific embodiment, each of the above-mentioned neural network components needs to meet the standardized definition specifications and be defined at the corresponding position of the XML instance text.

[0060] A convolutional neural network can be defined through a network layer sub-element, in which the output size of the convolutional neural network, the name of the convolutional neural network, the convolution kernel size of the convolutional neural network, the rate of the convolutional neural network, whether it can be reused, etc. can be defined.

[0061] For example, the network layer sub-element is <layer>, set the attribute value of the type attribute in the network layer sub-element to Convolution, the attribute of the output size of the convolutional neural network to out, the attribute of the name of the convolutional neural network to name, the attribute of the convolutional neural network convolution kernel size to kernel, and the attribute of whether it can be reused to reuse. At this time, the following definitions can be made:

[0062] <layer type="Convolution”out="64”name="conv1_1” / > ;

[0063] <layer type="Convolution”out="1024”kernel="1,1”name="conv7” / > ;

[0064] <layer type="Convolution”out="1024”rate="6”name="conv6” / > ;

[0065] <layer type="Convolution”out="[0]”reuse="True”name="cls_pred” / > .

[0066] Here, the network layer sub-element corresponding to the convolutional neural network can be a sub-element of the root element in the XML instance text, or a sub-element of the block sub-element of the root element in the XML instance text.

[0067] The pooling layer can be defined through the network layer sub-element, in which the pooling kernel size, the name of the pooling layer, the padding mode of the pooling layer, the pooling algorithm (maximum pooling or average pooling), the step size of the pooling layer, etc. can be defined.

[0068] For example, the network layer sub-element is <layer>, set the attribute value of the type attribute in the network layer sub-element to Subsampling, the attribute of the pooling kernel size to kernel, the attribute of the pooling layer name to name, the attribute of the pooling layer padding mode to mode, the attribute of the pooling algorithm to algorithm, and the attribute of the pooling layer step size to stride. At this time, the following definitions can be made:

[0069] <layer type="Subsampling”kernel="2,2”stride="2,2”mode="SAME”algorithm="MAX”name="pool2” / > ;

[0070] <layer type="Subsampling”kernel="3,3”mode="SAME”algorithm="MAX”name="pool5” / > .

[0071] Here, the network layer sub-element corresponding to the pooling layer may be a sub-element of the root element in the XML instance text, or a sub-element of the block sub-element of the root element in the XML instance text.

[0072] The Dropout layer can be defined through the network layer sub-element, in which the dropout probability of the Dropout layer, the Dropout layer, etc. can be defined.

[0073] For example, the network layer sub-element is <layer>, set the attribute value of the type attribute in the network layer sub-element to Dropout, the attribute of the dropout probability of the Dropout layer to rate, and the attribute of the name of the Dropout layer to name. At this time, the following definition can be given:

[0074] <layer type="Dropout”rete="0.1” / > ;

[0075] <layer type="Dropout”rete="0.1”name="dropout2” / > .

[0076] Here, the network layer sub-element corresponding to the Dropout layer may be a sub-element of the root element in the XML instance text, or a sub-element of the block sub-element of the root element in the XML instance text.

[0077] The optimizer module can be defined in the hyperparameter sub-element, that is, the optimizer sub-element is defined in the hyperparameter sub-element, and the optimizer sub-element can be <updater>, in the optimizer sub-element, you can define the optimizer type, and the attribute of the optimizer type can be type. For example, the hyperparameter sub-element is <parameters>, at this time, the following definition can be given:

[0078] <parameters>

[0079] <updater type="Adam” / >

[0080] < / parameters> .

[0081] In addition, more network structures or more network parameters can be defined in the XML instance text, which will not be described here one by one.

[0082] Step 120, parse the XML instance text to generate a target neural network model corresponding to the deep learning framework attribute value.

[0083] Specifically, the XML instance text is parsed based on a preset parsing algorithm, and the XML instance text is converted into a computational graph instance, and then further converted into a target neural network model corresponding to the attribute value of the deep learning framework. The execution process of the preset parsing algorithm refers to the following embodiments, which will not be repeated here.

[0084] It can be understood that by parsing the deep learning framework attribute value in the XML instance text, the deep learning framework to be embedded can be determined, and then the target neural network model corresponding to the deep learning framework attribute value can be generated.

[0085] Here, the target neural network model is a trainable neural network model. For example, if the deep learning framework attribute value is Pytorch, a trainable Pytorch neural network model is generated.

[0086] The XML-based neural network model generation method provided by the embodiment of the present invention obtains an XML instance text to be parsed, wherein the XML instance text defines deep learning framework attributes and their deep learning framework attribute values; the XML instance text is parsed to generate a target neural network model corresponding to the deep learning framework attribute values. In the above manner, the deep learning framework attributes and their deep learning framework attribute values ​​are defined in the XML instance text, so that when the XML instance text is subsequently parsed, a neural network model corresponding to the defined deep learning framework can be generated. Based on this, for different deep learning frameworks, it is only necessary to modify the deep learning framework attribute values. The XML instance text is compatible with multiple deep learning frameworks, and the user does not need to design, implement and debug their own unique model for each deep learning framework, thereby greatly reducing the workload of generating the neural network model.

[0087] Based on the above embodiments, Figure 2 The second flow chart of the XML-based neural network model generation method provided by the present invention is as follows: Figure 2 As shown, in this method, the above step 120 includes:

[0088] Step 121, based on the attributes of the root element in the XML instance text, initialization is performed to obtain an initial neural network model.

[0089] Specifically, when parsing an XML instance text, the attributes of the root element in the XML instance text are parsed first, and then an initial neural network model is initialized based on the parsing results. When subsequently parsing the child elements of the root element, the initial neural network model is continuously updated until all the child elements of the root element are parsed, thereby taking the last updated initial neural network model as the final generated target neural network model.

[0090] Here, the attributes of the root element may include but are not limited to one or more of the following: a deep learning framework attribute platform, a backpropagation attribute backpropagate, a model category attribute type, a pretrain attribute pretrain, and the like.

[0091] The platform attribute is used to specify which deep learning framework to embed, that is, the platform attribute is a deep learning framework attribute. The backpropagate attribute is used to specify whether to backpropagate. The type attribute is used to specify the type of the generated neural network model. The pretrain attribute is used to specify whether to perform pre-training.

[0092] For example, the root element is <model>, at this point, the root element can be defined as follows:

[0093] <model type="SSD”pretrain="false”backpropagate="true”platform="Pytorch”> .

[0094] In a specific embodiment, the deep learning framework attribute is defined in the root element, and the step 121 includes:

[0095] Based on the deep learning framework attribute value of the root element, initialization is performed to obtain an initial neural network model corresponding to the deep learning framework attribute value.

[0096] For example, the root element of the XML instance text is <model>, the deep learning framework attribute is platform, and the deep learning framework attribute value is Pytorch. At this time, the root element is<model platform="Pytorch”> Based on this, when parsing the attributes of the root element in the XML instance text, the initial neural network model corresponding to Pytorch can be initialized based on the parsing results. The target neural network model obtained by continuously updating the initial neural network model is also the neural network model corresponding to Pytorch.

[0097] Step 122, determining the first target sub-element to be parsed in the root element, and calling a corresponding parsing algorithm based on the type of the first target sub-element to parse the first target sub-element to update the initial neural network model.

[0098] Specifically, the root element may include, but is not limited to, one or more of the following: a hyperparameter sub-element, a network layer sub-element, an element-by-element operation layer sub-element, a block sub-element, a data set sub-element, and the like.

[0099] The hyperparameters sub-element is used to set hyperparameters. For example, the hyperparameters sub-element is <parameters>, which can be defined as

[0100] <parameters>

[0101] <updater type="Adam” / >

[0102] < / parameters> .

[0103] The network layer sub-element is used to build the main components of the neural network model, which can be a recurrent neural network (RNN), convolutional neural network (CNN), fully connected layer (Dense), Dropout layer, Transformer and other neural network layers. For example, the network layer sub-element is <layer>, which can be defined as

[0104] <layer type="Convolution”out="64”name="conv1_1” / > .

[0105] The element-by-element operation layer sub-element is used to represent operators that operate on two or more input data, such as concatenation (CONCAT), element-by-element addition (ELEMENT_ADD), element-by-element multiplication (ELEMENT_MUL), element-by-element maximum (ELEMENT_MAX), element-by-element average (ELEMENT_MEAN), etc. For example, the element-by-element operation layer sub-element is <vertex>.

[0106] The block sub-elements include network layer sub-elements, element-by-element operation layer sub-elements, etc. The block sub-elements can define the block type, block name, block activation function, block output size, block core size, etc.

[0107] For example, a block child element is <block>, the attribute of the block type is type, the attribute of the block name is name, the attribute of the block activation function is activation, the attribute of the block output size is out, and the attribute of the block kernel size is kernel. At this time, the following definitions can be made:

[0108] <block type="VGG16”name="vgg”kernel="3,3”activation="relu”>< / block> ;

[0109] <block kernel="3,3”out="84,16”parent="vgg / conv4_3”name="output1”> .

[0110] The dataset sub-element is used to specify the dataset and set related parameters. For example, the dataset sub-element is <dataset>.

[0111] In addition, the dataset sub-element is usually the first sub-element of the root element, and the hyperparameter sub-element is usually the second sub-element of the root element.

[0112] Specifically, if the first target sub-element is a hyperparameter sub-element, the hyperparameter parsing algorithm is called to parse the first target sub-element to update the initial neural network model; if the first target sub-element is a network layer sub-element, the network layer parsing algorithm is called to parse the first target sub-element to update the initial neural network model; if the first target sub-element is an element-by-element operation layer sub-element, the element-by-element operation layer parsing algorithm is called to parse the first target sub-element to update the initial neural network model; if the first target sub-element is a block sub-element, the block parsing algorithm is called to parse the first target sub-element to update the initial neural network model; if the first target sub-element is a data set sub-element, the data set parsing algorithm is called to parse the first target sub-element to update the initial neural network model. For each parsing algorithm, refer to the following embodiments and will not be repeated here.

[0113] Step 123, returns to the step of determining the first target sub-element to be parsed in the root element, until all sub-elements in the root element are parsed, and the last updated initial neural network model is determined as the target neural network model.

[0114] Specifically, each sub-element in the root element is parsed, and the initial neural network model is updated while parsing each sub-element. It can be understood that parsing all sub-elements is to update the same initial neural network model, that is, the update effect of each sub-element is superimposed.

[0115] According to the XML-based neural network model generation method provided by the embodiment of the present invention, the initial neural network model is initialized based on the attributes of the root element in the XML instance text; the first target sub-element to be parsed in the root element is determined, and the corresponding parsing algorithm is called based on the type of the first target sub-element to parse the first target sub-element to update the initial neural network model; the step of determining the first target sub-element to be parsed in the root element is returned until all sub-elements in the root element are parsed, and the last updated initial neural network model is determined as the target neural network model. In the above manner, regardless of the deep learning framework, the embodiment of the present invention parses the root element and its sub-elements of the XML instance text, and there is no need to design a corresponding parsing algorithm for each deep learning framework, so that the algorithm migration can be achieved, thereby further reducing the workload of generating the neural network model.

[0116] Based on any of the above embodiments, in this method, the root element includes a super parameter sub-element, and the execution steps of the parsing algorithm corresponding to the super parameter sub-element are as follows:

[0117] The attribute value of the hyperparameter attribute in the hyperparameter sub-element is read, and the initial neural network model is updated based on the read attribute value of the hyperparameter attribute.

[0118] Here, the hyperparameter sub-element is used to set the hyperparameter. The attributes of the hyperparameter sub-element include but are not limited to: seed number, optimization algorithm optimization_algorithm, learning rate learning_rate, training round epochs, batch size batch_size, weight type weight and optimizer updater, etc.

[0119] For example, the hyperparameter sub-element is <parameters>, which can be defined as

[0120]

[0121]

[0122] Specifically, the attribute value of the hyperparameter attribute in the hyperparameter sub-element is read, and the attribute value of the hyperparameter attribute is parsed, and the initial neural network model is updated based on the parsing result.

[0123] The method provided by the embodiment of the present invention provides an algorithm for parsing super-parameter sub-elements to provide support for the parsing algorithm of XML instance text.

[0124] Based on any of the above embodiments, in the method, the root element includes a network layer sub-element, and the execution steps of the parsing algorithm corresponding to the network layer sub-element are as follows:

[0125] Reading the attribute value of the network layer attribute in the network layer sub-element, and constructing the target neural network layer to be added based on the attribute value of the read network layer attribute;

[0126] The target neural network layer is added to the initial neural network model.

[0127] Specifically, the type attribute value in the network layer sub-element is read to determine the type of the neural network layer to be added, and then based on the type of the neural network layer, the attribute values ​​of other network layer attributes in the network layer sub-element are read to construct the target neural network layer to be added based on the read attribute values.

[0128] The root element also includes an element-by-element operation layer sub-element, and the execution steps of the parsing algorithm corresponding to the element-by-element operation layer sub-element are as follows:

[0129] Read the attribute value of the element-by-element operation layer attribute in the element-by-element operation layer sub-element, and construct the target element-by-element operation layer to be added based on the read attribute value of the element-by-element operation layer attribute;

[0130] The target element-wise operation layer is added to the initial neural network model.

[0131] Specifically, the type attribute value in the element-by-element operation layer sub-element is read to determine the type of element-by-element operation to be added, and then based on the type of element-by-element operation, the attribute values ​​of other element-by-element operation layer attributes in the element-by-element operation layer sub-element are read to construct the element-by-element operation layer to be added based on the read attribute values.

[0132] The method provided by the embodiment of the present invention provides an algorithm for parsing network layer sub-elements and element-by-element operation layer sub-elements, so as to provide support for the parsing algorithm of XML instance text.

[0133] Based on any of the above embodiments, in this method, the root element includes a block sub-element, and the execution steps of the parsing algorithm corresponding to the block sub-element are as follows:

[0134] Based on the attributes of the block sub-element, initialization is performed to obtain an initial block;

[0135] Determine a second target sub-element to be parsed currently in the block sub-element, and call a corresponding parsing algorithm based on the type of the second target sub-element to parse the second target sub-element, so as to update the initial block;

[0136] Return to the step of determining the second target sub-element currently to be parsed in the block sub-element until all sub-elements in the block sub-element are parsed, and add the last updated initial block to the initial neural network model.

[0137] Specifically, the type attribute value in the block sub-element is read to determine the type of the block to be added, and then based on the type of the block, the attribute values ​​of other block attributes in the block sub-element are read to initialize the initial block based on the read attribute values.

[0138] Afterwards, if the second target sub-element is a network layer sub-element, the network layer parsing algorithm is called to parse the second target sub-element to update the initial block; if the second target sub-element is an element-by-element operation layer sub-element, the element-by-element operation layer parsing algorithm is called to parse the second target sub-element to update the initial block. For each parsing algorithm, refer to the above embodiment and will not be repeated here.

[0139] The method provided by the embodiment of the present invention provides an algorithm for parsing block sub-elements to provide support for the parsing algorithm of XML instance text.

[0140] Based on any of the above embodiments, the XML instance text includes a root element, and the root element includes a hyperparameter sub-element, a network layer sub-element, an element-by-element operation layer sub-element, a block sub-element, and a data set sub-element, and the block sub-element includes a network layer sub-element and an element-by-element operation layer sub-element.

[0141] Based on any of the above embodiments, the above step 110 includes:

[0142] Obtaining an input XML text, and verifying the XML text based on a preset standardization definition specification;

[0143] If the XML text complies with the standardization definition specification, determining the XML text as the XML instance text;

[0144] If the XML text does not conform to the standard definition specification, an error prompt message is issued, and the error prompt message is used to prompt the XML text to be modified;

[0145] Return to the step of obtaining the input XML text until the modified XML text complies with the standardization definition specification.

[0146] Here, the standardized definition specifications are used to assist and constrain users in writing corresponding XML texts, that is, to assist and constrain users in writing corresponding neural network structures. Further, users write corresponding XML texts based on the standardized definition specifications and their own needs for neural network structures.

[0147] The standardization definition specification can be used to verify the input XML text. Specifically, it is checked according to the standardization definition specification whether the XML text meets the encoding requirements and whether it complies with the relevant restrictions defined in the specification.

[0148] The standardized definition specification includes, but is not limited to, one or more of the following: network organizational structure specification, included element specification, element length specification, element type specification, etc., which is not specifically limited in the embodiment of the present invention.

[0149] It is understandable that by judging whether the input XML text has problems according to the standardized definition specifications, the writing of XML text becomes rule-based and there will not be a myriad of XML texts that change constantly.

[0150] Specifically, the standardization definition specifications include the general information exchange protocol specifications corresponding to the text organization structure of the XML text, and also include some restrictions on elements formulated for neural networks.

[0151] In a specific embodiment, the standardized definition specification is as follows:

[0152] The root element may include, but is not limited to, one or more of the following: hyperparameter sub-element, network layer sub-element, element-by-element operation layer sub-element, block sub-element, dataset sub-element, etc.

[0153] The hyperparameters sub-element is used to set hyperparameters. For example, the hyperparameters sub-element is <parameters>, which can be defined as

[0154] <parameters>

[0155] <updater type="Adam” / >

[0156] < / parameters> .

[0157] The network layer sub-element is used to build the main components of the neural network model, which can be a recurrent neural network (RNN), convolutional neural network (CNN), fully connected layer (Dense), Dropout layer, Transformer and other neural network layers. For example, the network layer sub-element is <layer>, which can be defined as

[0158] <layer type="Convolution”out="64”name="conv1_1” / > .

[0159] The element-by-element operation layer sub-element is used to represent operators that operate on two or more input data, such as concatenation (CONCAT), element-by-element addition (ELEMENT_ADD), element-by-element multiplication (ELEMENT_MUL), element-by-element maximum (ELEMENT_MAX), element-by-element average (ELEMENT_MEAN), etc. For example, the element-by-element operation layer sub-element is <vertex>.

[0160] The block sub-elements include network layer sub-elements, element-by-element operation layer sub-elements, etc. The block sub-elements can define the block type, block name, block activation function, block output size, block core size, etc.

[0161] For example, a block child element is <block>, the attribute of the block type is type, the attribute of the block name is name, the attribute of the block activation function is activation, the attribute of the block output size is out, and the attribute of the block kernel size is kernel. At this time, the following definitions can be made:

[0162] <block type="VGG16”name="vgg”kernel="3,3”activation="relu”>< / block> ;

[0163] <block kernel="3,3”out="84,16”parent="vgg / conv4_3”name="output1”> .

[0164] The dataset sub-element is used to specify the dataset and set related parameters. For example, the dataset sub-element is <dataset>.

[0165] In addition, the dataset sub-element is usually the first sub-element of the root element, and the hyperparameter sub-element is usually the second sub-element of the root element.

[0166] Here, the error prompt information is used to prompt the user to modify the XML text, that is, the user can modify the XML text according to the error prompt information. The error prompt information may include the position where the XML text does not conform to the specification, so that the user can quickly determine the position where the XML text is written incorrectly, thereby speeding up the modification of the XML text. In addition, after the user completes the modification, the modified XML text will be re-entered.

[0167] The XML-based neural network model generation method provided by the embodiment of the present invention obtains the input XML text and verifies the XML text based on the preset standardized definition specification; if the XML text conforms to the standardized definition specification, the XML text is determined as an XML instance text; if the XML text does not conform to the standardized definition specification, an error prompt message is issued, and the error prompt message is used to prompt the XML text to be modified; return to the step of obtaining the input XML text until the modified XML text conforms to the standardized definition specification. In the above manner, the XML text is verified based on the standardized definition specification to ensure that the XML implementation text that conforms to the standardized definition specification is parsed, thereby ensuring that the target neural network model can be obtained.

[0168] The generation device of the neural network model provided by the present invention is described below. The generation device of the neural network model described below and the XML-based neural network model generation method described above can refer to each other.

[0169] In this embodiment, the generating device of the neural network model includes:

[0170] An acquisition module, used to acquire an XML instance text to be parsed, wherein the XML instance text defines a deep learning framework attribute and its deep learning framework attribute value;

[0171] A parsing module is used to parse the XML instance text and generate a target neural network model corresponding to the deep learning framework attribute value.

[0172] The neural network model generation device provided by the embodiment of the present invention obtains an XML instance text to be parsed, wherein the XML instance text defines deep learning framework attributes and their deep learning framework attribute values; the XML instance text is parsed to generate a target neural network model corresponding to the deep learning framework attribute values. In the above manner, the deep learning framework attributes and their deep learning framework attribute values ​​are defined in the XML instance text, so that when the XML instance text is subsequently parsed, a neural network model corresponding to the defined deep learning framework can be generated. Based on this, for different deep learning frameworks, it is only necessary to modify the deep learning framework attribute values. The XML instance text is compatible with multiple deep learning frameworks, and the user does not need to design, implement and debug his own unique model for each deep learning framework, thereby greatly reducing the workload of generating the neural network model.

[0173] Based on any of the above embodiments, the parsing module is further used for:

[0174] Initializing based on the attributes of the root element in the XML instance text to obtain an initial neural network model;

[0175] Determine a first target sub-element to be parsed in the root element, and call a corresponding parsing algorithm based on the type of the first target sub-element to parse the first target sub-element to update the initial neural network model;

[0176] Return to the step of determining the first target sub-element to be parsed in the root element until all sub-elements in the root element are parsed, and the last updated initial neural network model is determined as the target neural network model.

[0177] Based on any of the above embodiments, the root element includes a super parameter sub-element, and the execution steps of the parsing algorithm corresponding to the super parameter sub-element are as follows:

[0178] The attribute value of the hyperparameter attribute in the hyperparameter sub-element is read, and the initial neural network model is updated based on the read attribute value of the hyperparameter attribute.

[0179] Based on any of the above embodiments, the root element includes a network layer sub-element, and the execution steps of the parsing algorithm corresponding to the network layer sub-element are as follows:

[0180] Reading the attribute value of the network layer attribute in the network layer sub-element, and constructing the target neural network layer to be added based on the attribute value of the read network layer attribute;

[0181] Adding the target neural network layer to the initial neural network model;

[0182] The root element also includes an element-by-element operation layer sub-element, and the execution steps of the parsing algorithm corresponding to the element-by-element operation layer sub-element are as follows:

[0183] Read the attribute value of the element-by-element operation layer attribute in the element-by-element operation layer sub-element, and construct the target element-by-element operation layer to be added based on the read attribute value of the element-by-element operation layer attribute;

[0184] The target element-wise operation layer is added to the initial neural network model.

[0185] Based on any of the above embodiments, the root element includes a block sub-element, and the execution steps of the parsing algorithm corresponding to the block sub-element are as follows:

[0186] Based on the attributes of the block sub-element, initialization is performed to obtain an initial block;

[0187] Determine a second target sub-element to be parsed currently in the block sub-element, and call a corresponding parsing algorithm based on the type of the second target sub-element to parse the second target sub-element, so as to update the initial block;

[0188] Return to the step of determining the second target sub-element currently to be parsed in the block sub-element until all sub-elements in the block sub-element are parsed, and add the last updated initial block to the initial neural network model.

[0189] Based on any of the above embodiments, the deep learning framework attributes are defined in the root element;

[0190] The initialization based on the attributes of the root element in the XML instance text to obtain an initial neural network model includes:

[0191] Based on the deep learning framework attribute value of the root element, initialization is performed to obtain an initial neural network model corresponding to the deep learning framework attribute value.

[0192] Based on any of the above embodiments, the XML instance text includes a root element, and the root element includes a hyperparameter sub-element, a network layer sub-element, an element-by-element operation layer sub-element, a block sub-element, and a data set sub-element, and the block sub-element includes a network layer sub-element and an element-by-element operation layer sub-element.

[0193] Based on any of the above embodiments, the acquisition module is further used for:

[0194] Obtaining an input XML text, and verifying the XML text based on a preset standardization definition specification;

[0195] If the XML text complies with the standardization definition specification, determining the XML text as the XML instance text;

[0196] If the XML text does not conform to the standard definition specification, an error prompt message is issued, and the error prompt message is used to prompt the XML text to be modified;

[0197] Return to the step of obtaining the input XML text until the modified XML text complies with the standardization definition specification.

[0198] Figure 3 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 3 As shown, the electronic device may include: a processor 310, a communication interface 320, a memory 330 and a communication bus 340, wherein the processor 310, the communication interface 320 and the memory 330 communicate with each other through the communication bus 340. The processor 310 may call the logic instructions in the memory 330 to execute the XML-based neural network model generation method, the method comprising: obtaining an XML instance text to be parsed, wherein the XML instance text defines a deep learning framework attribute and its deep learning framework attribute value; parsing the XML instance text to generate a target neural network model corresponding to the deep learning framework attribute value.

[0199] In addition, the logic instructions in the above-mentioned memory 330 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0200] On the other hand, the present invention also provides a computer program product, which includes a computer program, and the computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the XML-based neural network model generation method provided by the above-mentioned methods, and the method includes: obtaining an XML instance text to be parsed, wherein the XML instance text defines deep learning framework attributes and their deep learning framework attribute values; parsing the XML instance text to generate a target neural network model corresponding to the deep learning framework attribute values.

[0201] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it is implemented to execute the XML-based neural network model generation method provided by the above-mentioned methods, the method comprising: obtaining an XML instance text to be parsed, wherein the XML instance text defines deep learning framework attributes and their deep learning framework attribute values; parsing the XML instance text to generate a target neural network model corresponding to the deep learning framework attribute values.

[0202] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0203] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0204] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.< / dataset> < / vertex> < / layer> < / parameters> < / parameters> < / dataset> < / vertex> < / layer> < / parameters> < / model> < / model> < / parameters> < / updater> < / layer> < / layer> < / layer> < / model>

Claims

1. A method for generating a neural network model based on XML, characterized in that: include: Obtain an XML instance text to be parsed, wherein the XML instance text defines a deep learning framework attribute and its deep learning framework attribute value; Parsing the XML instance text to generate a target neural network model corresponding to the deep learning framework attribute value; The step of parsing the XML instance text to generate a target neural network model corresponding to the deep learning framework attribute value includes: Initializing based on the attributes of the root element in the XML instance text to obtain an initial neural network model; Determine a first target sub-element to be parsed in the root element, and call a corresponding parsing algorithm based on the type of the first target sub-element to parse the first target sub-element to update the initial neural network model; Return to the step of determining the first target sub-element to be parsed in the root element until all sub-elements in the root element are parsed, and the last updated initial neural network model is determined as the target neural network model.

2. The XML-based neural network model generation method according to claim 1, characterized in that: The root element includes a hyperparameter sub-element, and the execution steps of the parsing algorithm corresponding to the hyperparameter sub-element are as follows: The attribute value of the hyperparameter attribute in the hyperparameter sub-element is read, and the initial neural network model is updated based on the read attribute value of the hyperparameter attribute.

3. The XML-based neural network model generation method according to claim 1, characterized in that: The root element includes a network layer sub-element, and the execution steps of the parsing algorithm corresponding to the network layer sub-element are as follows: Reading the attribute value of the network layer attribute in the network layer sub-element, and constructing the target neural network layer to be added based on the attribute value of the read network layer attribute; Adding the target neural network layer to the initial neural network model; The root element also includes an element-by-element operation layer sub-element, and the execution steps of the parsing algorithm corresponding to the element-by-element operation layer sub-element are as follows: Read the attribute value of the element-by-element operation layer attribute in the element-by-element operation layer sub-element, and construct the target element-by-element operation layer to be added based on the read attribute value of the element-by-element operation layer attribute; The target element-wise operation layer is added to the initial neural network model.

4. The XML-based neural network model generation method according to claim 1, characterized in that: The root element includes a block sub-element, and the execution steps of the parsing algorithm corresponding to the block sub-element are as follows: Based on the attributes of the block sub-element, initialization is performed to obtain an initial block; Determine a second target sub-element to be parsed currently in the block sub-element, and call a corresponding parsing algorithm based on the type of the second target sub-element to parse the second target sub-element, so as to update the initial block; Return to the step of determining the second target sub-element currently to be parsed in the block sub-element until all sub-elements in the block sub-element are parsed, and add the last updated initial block to the initial neural network model.

5. The XML-based neural network model generation method according to claim 1, characterized in that: The deep learning framework attributes are defined in the root element; The initialization based on the attributes of the root element in the XML instance text to obtain an initial neural network model includes: Based on the deep learning framework attribute value of the root element, initialization is performed to obtain an initial neural network model corresponding to the deep learning framework attribute value.

6. The XML-based neural network model generation method according to any one of claims 1 to 5, characterized in that: The XML instance text includes a root element, and the root element includes a hyperparameter sub-element, a network layer sub-element, an element-by-element operation layer sub-element, a block sub-element, and a data set sub-element. The block sub-element includes a network layer sub-element and an element-by-element operation layer sub-element.

7. The XML-based neural network model generation method according to any one of claims 1 to 5, characterized in that: The step of obtaining the XML instance text to be parsed includes: Obtaining an input XML text, and verifying the XML text based on a preset standardization definition specification; If the XML text complies with the standardization definition specification, determining the XML text as the XML instance text; If the XML text does not conform to the standard definition specification, an error prompt message is issued, and the error prompt message is used to prompt the XML text to be modified; Return to the step of obtaining the input XML text until the modified XML text complies with the standardization definition specification.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the XML-based neural network model generation method as described in any one of claims 1 to 7 are implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the XML-based neural network model generation method as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Deep learning model file conversion method and system, computer equipment and computer readable storage medium

    CN111275199A