Method and device for training neural network model

By determining the depth and width in the neural network model and selecting the appropriate random number to select the sub-neural network model for training, the problem of how to adapt to the information processing of different electronic devices is solved, and efficient model training and application is achieved.

CN120031089APending Publication Date: 2025-05-23TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311584824.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

How to quickly and efficiently train neural network models of various sizes to adapt to electronic devices with different performances and realize information processing on different electronic devices.

Method used

By determining the depth and width of the first neural network model, and selecting random numbers based on these dimensions to select sub-neural network models, performing sample processing and parameter updates, ensuring that the model has good performance during training and application.

Benefits of technology

After completing the training of the neural network model, the sub-neural network model of any size is directly selected from the complete large-size neural network model according to the processing capabilities of different electronic devices, which avoids the need for secondary training and improves training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031089A_ABST
    Figure CN120031089A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method and device for training a neural network model, a computer program product and a storage medium. The method comprises the following steps: determining the depth and width of a first neural network model, and selecting a sub-neural network model with any depth and width from the first neural network model as a second neural network model, taking the neural network model except the second neural network model in the first neural network model as a third neural network model; and obtaining samples used for training a number of first neural network models and sample labels, so as to train the first neural network models. According to the method disclosed by the invention, the trained neural network model can be conveniently adapted to various electronic devices with different performances.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, and more specifically, to a method, device, computer program product, and storage medium for training a neural network model, as well as a method, device, computer program product, and storage medium for performing information processing based on a neural network model in an electronic device. Background Art

[0002] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.

[0003] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. Basic artificial intelligence technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, pre-trained models are also called large models and basic models. After fine-tuning, they can be widely used in downstream tasks in various major directions of artificial intelligence. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0004] Neural Network (NN), as an important branch of artificial intelligence, is a network structure that imitates the behavioral characteristics of animal neural networks to process information. The structure of a neural network is composed of a large number of nodes (or neurons) connected to each other. It processes information by learning and training the input information based on a specific operation model. A neural network includes an input layer, a hidden layer, and an output layer. The input layer is responsible for receiving input signals, the output layer is responsible for outputting the calculation results of the neural network, and the hidden layer is responsible for the calculation process such as learning and training. It is the memory unit of the network. The memory function of the hidden layer is represented by a weight matrix. Usually, each neuron corresponds to a weight coefficient.

[0005] The size of a neural network model is usually reflected by two key dimensions: width and depth, which have an important impact on the learning ability and performance of the model. Depth refers to the number of layers in a neural network. A deeper network has better nonlinear expression capabilities and can learn more complex transformations, thereby fitting more complex features. However, deepening the network can also bring about problems such as gradient instability and network degradation. When the depth reaches a certain level, the performance may not improve, or may even decline. Width refers to the number of neurons in each layer. Sufficient width can ensure that each layer learns rich features. Too narrow a width will lead to insufficient feature extraction, insufficient learning of information, and limited model performance. Width will also contribute a lot of computation, and a network that is too wide may increase the computational burden of the model. When designing a neural network model, you can start with the two dimensions of width and depth to balance the performance of the model with computational efficiency.

[0006] As the types of electronic devices become more and more diverse, how to quickly and efficiently train neural network models of various sizes to apply to various electronic devices with different performances, so as to realize information processing on different electronic devices, is one of the key research directions in this field. Summary of the invention

[0007] In order to make the trained neural network model conveniently adaptable to various electronic devices with different performances, the present disclosure provides a method for training a neural network model, including: determining the depth and width of a first neural network model as the first depth and the first width, respectively, wherein the values ​​of the first depth and the first width are greater than 1, the first depth indicates the number of multiple identical network modules cascaded in the first neural network model, and the first width indicates the number of multiple identical network modules connected in parallel in the first neural network model; determining a first random number and a second random number, wherein the first random number is less than or equal to the first depth, and the second random number is less than or equal to the first width; selecting a part of the first neural network model with the second depth and the second width as the second neural network model in a predetermined order, and selecting a neural network model other than the second neural network model in the first neural network model as a third neural network model; wherein the second depth is equal to the first random number, and the second width is equal to the second random number; obtaining samples and sample labels for training the first neural network model; processing the samples using the first neural network model; and updating the parameters of the second neural network model based on the processing results of the samples by the second neural network model and the sample labels, and updating the parameters of the third neural network model based on the processing results of the samples by the first neural network model and the sample labels, so as to obtain the first neural network model after parameter update.

[0008] An embodiment of the present disclosure also provides a method for performing information processing based on a neural network model in an electronic device, comprising: obtaining information to be processed; obtaining a trained first neural network model; selecting, in a predetermined order and according to the processing capability of the electronic device, a part of the first neural network model having a third depth and a third width as an information processing neural network model, wherein the first neural network model has a first depth and a first width, the values ​​of the first depth and the first width are greater than 1, the first depth indicates the number of multiple identical network modules cascaded in the first neural network model, the first width indicates the number of multiple identical network modules connected in parallel in the first neural network model, the third depth is less than or equal to the first depth, and the third width is less than or equal to the first width; processing the information to be processed based on the information processing neural network model to obtain an information processing result for the information to be processed.

[0009] An embodiment of the present disclosure also provides an apparatus for training a neural network model, comprising: a dimension determination unit, configured to: determine the depth and width of a first neural network model, as a first depth and a first width, respectively, wherein the values ​​of the first depth and the first width are greater than 1, the first depth indicates the number of multiple identical network modules cascaded in the first neural network model, and the first width indicates the number of multiple identical network modules connected in parallel in the first neural network model; a random number determination unit, configured to: determine a first random number and a second random number, wherein the first random number is less than or equal to the first depth, and the second random number is less than or equal to the first width; a model selection unit, configured to: select a model having a second depth and a second width in the first neural network model in a predetermined order. Part of the first neural network model is used as the second neural network model, and the neural network models other than the second neural network model in the first neural network model are used as the third neural network model; wherein the second depth is equal to the first random number, and the second width is equal to the second random number; the sample information acquisition unit is configured to: obtain samples and sample labels for training the first neural network model; the sample processing unit is configured to: process the samples using the first neural network model; and the model parameter updating unit is configured to: update the parameters of the second neural network model based on the processing results of the samples by the second neural network model and the sample labels, and update the parameters of the third neural network model based on the processing results of the samples by the first neural network model and the sample labels.

[0010] An embodiment of the present disclosure also provides an apparatus for performing information processing based on a neural network model in an electronic device, comprising: an information acquisition unit, configured to: acquire information to be processed; a model acquisition unit, configured to: acquire a trained first neural network model; a model selection unit, configured to: select, in a predetermined order, a part of the first neural network model having a third depth and a third width as an information processing neural network model according to the processing capability of the electronic device, wherein the first neural network model has a first depth and a first width, the values ​​of the first depth and the first width are greater than 1, the first depth indicates the number of multiple identical network modules cascaded in the first neural network model, the first width indicates the number of multiple identical network modules connected in parallel in the first neural network model, the third depth is less than or equal to the first depth, and the third width is less than or equal to the first width; a result acquisition unit, configured to: process the information to be processed based on the information processing neural network model to obtain an information processing result for the information to be processed.

[0011] An embodiment of the present disclosure further provides a computer program product, which includes computer software code. When the computer software code is executed by a processor, the above method is provided.

[0012] An embodiment of the present disclosure further provides a computer-readable storage medium having computer-executable instructions stored thereon, and when the instructions are executed by a processor, the above method is provided.

[0013] Since the method for training a neural network model disclosed in the present invention not only considers the processing results of samples by a sub-neural network model selected from the complete large-size neural network model to update the parameters of the sub-neural network model, but also considers the processing results of samples by the complete large-size neural network model to update the neural network models other than the sub-neural network model in the complete large-size neural network model, when training the neural network model, it is equivalent to training both the complete large-size neural network model and the small-size sub-neural network model. Therefore, after completing the training of the neural network model, according to the processing capabilities of different electronic devices, a sub-neural network model of any size directly selected from the complete large-size neural network model and applied to the electronic device will have good model performance, without the need for secondary training of the selected small-size sub-neural network model, which effectively improves the efficiency of training neural network models for application to various electronic devices with different performances. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some exemplary embodiments of the present disclosure, and a person of ordinary skill in the art can obtain other drawings based on these drawings without creative work.

[0015] Herein, in the accompanying drawings:

[0016] Figure 1 A schematic diagram of an application scenario according to an embodiment of the present disclosure is shown.

[0017] Figure 2 It is an example schematic diagram showing a scenario of information processing and training based on a neural network model according to an embodiment of the present disclosure.

[0018] Figure 3A is a schematic diagram showing a process of training a neural network model according to an embodiment of the present disclosure;

[0019] Figure 3B is a schematic diagram showing a process of selecting a neural network model according to the performance of an electronic device according to an embodiment of the present disclosure;

[0020] Figure 4A is a schematic flow chart showing a method for training a neural network model according to an embodiment of the present disclosure;

[0021] Figure 4B is a schematic diagram showing the structure of a first neural network model according to an embodiment of the present disclosure;

[0022] Figure 5 is a schematic diagram showing the structure of a sub-network according to an embodiment of the present disclosure;

[0023] Figure 6 is a schematic flow chart showing a method for performing information processing based on a neural network model in an electronic device according to an embodiment of the present disclosure;

[0024] Figure 7 is a schematic diagram showing the composition of an apparatus for training a neural network model according to an embodiment of the present disclosure;

[0025] Figure 8 is a schematic diagram showing the composition of an apparatus for performing information processing based on a neural network model in an electronic device according to an embodiment of the present disclosure; and

[0026] Fig. 9 is an illustration of an architecture of a computing device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0027] In order to make the purpose, technical solution and advantages of the present disclosure more obvious, the exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described here.

[0028] Furthermore, in the present specification and the drawings, steps and elements having substantially the same or similar features are denoted by the same or similar reference numerals, and repeated description of these steps and elements will be omitted.

[0029] In addition, in this specification and the accompanying drawings, elements are described in singular or plural form, depending on the embodiment. However, the singular and plural forms are appropriately selected for the proposed situation only for the convenience of explanation and are not intended to limit the present disclosure thereto. Therefore, the singular form may include the plural form, and the plural form may also include the singular form, unless the context clearly indicates otherwise.

[0030] In this specification and the accompanying drawings, substantially the same or similar steps or elements are represented by the same or similar reference numerals, and repeated descriptions of these steps and elements will be omitted. At the same time, in the description of the present disclosure, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance or ranking.

[0031] Various neural networks (or neural network models) that can be used in the embodiments of the present disclosure below can all be artificial intelligence models, especially artificial intelligence-based neural network models. Typically, artificial intelligence-based neural network models are implemented as acyclic graphs in which neurons are arranged in different layers. Typically, a neural network model includes an input layer and an output layer, which are separated by at least one hidden layer. The hidden layer transforms the input received by the input layer into a representation useful for generating outputs in the output layer. The network nodes (i.e., neurons) are fully connected to the nodes in the adjacent layers via edges, and there are no edges between the nodes in each layer. The data received at the nodes of the input layer of the neural network are propagated to the nodes of the output layer via any one of the hidden layer, activation layer, pooling layer, convolutional layer, etc. The input and output of the neural network model can take various forms, and the present disclosure is not limited to this.

[0032] The embodiments of the present disclosure will be further described below in conjunction with the accompanying drawings.

[0033] First refer to Figure 1 Describe the application scenarios of the method and corresponding device according to the embodiments of the present disclosure. Figure 1 A schematic diagram of an application scenario 100 according to an embodiment of the present disclosure is shown, wherein a server 110 and multiple terminals 120 are schematically shown.

[0034] The neural network model for information processing in the embodiment of the present disclosure can be integrated into various electronic devices, for example, Figure 1 The server 110 and any electronic device in the multiple terminals 120. For example, the information processing neural network model can be integrated in the terminal 120. The terminal 120 can be a mobile phone, a tablet computer, a laptop computer, a desktop computer, a personal computer (PC, Personal Computer), a smart speaker or a smart watch, etc., but is not limited to this. For another example, the information processing neural network model can also be integrated in the server 110. The server 110 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms. The terminal and the server can be directly or indirectly connected via wired or wireless communication, and the present disclosure is not limited here.

[0035] It is understandable that the electronic device that performs information processing based on the neural network model according to the embodiment of the present disclosure can be a terminal, a server, or a system composed of a terminal and a server. The method for performing information processing based on the neural network model in an electronic device according to the embodiment of the present disclosure can be executed on a terminal, on a server, or jointly by a terminal and a server.

[0036] The information processing neural network model provided by the embodiments of the present disclosure can be used to perform various information processing tasks, including: information extraction (e.g., key information extraction, information search, feature extraction, speech separation, etc.), information classification (e.g., image classification, pathology classification, junk information identification), information restoration (e.g., image repair, incomplete information prediction, etc.), information style transfer (e.g., image style transfer, timbre conversion, etc.), information enhancement (e.g., image clarity enhancement, audio denoising, etc.), information mining (e.g., big data mining, network information mining, etc.), information identification (e.g., disease diagnosis, junk information identification, etc.), machine translation, etc., and are not limited thereto. In the present disclosure, the information processed based on the neural network model may include: images, text, audio, video, number series, etc.

[0037] The information processing neural network model provided by the embodiments of the present disclosure may also relate to artificial intelligence cloud services in the field of cloud technology. Among them, cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data calculation, storage, processing, and sharing. Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model, which can form a resource pool, be used as needed, and be flexible and convenient. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the background system for logical processing. Data at different levels will be processed separately, and various types of industry data require the support of a powerful system background, which can only be achieved through cloud computing.

[0038] Among them, artificial intelligence cloud services are generally also referred to as AIaaS (AI as a Service, Chinese for "AI as a Service"). This is currently the mainstream service method of an artificial intelligence platform. Specifically, the AIaaS platform will split several common AI services and provide independent or packaged services in the cloud. This service model is similar to opening an AI-themed mall: all developers can access and use one or more artificial intelligence services provided by the platform through the application programming interface (API, Application Programming Interface). Some senior developers can also use the AI framework and AI infrastructure provided by the platform to deploy and operate exclusive cloud artificial intelligence services.

[0039] Figure 2 It is an example schematic diagram showing a scenario 200 for information processing and training based on a neural network model according to an embodiment of the present disclosure.

[0040] In the training stage, the server 110 can train the neural network model based on training samples. After training is completed, the neural network model can be pruned, and then the pruned neural network model can be deployed to one or more servers (or cloud services) to provide artificial intelligence services related to information processing based on the neural network model.

[0041] It is worth noting that all training samples used in this disclosure comply with the legality, morality and privacy requirements of laws and regulations. Specifically, the sources of all training samples are legal and have been explicitly permitted by the user during the collection process. In addition, all training samples used in this disclosure comply with the privacy protection principle. These training samples have been strictly screened and cleaned, and will not be disclosed to any third party without explicit authorization.

[0042] In the stage of information processing based on the neural network model, it is assumed that a client or application (e.g., an image processing application, a text processing application, etc.) that interacts with the server 110 for information processing has been installed on the user terminal 120 that performs information processing. The user terminal 120 can send an information processing request to the server 110 corresponding to the application through the network to request the neural network deployed on the server 110 to process the information, wherein the neural network deployed on the server 110 can be a pruned neural network model. According to an embodiment of the present disclosure, after the server 110 receives the information processing request, the pruned neural network model can be used to perform information processing in response to the information processing request, and the predicted information processing result can be fed back to the user terminal 120. The user terminal 120 can receive the information processing result. Afterwards, the user terminal 120 can perform further analysis or processing based on the information processing result.

[0043] It is worth noting that Figure 2 The training sample data shown in can also be updated in real time. For example, the user can score the information processing result. For example, if the user believes that the rationality and accuracy of the information processing result are high, the user can give a high score to the information processing result, and the server 110 can use the information processing result as a positive sample for real-time training of the neural network model. If the user gives a low score to the information processing result, the server 110 can use the information processing result as a negative sample.

[0044] Figure 2 The training sample set shown in can also be set in advance. For example, refer to Figure 2 The server can obtain training data (eg, image training samples, text training samples, audio training samples, video training samples, etc.) from the database, and then generate a training sample set for the neural network model. Of course, the present disclosure is not limited to this.

[0045] Figure 3A is a schematic diagram illustrating a process of training a neural network model according to an embodiment of the present disclosure.

[0046] like Figure 3AAs shown, the depth of the first neural network model N0 can be determined to be the first depth D, and the width can be determined to be the first width W, wherein the values ​​of D and W are greater than 1, and the depth D of the first neural network model N0 is used to represent the number of multiple identical network modules cascaded in the first neural network model. The width W of the first neural network model N0 is used to represent the number of multiple identical network modules connected in parallel in the first neural network model. Figure 3A Each of the same network modules is represented by a circle, Figure 3A In the example, the depth D=4 and W=4 of the first neural network model N0 are taken as examples for illustration but not limitation.

[0047] In the process of training the first neural network model N0, a first random number d and a second random number w can be first determined, wherein the first random number and the second random number are positive integers, and the first random number is less than or equal to D, and the second random number is less than or equal to W.

[0048] Then, in a predetermined order, the part with the second depth and the second width in the first neural network model N0 can be selected as the second neural network model, and the neural network model other than the second neural network model in the first neural network model N0 can be selected as the third neural network model; wherein the second depth is equal to the first random number, and the second width is equal to the second random number. Assuming that the first random number is 3 and the second random number is 4, then, Figure 3A As shown on the right, the first neural network model N0 may include the second neural network model shown in the black solid circle in the figure (the depth of the second neural network model is 3 and the width is 4), and the third neural network model shown in the white hollow circle in the figure (the depth of the third neural network model is 1 and the width is 4).

[0049] By inputting the training sample X into the first neural network model N0, the result Y processed by the first neural network model N0 and the result Y' processed by the second neural network model can be obtained respectively. By simultaneously considering the difference between the result Y and the result Y' and the training sample label, the first neural network model N0 can be trained.

[0050] For samples of different batches, second neural network models of different sizes can be obtained by random sampling, so as to train the first neural network model (including the second neural network model). After training the model based on multiple batches of samples, it is equivalent to training the complete first neural network model and multiple second neural network models of different sizes. Therefore, the small-sized neural network sub-model obtained by trimming the trained first neural network model also has good performance and no secondary training is required.

[0051] Figure 3B is a schematic diagram illustrating a process of selecting a neural network model according to the performance of an electronic device according to an embodiment of the present disclosure.

[0052] like Figure 3B As shown, after completing Figure 3A After training of the first neural network model N0,

[0053] The part with the third depth and the third width in the first neural network model can be selected as the information processing neural network model according to the processing capability of the electronic device, so that the electronic device can use a suitable information processing neural network model to perform information processing (for example, image recognition, audio separation, etc.), wherein the third depth is less than or equal to the depth D of the first neural network model N0, and the third width is less than or equal to the width W of the first neural network model N0.

[0054] According to an embodiment of the present disclosure, when the electronic device is a mobile phone 301, a small-sized sub-neural network model N1 can be selected from the first neural network model as an information processing neural network model; when the electronic device is a computer 302, a medium-sized sub-neural network model N2 can be selected from the first neural network model as an information processing neural network model; when the electronic device is a server 303, a large-sized sub-neural network model N3 can be selected from the first neural network model as an information processing neural network model (the first neural network model can also be used directly).

[0055] It should be understood that the second neural network model selected in the training phase and the information processing neural network model selected in the information processing phase (i.e., the model inference phase) are selected in the same way in terms of depth and width. Figure 3A and 3B As shown, if a neural network model with a width of w and a depth of d is to be selected from the first neural network model as the second neural network model or the information processing neural network model, the depth and / or width in the first neural network model can be numbered and selected in ascending order. By making the selection method of the second neural network model and the information processing neural network model the same, the consistency of the model in the training stage and the information processing stage can be ensured, and the processing result of the information is more accurate.

[0056] Figure 4A 4 is a schematic flowchart illustrating a method 400 for training a neural network model according to an embodiment of the present disclosure.

[0057] In which, in step S410, the depth and width of the first neural network model are determined as the first depth and the first width, respectively, wherein the values ​​of the first depth and the first width are greater than 1, the first depth indicates the number of multiple identical network modules cascaded in the first neural network model, and the first width indicates the number of multiple identical network modules connected in parallel in the first neural network model.

[0058] It should be understood that the network module can be various sub-neural network models or neural network layers. For example, each of the network modules can include: a convolutional neural network, a neural network based on an attention mechanism, a recurrent neural network, a recursive neural network, a feedforward neural network, a generative adversarial neural network, a deep neural network, or a combination thereof.

[0059] In step S420, a first random number and a second random number are determined, wherein the first random number is less than or equal to the first depth, and the second random number is less than or equal to the first width.

[0060] In step S430, in a predetermined order, a part of the first neural network model having a second depth and a second width is selected as a second neural network model, and a neural network model other than the second neural network model in the first neural network model is selected as a third neural network model; wherein the second depth is equal to the first random number, and the second width is equal to the second random number.

[0061] According to an embodiment of the present disclosure, multiple identical network modules cascaded in the first neural network model can be numbered in a cascade order to obtain a first serial number, and multiple identical network modules in parallel in the first neural network model can be numbered in a parallel order to obtain a second serial number; then, the part of the first neural network model whose first serial number is less than or equal to the first random number and whose second serial number is less than or equal to the second random number is selected as the second neural network model.

[0062] According to an embodiment of the present disclosure, the first neural network model may include a plurality of cascaded first sub-networks, the number of the plurality of first sub-networks being equal to the first depth, wherein each of the plurality of first sub-networks includes a plurality of parallel network modules whose number is equal to the first width, wherein each of the network modules is initialized based on different parameters. Similarly, the second neural network model may include a plurality of cascaded second sub-networks, the number of the plurality of second sub-networks being equal to the second depth, wherein each of the plurality of second sub-networks includes a plurality of parallel network modules whose number is equal to the second width.

[0063] For a plurality of cascaded first sub-networks, the output features of a previous first sub-network may be used as input features of a next first sub-network, and the dimensions of the output features and the input features may be the same.

[0064] It should be understood that the first sub-network and the second sub-network may include only the network module or other network modules in addition to the network module. For example, at least one of the first sub-network and the second sub-network may also include at least one of a pre-processing module and a post-processing module different from the network module. Optionally, the pre-processing module may include at least one of a feature extraction network and a dimensionality transformation network, and the post-processing module may include at least one of a dimensionality transformation network and a residual network.

[0065] Optionally, the dimension of the input features of each of the first sub-networks or each of the second sub-networks may be the same as the dimension of the output features.

[0066] It should be understood that the width of each second sub-network in the selected second neural network model can be the same or different. In the case where the width of each second sub-network is the same (for example, Figure 3A In the example shown, the width of each second subnetwork is 4), and only one second random number can be generated to determine the width of the second subnetwork. In the case where the width of each second subnetwork is different, the first random number can be determined first, and the first random number is used to determine the number of first subnetworks to be selected in the first neural network model. Then, for each first subnetwork to be selected, a second random number is generated, wherein the second random number is used to determine the number of network modules in parallel in the first subnetwork to be selected. Then, from the multiple first subnetworks in the cascade, a first subnetwork whose cascade level is less than or equal to the second depth can be selected, and the second depth is equal to the first random number; for each selected first subnetwork, a network module in the subnetwork whose parallel order is before the second width is selected, and the second width is equal to the second random number generated for the first subnetwork.

[0067] For example, Figure 4B As shown, the first random number can be determined to be 3, that is, the second neural network model includes three cascaded first sub-networks. Then, for each first sub-network to be selected, a second random number is generated to determine the number of network modules in parallel in the first sub-network to be selected. Figure 4BIn the example shown, the second random numbers are 4, 2, and 3, respectively. By selecting a first subnetwork whose cascade level is less than or equal to the first random number (here 3) from the multiple first subnetworks in the cascade, and for each selected first subnetwork, selecting a network module in the subnetwork whose parallel connection order is before the second random number (respectively 4, 2, and 3), the second neural network model in the first neural network model N0 (that is, Figure 4B The part indicated by the black solid circle in the figure).

[0068] In step S440, samples and sample labels for training the first neural network model are obtained.

[0069] It should be understood that the samples and sample labels in the present disclosure may be pre-arranged and labeled by human experience, or may be from an existing model training database. The sample may include at least one of an image sample, a text sample, an audio sample, and a video sample.

[0070] In step S450, the sample is processed using the first neural network model.

[0071] Optionally, the samples may be preprocessed (for example, video samples may be processed into both images and audio, etc.), and then the processed samples may be further processed using the first neural network model.

[0072] According to an embodiment of the present disclosure, when the first neural network model includes a plurality of cascaded first sub-networks and the second neural network model includes a plurality of cascaded second sub-networks, for any one of the first sub-network and the second sub-network, the plurality of network modules in the sub-network can be used to process the information corresponding to the sample respectively to obtain a plurality of network module processing results, and the weighted sum of the plurality of network module processing results can be performed to obtain the processing result of the sub-network; and the processing result of the last one of the first sub-networks in the first neural network model is used as the processing result of the first neural network model for the sample, and the processing result of the last one of the second sub-networks in the second neural network model is used as the processing result of the second neural network model for the sample.

[0073] According to an embodiment of the present disclosure, each of the network module processing results can be divided into two processing sub-results, wherein the first of the two processing sub-results is used to determine the weight corresponding to the network module, and the second of the two network module processing sub-results is used to multiply the weight to obtain the result of the sub-network processing of the sample.

[0074] It should be understood that the first processing sub-result and the second processing sub-result may be sub-processing results divided according to different channels, or may be the same processing results, which is not limited herein.

[0075] In step S460, the parameters of the second neural network model are updated based on the processing results of the second neural network model on the sample and the sample label, and the parameters of the third neural network model are updated based on the processing results of the first neural network model on the sample and the sample label, so as to obtain the first neural network model with updated parameters.

[0076] According to an embodiment of the present disclosure, the value of the first loss function can be calculated based on the processing results of the first neural network model on the sample and the sample label, and the parameters of the third neural network model can be updated based on the value of the first loss function. The value of the second loss function can be calculated based on the processing results of the second neural network model on the sample and the sample label, and the parameters of the second neural network model can be updated based on the value of the second loss function.

[0077] For example, the loss function L used to train the first neural network model can be expressed by formula (1):

[0078] L=L obj (θ W,D )+L obj (θ w,d ) (1)

[0079] Among them, using the second loss function L obj (θ w,d ) Update the network parameters of the second neural network model, using the first loss function L obj (θ W,D ) Update the network parameters of the third neural network model, the first loss function L obj (θ W,D ) is used to reflect the error between the processing result of the first neural network model on the sample and the sample label, and the second loss function L obj (θ w,d ) is used to reflect the error between the processing result of the second neural network model on the sample and the sample label.

[0080] It should be understood that the loss function L obj The specific form of () can be determined according to the application scenario of the first neural network model. For example, for the image processing application scenario, the first loss function L obj (θ W,D ) is used to reflect the difference between the processing results of the first neural network model on the image samples and the image sample labels, and the second loss function L obj(θ w,d ) is used to reflect the difference between the processing results of the second neural network model on the image sample and the image sample labels. For text processing application scenarios, the first loss function L obj (θ W,D ) is used to reflect the difference between the processing result of the first neural network model on the text sample and the text sample label, and the second loss function L obj (θ w,d ) is used to reflect the difference between the processing results of the second neural network model on the text sample and the text sample labels.

[0081] When the updated first neural network model meets predetermined conditions (e.g., the loss function converges, the number of training times is reached, etc.), it can be determined that the training of the first neural network model is completed, and the trained first neural network model is obtained. At this time, the trained first neural network model can be used for feature extraction, speech separation, image classification, disease diagnosis, machine translation, etc.

[0082] According to an embodiment of the present disclosure, a method for training a neural network model for image processing includes: determining the depth and width of the neural network model for image processing, respectively as a first depth and a first width, wherein the values ​​of the first depth and the first width are greater than 1, the first depth indicates the number of multiple identical image processing network modules cascaded in the neural network model for image processing, and the first width indicates the number of multiple identical image processing network modules connected in parallel in the neural network model for image processing; determining a first random number and a second random number, wherein the first random number is less than or equal to the first depth, and the second random number is less than or equal to the first width; and selecting, in a predetermined order, a portion of the neural network model for image processing having a second depth and a second width as a first random number. An image processing submodel, and using the neural network model other than the first image processing submodel in the neural network model for image processing as a second image processing submodel; wherein the second depth is equal to the first random number, and the second width is equal to the second random number; obtaining image samples and image sample labels; processing the image samples using the neural network model for image processing; and updating the parameters of the first image processing submodel based on the processing results of the image samples by the first image processing submodel and the image sample labels, and updating the parameters of the second image processing submodel based on the processing results of the image samples by the neural network model for image processing and the image sample labels, so as to obtain the neural network model for image processing after parameter update. Then, optionally, the trained neural network model for image processing can be used to extract features of image data to further complete face recognition tasks or disease diagnosis and treatment tasks, etc.

[0083] According to an embodiment of the present disclosure, a method for training a neural network model for audio processing includes: determining the depth and width of the neural network model for audio processing, respectively as a first depth and a first width, wherein the values ​​of the first depth and the first width are greater than 1, the first depth indicates the number of multiple identical audio processing network modules cascaded in the neural network model for audio processing, and the first width indicates the number of multiple identical audio processing network modules connected in parallel in the neural network model for audio processing; determining a first random number and a second random number, wherein the first random number is less than or equal to the first depth, and the second random number is less than or equal to the first width; and selecting, in a predetermined order, a portion of the neural network model for audio processing having a second depth and a second width as a first random number. An audio processing submodel, and using the neural network model other than the first audio processing submodel in the neural network model for audio processing as a second audio processing submodel; wherein the second depth is equal to the first random number, and the second width is equal to the second random number; obtaining audio samples and audio sample labels; processing the audio samples using the neural network model for audio processing; and updating the parameters of the first audio processing submodel based on the processing results of the audio samples by the first audio processing submodel and the audio sample labels, and updating the parameters of the second audio processing submodel based on the processing results of the audio samples by the neural network model for audio processing and the audio sample labels, so as to obtain the neural network model for audio processing after parameter update. Then, optionally, the trained neural network model for audio processing can be used to extract features of audio data to further complete audio separation, audio classification, audio recognition, etc.

[0084] According to an embodiment of the present disclosure, a method for training a neural network model for text processing includes: determining the depth and width of the neural network model for text processing, respectively as a first depth and a first width, wherein the values ​​of the first depth and the first width are greater than 1, the first depth indicates the number of multiple identical text processing network modules cascaded in the neural network model for text processing, and the first width indicates the number of multiple identical text processing network modules connected in parallel in the neural network model for text processing; determining a first random number and a second random number, wherein the first random number is less than or equal to the first depth, and the second random number is less than or equal to the first width; selecting, in a predetermined order, a portion of the neural network model for text processing having a second depth and a second width as a first random number. A text processing submodel, and using the neural network model other than the first text processing submodel in the neural network model for text processing as a second text processing submodel; wherein the second depth is equal to the first random number, and the second width is equal to the second random number; obtaining a text sample and a text sample label; processing the text sample using the neural network model for text processing; and updating the parameters of the first text processing submodel based on the processing result of the text sample by the first text processing submodel and the text sample label, and updating the parameters of the second text processing submodel based on the processing result of the text sample by the neural network model for text processing and the text sample label, so as to obtain the neural network model for text processing after parameter update. Then, optionally, the trained neural network model for text processing can be used to extract features of text data to further complete semantic analysis, machine translation, etc.

[0085] Figure 5 It is a schematic diagram showing the structure of a sub-network according to an embodiment of the present disclosure, wherein the sub-network can be either a plurality of cascaded first sub-networks included in the above-mentioned first neural network model, or a plurality of cascaded second sub-networks included in the above-mentioned second neural network model.

[0086] exist Figure 5 In the example of , the width of the subnetwork can be defined as the number of fully connected (FC) layers. The traditional residual recurrent neural network (RNN) contains only one Figure 5 The FC layer shown in , and the input feature E is processed in series using the normalization layer (Norm), the RNN layer, and the FC layer to obtain the processed features, and the processed features are added to the input feature E to obtain the output feature F, where the dimensions of the input feature E and the output feature F are the same, that is, E∈R N×T , and F∈R N ×TThe present disclosure proposes that when the width of the subnetwork is not 1 (e.g. Figure 5 As shown in the figure, the number of FC layers is 3 and the width of the subnetwork is 3). For each FC layer, the results processed by the normalization layer and the RNN layer are processed to obtain two processed features: and Among them, G i After being processed by the Transformed Average Connection (TAC) module, each feature is obtained The corresponding weight By adding each feature A i The corresponding weight Q i Multiplying and adding it to feature E, we can get the output feature F, where Figure 5 The three FC layers in have the same structure but are initialized with different parameters.

[0087] Figure 5 The processing shown is equivalent to the processing when the neural network model depth is 1. When the model depth is greater than 1, multiple Figure 5 The sub-networks shown are cascaded, that is, the output features F of each sub-network are used as the input features E of the next sub-network.

[0088] It should be understood that Figure 5 Only one example of the structure of the sub-network is shown. In fact, the sub-network may also be other various network structures, which are not limited here.

[0089] Figure 6 is a schematic flowchart illustrating a method 600 for performing information processing based on a neural network model in an electronic device according to an embodiment of the present disclosure.

[0090] Wherein, in step S610, information to be processed is obtained.

[0091] It should be understood that the information to be processed may include at least one of image information, text information, audio information, and video information. Method 600 can be used to process one type of data (for example, machine translation only processes text information) or to process multiple types of data at the same time (for example, performing pathological analysis based on medical images and medical reports at the same time).

[0092] In step S620, a trained first neural network model is obtained.

[0093] It should be understood that the first neural network model can be trained by the above-mentioned method 400 for training a neural network model.

[0094] In step S630, in a predetermined order and according to the processing capability of the electronic device, a part of the first neural network model having a third depth and a third width is selected as an information processing neural network model, wherein the first neural network model has a first depth and a first width, the values ​​of the first depth and the first width are greater than 1, the first depth indicates the number of multiple identical network modules cascaded in the first neural network model, the first width indicates the number of multiple identical network modules connected in parallel in the first neural network model, the third depth is less than or equal to the first depth, and the third width is less than or equal to the first width.

[0095] According to an embodiment of the present disclosure, multiple identical network modules cascaded in the first neural network model can be numbered in a cascade order to obtain a first serial number, and multiple identical network modules in parallel in the first neural network model can be numbered in a parallel order to obtain a second serial number; and the part of the first neural network model in which the first serial number is less than or equal to the third depth and the second serial number is less than or equal to the third width is selected as the information processing neural network model.

[0096] In step S640, the information to be processed is processed based on the information processing neural network model to obtain an information processing result for the information to be processed.

[0097] According to an embodiment of the present disclosure, the information processing neural network model includes a plurality of cascaded third sub-networks, the number of the plurality of third sub-networks being equal to the third depth, wherein each of the plurality of third sub-networks includes a plurality of parallel network modules whose number is equal to the third width.

[0098] It should be understood that the third sub-network may include only the network module or other network modules in addition to the network module. For example, the third sub-network may include at least one of a pre-processing module and a post-processing module that are different from the network module. Optionally, the pre-processing module may include at least one of a feature extraction network and a dimensionality transformation network, and the post-processing module may include at least one of a dimensionality transformation network and a residual network.

[0099] It should be understood that the width of each third sub-network in the selected information processing neural network model can be the same (for example, Figure 3B The example shown), but can also be different (similar to Figure 4B example shown).

[0100] For each of the third sub-networks, the multiple network modules in the third sub-network can be used to process the information to be processed respectively to obtain multiple network module processing results; the multiple network module processing results are weighted summed to obtain the processing result of the third sub-network; and the processing result of the last third sub-network in the information processing neural network model is used as the information processing result for the information to be processed.

[0101] After the result of processing the information to be processed is obtained based on step S640, the processing result can be displayed on a screen such as Figure 1 On the terminal 120 shown. Optionally, as required, Figure 1 The server 110 or terminal 120 shown may further analyze or display the processing result.

[0102] During the training phase of the first neural network model, the indicators of the second neural network model selected at each depth and width can be counted, so that in the inference phase, according to the processing capability of the electronic device, an information processing neural network model of any width and depth can be selected from the first neural network model to match the size or complexity constraints of the electronic device.

[0103] For example, for the audio separation task of separating human voice from background music, the indicators of the second neural network model at various depths and widths can be obtained as shown in Table 1, including the signal-to-noise ratio (in dB) reflecting the accuracy of the model, the number of model parameters (in millions (M)), and the model complexity (measured by the number of multiply-accumulate operations (MACs) in gigabytes (G)). Assuming that a network model with a model complexity less than 4G is required, a model with a depth of 3 and a width of 2 can be selected.

[0104] Table 1

[0105]

[0106] It can also be seen from Table 1 that increasing the model depth can better improve the accuracy of the model processing information than increasing the model width. Therefore, when the model complexity index is fixed, a model with a larger depth value can be selected as much as possible.

[0107] Figure 7 is a schematic diagram showing the composition of an apparatus 700 for training a neural network model according to an embodiment of the present disclosure.

[0108] According to an embodiment of the present disclosure, the device 700 for training a neural network model may include: a dimension determination unit 710, a random number determination unit 720, a model selection unit 730, a sample information acquisition unit 740, a sample processing unit 750 and a model parameter updating unit 760.

[0109] In which, the dimension determination unit 710 can be configured to: determine the depth and width of the first neural network model, as the first depth and the first width, respectively, wherein the values ​​of the first depth and the first width are greater than 1, the first depth indicates the number of multiple identical network modules cascaded in the first neural network model, and the first width indicates the number of multiple identical network modules connected in parallel in the first neural network model.

[0110] The random number determination unit 720 may be configured to determine a first random number and a second random number, wherein the first random number is less than or equal to the first depth, and the second random number is less than or equal to the first width.

[0111] The model selection unit 730 can be configured to: select, in a predetermined order, a part of the first neural network model having a second depth and a second width as a second neural network model, and select neural network models other than the second neural network model in the first neural network model as a third neural network model; wherein the second depth is equal to the first random number, and the second width is equal to the second random number.

[0112] The sample information acquisition unit 740 may be configured to acquire samples and sample labels for training the first neural network model.

[0113] The sample processing unit 750 may be configured to process the sample using the first neural network model.

[0114] The model parameter updating unit 760 can be configured to update the parameters of the second neural network model based on the processing results of the second neural network model on the sample and the sample label, and update the parameters of the third neural network model based on the processing results of the first neural network model on the sample and the sample label.

[0115] It should be understood that Figure 7 The apparatus 700 for training a neural network model shown in FIG. Figure 4AThe various methods for training neural network models described herein. The dimension determination unit 710, the random number determination unit 720, the model selection unit 730, the sample information acquisition unit 740, the sample processing unit 750, and the model parameter updating unit 760 can respectively implement the processing of step S410, step S420, step S430, step S440, step S450, and step S460, which will not be described in detail herein.

[0116] Figure 8 is a schematic diagram showing the composition of an apparatus 800 for performing information processing based on a neural network model in an electronic device according to an embodiment of the present disclosure.

[0117] According to an embodiment of the present disclosure, an apparatus 800 for performing information processing based on a neural network model in an electronic device may include: an information acquisition unit 810 , a model acquisition unit 820 , a model selection unit 830 and a result acquisition unit 840 .

[0118] The information acquisition unit 810 may be configured to: acquire information to be processed.

[0119] The model acquisition unit 820 may be configured to: acquire a trained first neural network model.

[0120] The model selection unit 830 can be configured to: select a part of the first neural network model with a third depth and a third width as an information processing neural network model in a predetermined order according to the processing capability of the electronic device, wherein the first neural network model has a first depth and a first width, the values ​​of the first depth and the first width are greater than 1, the first depth indicates the number of multiple identical network modules cascaded in the first neural network model, the first width indicates the number of multiple identical network modules connected in parallel in the first neural network model, the third depth is less than or equal to the first depth, and the third width is less than or equal to the first width.

[0121] The result acquisition unit 840 may be configured to: process the information to be processed based on the information processing neural network model to obtain an information processing result for the information to be processed.

[0122] It should be understood that the device 800 for performing information processing based on the neural network model can be located at Figure 1 The server 110 shown may also be located at Figure 1 On the terminal 120 shown. Figure 8 The apparatus 800 for performing information processing based on a neural network model in an electronic device can be implemented as follows: Figure 6 The various methods for information processing based on neural network models described above. The first neural network model can be Figure 4A The information acquisition unit 810, the model acquisition unit 820, the model selection unit 830 and the result acquisition unit 840 can respectively implement the processing of step S610, step S620, step S630 and step S640, which will not be repeated here.

[0123] The present disclosure is based on a band separation recurrent neural network (BSRNN) network architecture for separating human voice from background music, and experiments are conducted on the MUSDB18-HQ benchmark dataset. The results are shown in Tables 2 and 3 below.

[0124] It can be seen from Table 2 below that, assuming that the depth D=12 and the width W=16 of the first neural network model, for the information processing neural network models (depth is d and width is w) of various sizes (reflected by depth and width) selected from the first neural network model, the neural network model trained using the method of the present invention (i.e., the neural network model trained using dynamic size) has more accurate audio signal processing results than the neural network model directly trained without using dynamic size.

[0125] Table 2

[0126]

[0127] It can be seen from Table 3 below that the time taken to train a complete neural network model (i.e., the first neural network model) using the method of the present invention (i.e., the last row of Table 3) is shorter than the sum of the time taken to train three sub-neural network models (information processing neural network models with a depth of d and a width of w) separately (i.e., the first 3 rows of Table 3).

[0128] Table 3

[0129]

[0130] It can be seen that the training method disclosed in the present invention can effectively improve the accuracy of model processing information and improve the efficiency of training neural network models for application in various electronic devices with different performances.

[0131] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuits or logic, general purpose hardware or controllers or other computing devices, or some combination thereof as non-limiting examples.

[0132] For example, the method or device according to the embodiment of the present disclosure may also be implemented by Fig. 9 The architecture of the computing device 3000 shown in FIG. Fig. 9 As shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage device in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for processing and / or communication of the method provided in the present disclosure and program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Fig. 9 The architecture shown is only exemplary and can be omitted according to actual needs when implementing different devices. Fig. 9 One or more components of a computing device are shown.

[0133] According to another aspect of the present disclosure, a computer-readable storage medium is also provided. Computer-readable instructions are stored on the computer storage medium. When the computer-readable instructions are executed by the processor, the method according to the embodiment of the present disclosure described with reference to the above figures can be executed. The computer-readable storage medium in the embodiment of the present disclosure can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM) or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of exemplary but not limiting description, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous connection dynamic random access memory (SLDRAM) and direct memory bus random access memory (DR RAM). It should be noted that memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory. It should be noted that memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0134] The embodiments of the present disclosure also provide a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. The processor of the computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the method according to the embodiments of the present disclosure.

[0135] In summary, the embodiments of the present disclosure provide a method, apparatus, computer program product, and storage medium for training a neural network model, as well as a method, apparatus, computer program product, and storage medium for performing information processing based on a neural network model in an electronic device.

[0136] The method for training a neural network model disclosed in the present invention includes: determining the depth and width of a first neural network model as the first depth and the first width, respectively, wherein the values ​​of the first depth and the first width are greater than 1, the first depth indicates the number of multiple identical network modules cascaded in the first neural network model, and the first width indicates the number of multiple identical network modules connected in parallel in the first neural network model; determining a first random number and a second random number, wherein the first random number is less than or equal to the first depth, and the second random number is less than or equal to the first width; selecting, in a predetermined order, a part of the first neural network model having the second depth and the second width as the second neural network model, and selecting a neural network model other than the second neural network model in the first neural network model as the third neural network model; wherein the second depth is equal to the first random number, and the second width is equal to the second random number; obtaining samples and sample labels for training the first neural network model; processing the samples using the first neural network model; and updating the parameters of the first neural network model in the following manner: updating the parameters of the second neural network model based on the processing result of the second neural network model on the samples and the sample labels, and updating the parameters of the third neural network model based on the processing result of the first neural network model on the samples and the sample labels.

[0137] Since the method for training a neural network model disclosed in the present invention not only considers the processing results of samples by a sub-neural network model selected from the complete large-size neural network model to update the parameters of the sub-neural network model, but also considers the processing results of samples by the complete large-size neural network model to update the neural network models other than the sub-neural network model in the complete large-size neural network model, when training the neural network model, it is equivalent to training both the complete large-size neural network model and the small-size sub-neural network model. Therefore, after completing the training of the neural network model, according to the processing capabilities of different electronic devices, a sub-neural network model of any size directly selected from the complete large-size neural network model and applied to the electronic device will have good model performance, without the need for secondary training of the selected small-size sub-neural network model, which effectively improves the efficiency of training neural network models for application to various electronic devices with different performances.

[0138] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment, or a part of a code, and the module, program segment, or a part of the code contains at least one executable instruction for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0139] The present disclosure uses specific words to describe the embodiments of the present disclosure. For example, "first / second embodiment", "one embodiment", and / or "some embodiments" refer to a certain feature, structure or characteristic related to at least one embodiment of the present disclosure. Therefore, it should be emphasized and noted that "one embodiment" or "an embodiment" or "an alternative embodiment" mentioned twice or more in different locations in this specification does not necessarily refer to the same embodiment. In addition, certain features, structures or characteristics in one or more embodiments of the present disclosure may be appropriately combined.

[0140] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0141] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the invention belongs. It should also be understood that terms such as those defined in common dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology and should not be interpreted in an idealized or extremely formal sense, unless explicitly defined as such herein.

[0142] The above is an explanation of the present invention and should not be considered as a limitation thereof. Although several exemplary embodiments of the present invention have been described, it will be readily appreciated by those skilled in the art that many modifications may be made to the exemplary embodiments without departing from the novel teachings and advantages of the present invention. Therefore, all such modifications are intended to be included within the scope of the present invention as defined in the claims. It should be understood that the above is an explanation of the present invention and should not be considered as being limited to the specific embodiments disclosed, and modifications to the disclosed embodiments and other embodiments are intended to be included within the scope of the appended claims. The present invention is defined by the claims and their equivalents.

Claims

1. A method for training a neural network model, include: Determine a depth and a width of the first neural network model as a first depth and a first width, respectively, wherein the values ​​of the first depth and the first width are greater than 1, the first depth indicates the number of multiple identical network modules cascaded in the first neural network model, and the first width indicates the number of multiple identical network modules connected in parallel in the first neural network model; Determine a first random number and a second random number, wherein the first random number is less than or equal to the first depth, and the second random number is less than or equal to the first width; In a predetermined order, a portion of the first neural network model having a second depth and a second width is selected as a second neural network model, and a neural network model other than the second neural network model in the first neural network model is selected as a third neural network model; wherein the second depth is equal to the first random number, and the second width is equal to the second random number; Obtaining samples and sample labels for training the first neural network model; Processing the sample using the first neural network model; and The parameters of the second neural network model are updated based on the processing results of the second neural network model on the sample and the sample label, and the parameters of the third neural network model are updated based on the processing results of the first neural network model on the sample and the sample label to obtain the first neural network model with updated parameters.

2. The method according to claim 1, in, Selecting, in a predetermined order, a portion of the first neural network model having a second depth and a second width as the second neural network model comprises: Numbering a plurality of identical network modules cascaded in the first neural network model in a cascade order to obtain a first sequence number, and numbering a plurality of identical network modules connected in parallel in the first neural network model in a parallel order to obtain a second sequence number; The part of the first neural network model in which the first serial number is less than or equal to the first random number and the second serial number is less than or equal to the second random number is selected as the second neural network model.

3. The method according to claim 1, in, The first neural network model includes a plurality of cascaded first sub-networks, the number of the plurality of first sub-networks is equal to the first depth, wherein each of the plurality of first sub-networks includes a plurality of parallel network modules whose number is equal to the first width, wherein each of the network modules is initialized based on different parameters; The second neural network model includes a plurality of cascaded second sub-networks, the number of the plurality of second sub-networks being equal to the second depth, wherein each of the plurality of second sub-networks includes a plurality of parallel network modules whose number is equal to the second width.

4. The method according to claim 3, in, At least one of the first sub-network and the second sub-network also includes at least one of a pre-processing module and a post-processing module that are different from the network module, the pre-processing module includes: at least one of a feature extraction network and a dimensionality transformation network, and the post-processing module includes: at least one of a dimensionality transformation network and a residual network.

5. The method according to claim 3, in, Processing the sample using the first neural network model includes: For any sub-network of the first sub-network and the second sub-network, use the plurality of network modules in the sub-network to process the information corresponding to the sample respectively to obtain a plurality of network module processing results, and perform weighted summation on the plurality of network module processing results to obtain the processing result of the sub-network; and The processing result of the last first sub-network in the first neural network model is used as the processing result of the first neural network model on the sample, and the processing result of the last second sub-network in the second neural network model is used as the processing result of the second neural network model on the sample.

6. The method according to claim 5, in, Performing weighted summation on the processing results of the multiple network modules includes: Each of the network module processing results is divided into two processing sub-results, wherein the first of the two processing sub-results is used to determine the weight corresponding to the network module, and the second of the two network module processing sub-results is used to multiply the weight to obtain the result of the sub-network processing the sample.

7. The method according to claim 3, in, Determining the first random number and the second random number includes: Determining a first random number, wherein the first random number is used to determine the number of first sub-networks to be selected in the first neural network model; and For each first sub-network to be selected, generating a second random number, wherein the second random number is used to determine the number of network modules connected in parallel in the first sub-network to be selected; Wherein, selecting, in a predetermined order, a portion having a second depth and a second width in the first neural network model as the second neural network model comprises: Among the plurality of first sub-networks in the cascade, a first sub-network whose cascade level is less than or equal to the second depth is selected, and the second depth is equal to the first random number; For each selected first sub-network, a network module whose parallel connection order is before the second width in the sub-network is selected, and the second width is equal to the second random number generated for the first sub-network.

8. The method according to claim 1, in, Updating the parameters of the third neural network model based on the processing result of the first neural network model on the sample and the sample label includes: calculating the value of a first loss function based on the processing result of the first neural network model on the sample and the sample label, and updating the parameters of the third neural network model based on the value of the first loss function; Updating the parameters of the second neural network model based on the processing results of the samples by the second neural network model and the sample labels includes: calculating the value of the second loss function based on the processing results of the samples by the second neural network model and the sample labels, and updating the parameters of the second neural network model based on the value of the second loss function.

9. A method for information processing based on a neural network model in an electronic device, include: Get the information to be processed; Obtaining a trained first neural network model; In a predetermined order, according to the processing capability of the electronic device, a portion of the first neural network model having a third depth and a third width is selected as an information processing neural network model, wherein the first neural network model has a first depth and a first width, the values ​​of the first depth and the first width are greater than 1, the first depth indicates the number of multiple identical network modules cascaded in the first neural network model, the first width indicates the number of multiple identical network modules connected in parallel in the first neural network model, the third depth is less than or equal to the first depth, and the third width is less than or equal to the first width; The information to be processed is processed based on the information processing neural network model to obtain an information processing result for the information to be processed.

10. The method according to claim 9, in, Selecting a portion having a third depth and a third width in the first neural network model as an information processing neural network model according to the processing capability of the electronic device includes: Numbering a plurality of identical network modules cascaded in the first neural network model in a cascade order to obtain a first sequence number, and numbering a plurality of identical network modules connected in parallel in the first neural network model in a parallel order to obtain a second sequence number; The part of the first neural network model in which the first serial number is less than or equal to the third depth and the second serial number is less than or equal to the third width is selected as the information processing neural network model.

11. The method according to claim 9, in, The information processing neural network model includes a plurality of cascaded third sub-networks, the number of the plurality of third sub-networks being equal to the third depth, wherein each of the plurality of third sub-networks includes a plurality of parallel network modules whose number is equal to the third width.

12. The method according to claim 11, in, Processing the information to be processed based on the information processing neural network model includes: For each of the third sub-networks, use the plurality of network modules in the third sub-network to process the information to be processed respectively, so as to obtain a plurality of network module processing results; Performing weighted summation on the processing results of the multiple network modules to obtain a processing result of the third sub-network; and The processing result of the last third sub-network in the information processing neural network model is used as the information processing result for the information to be processed.

13. A device for training a neural network model, include: A dimension determination unit is configured to: determine a depth and a width of the first neural network model as a first depth and a first width, respectively, wherein the values ​​of the first depth and the first width are greater than 1, the first depth indicates the number of multiple identical network modules cascaded in the first neural network model, and the first width indicates the number of multiple identical network modules in parallel in the first neural network model; A random number determination unit is configured to: determine a first random number and a second random number, wherein the first random number is less than or equal to the first depth, and the second random number is less than or equal to the first width; A model selection unit is configured to: select, in a predetermined order, a portion of the first neural network model having a second depth and a second width as a second neural network model, and select a neural network model other than the second neural network model in the first neural network model as a third neural network model; wherein the second depth is equal to the first random number, and the second width is equal to the second random number; A sample information acquisition unit is configured to: acquire samples and sample labels for training the first neural network model; A sample processing unit, configured to: process the sample using the first neural network model; and The model parameter updating unit is configured to update the parameters of the second neural network model based on the processing result of the second neural network model on the sample and the sample label, and to update the parameters of the third neural network model based on the processing result of the first neural network model on the sample and the sample label.

14. A device for performing information processing based on a neural network model in an electronic device, include: The information acquisition unit is configured to: acquire information to be processed; The model acquisition unit is configured to: acquire a trained first neural network model; A model selection unit is configured to: select, in a predetermined order and according to the processing capability of the electronic device, a portion of the first neural network model having a third depth and a third width as an information processing neural network model, wherein the first neural network model has a first depth and a first width, the values ​​of the first depth and the first width are greater than 1, the first depth indicates the number of multiple identical network modules cascaded in the first neural network model, the first width indicates the number of multiple identical network modules connected in parallel in the first neural network model, the third depth is less than or equal to the first depth, and the third width is less than or equal to the first width; The result acquisition unit is configured to: process the information to be processed based on the information processing neural network model to obtain an information processing result for the information to be processed.

15. A computer program product, comprising computer software code, which is used to implement the method according to any one of claims 1 to 12 when executed by a processor.

16. A computer-readable storage medium having computer-executable instructions stored thereon, wherein the instructions are used to implement the method according to any one of claims 1 to 12 when executed by a processor.