Neural Network-Based Information Processing Method and Device

By using meta-neural networks to predict the parameter probability distribution in parameter memory in complex decision-making systems, dynamically adjusting the parameter combination mode of the main neural network, the problem that a single neural network cannot adapt to multiple strategy combinations is solved, and better complex task processing effects are achieved.

CN110689117BActive Publication Date: 2025-07-01BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201910926738.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-09-27
Publication Date
2025-07-01
Estimated Expiration
2039-09-27

AI Technical Summary

Technical Problem

In complex decision-making systems, using a single neural network with fixed parameters cannot effectively learn a variety of different combinations of strategies, especially when input information changes, policy adjustment cannot be achieved.

Method used

The meta-neural network is used to predict the probability distribution of each parameter in the parameter memory, and the parameter memory is constructed through the full connection layer of the main neural network, and the parameter combination mode of the main neural network is dynamically adjusted to adapt to different input information.

Benefits of technology

Automatic dynamic adjustment of main neural network parameters is realized, and different parameter combination modes can be used for different input information, which improves the processing effect of neural networks in complex tasks and solves the problem of poor expression ability of a single strategy model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN110689117B_ABST
    Figure CN110689117B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of computer data processing. Embodiments of the present disclosure disclose an information processing method and apparatus, an electronic device, and a computer-readable medium based on a neural network. The information processing method based on the neural network includes: obtaining input information; based on the input information, using a meta neural network to predict the probability distribution of each parameter in a parameter memory, where the parameter memory is pre-constructed based on the fully connected layer of a main neural network; based on the probability distribution of each parameter in the parameter memory, determining a parameter combination pattern of the fully connected layer of the main neural network with respect to the input information; based on the parameter combination pattern of the fully connected layer with respect to the input information, updating the fully connected layer of the main neural network, and processing the input information based on the main neural network after updating the fully connected layer to obtain output information corresponding to the input information. This method realizes the automatic dynamic adjustment of the parameters of the main neural network based on the input information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technologies, specifically to the field of computer data processing technologies, and particularly to an information processing method and apparatus based on neural networks. Background Art

[0002] The goals of complex decision-making systems usually need to be accomplished by multiple strategies. Neural networks can be used to model complex decision-making systems to represent the non-linear relationships between their subsystems. However, complex decision-making systems usually contain multiple different strategies. If only one neural network with fixed parameters is used for modeling, it is impossible to learn multiple different parameter combination methods in the same network. That is to say, the input of a complex decision-making system is usually variable, and it is impossible to adjust the strategy in the case of different inputs using the same neural network.

[0003] The current solution is to use a Mixture of Experts (MOE), set different expert networks to model different strategies, and then use the same gating matrix to select from different expert networks. Summary of the Invention

[0004] Embodiments of the present disclosure propose an information processing method and apparatus based on neural networks, an electronic device, and a computer-readable medium.

[0005] In a first aspect, an embodiment of the present disclosure provides an information processing method based on neural networks, including: obtaining input information; based on the input information, using a meta neural network to predict the probability distribution of each parameter in a parameter memory, where the parameter memory is pre-constructed based on the fully connected layer of a main neural network; based on the probability distribution of each parameter in the parameter memory, determining a parameter combination mode of the fully connected layer of the main neural network with respect to the input information; based on the parameter combination mode of the fully connected layer with respect to the input information, updating the fully connected layer of the main neural network, and processing the input information based on the main neural network after updating the fully connected layer to obtain output information corresponding to the input information.

[0006] In some embodiments, the above method further includes: using the feature extraction layer of the main neural network to extract features from the input information to obtain an abstract representation of the input information; and the above using a meta neural network to predict the probability distribution of each parameter in the parameter memory based on the input information includes: using a meta neural network to predict the probability distribution of each parameter in the parameter memory based on the abstract representation of the input information; the above processing the input information based on the main neural network after updating the fully connected layer includes: processing the abstract representation of the input information based on the updated fully connected layer.

[0007] In some embodiments, for the above-mentioned abstract representation based on input information, a meta-neural network is used to predict the probability distribution of each parameter in the parameter memory, including: using the abstract representation of the input information as the dynamic parameter of the meta-neural network, and using the meta-neural network containing the dynamic parameter to predict the probability distribution of each parameter in the parameter memory.

[0008] In some embodiments, for the above-mentioned determination of the parameter combination pattern of the fully connected layer of the main neural network with respect to the input information based on the probability distribution of each parameter in the parameter memory, it includes: using the probability value in the probability distribution as the weight coefficient represented by the corresponding parameter in the fully connected layer, and performing a non-linear transformation after weighted summation of each node in the previous layer of the fully connected layer of the main neural network based on the weight coefficient, to obtain the parameter combination pattern of the fully connected layer of the main neural network with respect to the input information.

[0009] In some embodiments, the above-mentioned main neural network is pre-trained based on sample information of the same type as the input information.

[0010] In a second aspect, an embodiment of the present disclosure provides an information processing device based on a neural network, including: an acquisition unit configured to acquire input information; a prediction unit configured to, based on the input information, use a meta-neural network to predict the probability distribution of each parameter in a parameter memory, where the parameter memory is pre-constructed based on the fully connected layer of the main neural network; a determination unit configured to, based on the probability distribution of each parameter in the parameter memory, determine the parameter combination pattern of the fully connected layer of the main neural network with respect to the input information; and a processing unit configured to, based on the parameter combination pattern of the fully connected layer with respect to the input information, update the fully connected layer of the main neural network, and process the input information based on the main neural network after updating the fully connected layer, to obtain output information corresponding to the input information.

[0011] In some embodiments, the above-mentioned device further includes: an extraction unit configured to use the feature extraction layer of the main neural network to perform feature extraction on the input information to obtain an abstract representation of the input information; and the above-mentioned prediction unit is further configured to: based on the abstract representation of the input information, use a meta-neural network to predict the probability distribution of each parameter in the parameter memory; the above-mentioned processing unit is further configured to: process the abstract representation of the input information based on the updated fully connected layer.

[0012] In some embodiments, the above-mentioned prediction unit is further configured to: use the abstract representation of the input information as the dynamic parameter of the meta-neural network, and use the meta-neural network containing the dynamic parameter to predict the probability distribution of each parameter in the parameter memory.

[0013] In some embodiments, the above-mentioned determination unit is further configured to determine the parameter combination pattern of the fully connected layer of the main neural network with respect to the input information in the following manner: taking the probability values in the probability distribution as the weight coefficients represented by the corresponding parameters in the fully connected layer, and performing a non-linear transformation after performing a weighted sum of the nodes in the previous layer of the fully connected layer of the main neural network based on the weight coefficients, so as to obtain the parameter combination pattern of the fully connected layer of the main neural network with respect to the input information.

[0014] In some embodiments, the above-mentioned main neural network is pre-trained based on sample information of the same type as the input information.

[0015] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the neural network-based information processing method provided in the first aspect.

[0016] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable medium, on which a computer program is stored, wherein when the program is executed by a processor, it implements the neural network-based information processing method provided in the first aspect.

[0017] The neural network-based information processing method and apparatus, electronic device, and computer-readable medium of the above embodiments of the present disclosure obtain input information, and then, based on the input information, use a meta-neural network to predict the probability distribution of each parameter in the parameter memory, where the parameter memory is pre-constructed based on the fully connected layer of the main neural network. Then, based on the probability distribution of each parameter in the parameter memory, determine the parameter combination pattern of the fully connected layer of the main neural network with respect to the input information, adjust the parameters of the fully connected layer of the main neural network according to the parameter combination pattern, and then process the input information based on the main neural network after adjusting the parameters of the fully connected layer to obtain the output information corresponding to the input information, realizing the automatic dynamic adjustment of the parameters of the main neural network based on the input information, and can use different parameter combination patterns for different inputs to perform the information processing task of the neural network, solving the problem of poor model expression ability of a single strategy, and achieving good processing effects for complex tasks. Description of the Drawings

[0018] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present disclosure will become more apparent:

[0019] Figure 1 It is an exemplary system architecture diagram to which the embodiments of the present disclosure can be applied;

[0020] Figure 2It is a flowchart of an embodiment of the neural network-based information processing method according to the present disclosure;

[0021] Figure 3 It is a flowchart of another embodiment of the neural network-based information processing method according to the present disclosure;

[0022] Figure 4 It is a schematic diagram of the algorithm principle used in the neural network-based information processing method according to the present disclosure;

[0023] Figure 5 It is a schematic structural diagram of an embodiment of the neural network-based information processing apparatus according to the present disclosure;

[0024] Figure 6 It is a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present disclosure. Detailed implementation manners

[0025] The present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the related invention, rather than limiting the invention. Additionally, it should be noted that for the sake of description, only parts related to the relevant invention are shown in the drawings.

[0026] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The present disclosure will be described in detail below with reference to the drawings and embodiments.

[0027] Figure 1 An exemplary system architecture 100 to which the neural network-based information processing method or the neural network-based information processing apparatus according to the present disclosure can be applied is shown.

[0028] As Figure 1 shown, the system architecture 100 may include, as Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium to provide a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0029] The terminal devices 101, 102, 103 interact with the server 105 through the network 104 to receive or send messages, etc. Various client applications may be installed on the terminal devices 101, 102, 103. For example, image processing applications, information analysis applications, voice assistant applications, shopping applications, financial applications, etc.

[0030] The terminal devices 101, 102, and 103 can be hardware or software. When the terminal devices 101, 102, and 103 are hardware, they can be various electronic devices, including but not limited to smartphones, tablet computers, e-book readers, laptop computers, desktop computers, and so on. When the terminal devices 101, 102, and 103 are software, they can be installed in the above-listed electronic devices. They can be implemented as multiple software or software modules (such as multiple software or software modules for providing distributed services), or can be implemented as a single software or software module. Specific limitations are not made here.

[0031] The server 105 can be a server that provides various services, such as a backend server that provides backend support for applications installed on the terminal devices 101, 102, and 103. For example, the server 105 can receive the information to be processed sent by the terminal devices 101, 102, and 103, perform information processing tasks using a neural network-based model, and return the processing results to the terminal devices 101, 102, and 103.

[0032] In some specific examples, the terminal devices 101, 102, and 103 can send information processing requests related to tasks such as speech recognition, text classification, dialogue behavior classification, and image recognition to the server 105. A neural network model trained for the corresponding tasks can run on the server 105, and this neural network model is used to process the information.

[0033] It should be noted that the neural network-based information processing method provided by the embodiments of the present disclosure is generally executed by the server 105. Correspondingly, the neural network-based information processing device is generally set in the server 105.

[0034] It should also be pointed out that in some scenarios, the server 105 can obtain the information to be processed from a database, a memory, or other devices. At this time, the exemplary system architecture 100 may not have the terminal devices 101, 102, and 103 and the network 104.

[0035] It should be noted that the server 105 can be hardware or software. When the server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or can be implemented as a single server. When the server 105 is software, it can be implemented as multiple software or software modules (such as multiple software or software modules for providing distributed services), or can be implemented as a single software or software module. Specific limitations are not made here.

[0036] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0037] Continue to refer to Figure 2 , which shows a process 200 of an embodiment of a neural network-based information processing method according to the present disclosure. The neural network-based information processing method includes the following steps:

[0038] Step 201, obtain input information.

[0039] In this embodiment, the execution subject of the above neural network-based information processing method may first obtain the information to be processed from an electronic device that issues an information processing request as the input information, or may read the information to be processed from a database or a memory as the input information. The input information is the information to be input into the main neural network and used by the main neural network to execute a processing task, and may include data such as text, images, or audio and video.

[0040] In practice, the above execution subject may be a device for executing a specified information processing task. Here, the specified information processing task may be, for example, tasks that can be modeled using a neural network model such as speech recognition, speech synthesis, semantic analysis, image recognition, target tracking, behavior intention analysis, trend prediction, etc. The input information may be the processing object of the specified information processing task. For example, in a speech recognition task, the input information is a speech signal, and in an image recognition task, the input information is an image.

[0041] Step 202, based on the input information, use a meta neural network to predict the probability distribution of each parameter in the parameter memory.

[0042] Among them, the parameter memory is pre-constructed based on the fully connected layer of the main neural network. The main neural network is a multi-layer neural network for executing information processing tasks, and may include a feature extraction layer and a fully connected layer. The feature extraction layer is used to extract the features of the data input into the main neural network, and the fully connected layer integrates the features extracted by the feature extraction layer according to a certain strategy, so as to integrate the features of different dimensions extracted.

[0043] The above execution subject may pre-construct a parameter memory based on the parameters of the fully connected layer of the main neural network. The parameter memory includes all the parameters of the fully connected layer of the main neural network. In some alternative implementation manners, the parameter memory may be a two-dimensional parameter matrix, and each element in the matrix corresponds to a parameter of the fully connected layer.

[0044] Based on the input information, a meta neural network outside the main neural network may be used to predict the probability distribution of each parameter in the fully connected layer of the main neural network, that is, use the meta neural network to predict the values of the parameters corresponding to the current input information in the fully connected layer of the main neural network.

[0045] In this embodiment, the meta neural network can preprocess the input information. For example, after extracting features, the meta neural network outputs the probability distribution of each parameter in the parameter memory after preprocessing the input information.

[0046] The meta neural network can be pre-trained. When training the meta neural network, the fully connected layer parameters dynamically adjusted for different input information can be obtained to construct a sample set, and the meta neural network is trained based on the sample set using a supervised or unsupervised learning method.

[0047] Step 203: Based on the probability distribution of each parameter in the parameter memory, determine the parameter combination mode of the fully connected layer of the main neural network with respect to the input information.

[0048] The parameters of the fully connected layer are the weights of each node in its previous layer. Each node in the fully connected layer integrates the data of the previous layer by weighted summation of multiple nodes in the previous layer. Usually, the number of nodes in the previous layer of the fully connected layer and the number of nodes in the fully connected layer are relatively large, so the number of parameters of the fully connected layer is relatively large. In this embodiment, the values of the parameters of the fully connected layer (i.e., the weights corresponding to each node in the previous layer of the fully connected layer) change dynamically with the input information. Thus, different combination modes of the parameters of the fully connected layer can represent different strategy modes adopted by the main neural network when further processing the features extracted by the feature extraction layer to obtain the final result. Here, the parameter combination mode refers to the combination of each node in the previous layer of the fully connected layer and its corresponding weights in the fully connected layer. Specifically, the previous layer of the fully connected layer includes m nodes x1, x2,..., x m , and the fully connected layer includes n nodes y1, y2,..., y n , then the fully connected layer includes at least m×n parameters, expressed as a two-dimensional parameter matrix W m×n , and the parameter combination mode is different values of the two-dimensional parameter matrix W m×n combined with the corresponding previous layer nodes and this layer nodes.

[0049] For example, taking the previous layer of the fully connected layer containing 3 nodes x1, x2, x3 and the fully connected layer including two nodes y1, y2 as an example, the parameters of the fully connected layer are expressed as w 11 , w 12 , w 13 , w 21 , w 22 , w 23 , and the parameter combination mode of the fully connected layer is: y1 = w 11 × x1 + w 12 × x2 + w 13 × x3, y2 = w 21 × x1 + w 22 × x2 + w 23 × x3.

[0050] In this embodiment, the probability values corresponding to the parameters in the probability distribution output by the meta-neural network can be used as the values of the parameters in the parameter memory, and the nodes corresponding to the parameters are combined to generate the parameter combination pattern of the fully connected layer for the current input information. That is, according to the predicted probability distribution, the parameters in the above parameter matrix W m×n are assigned values, and Y = W m×n X is used as the parameter combination pattern for the input information, where Y = {y1, y2,..., y n}, and X = {x1, x2,..., x m}.

[0051] In some alternative implementation manners of this embodiment, the probability values in the probability distribution can be used as the weight coefficients represented by the corresponding parameters in the fully connected layer. After performing weighted summation on each node in the layer above the fully connected layer of the main neural network based on the weight coefficients and then performing a non-linear transformation, the parameter combination pattern of the fully connected layer of the main neural network for the input information is obtained.

[0052] Specifically, after using the probability values corresponding to the parameters in the probability distribution output by the meta-neural network as the values of the parameters (i.e., the weight coefficients of each layer of the fully connected layer) in the parameter memory and performing weighted summation on the nodes in the layer above the fully connected layer using the corresponding weight coefficients, a non-linear transformation can also be performed. For example, activation functions such as sigmoid, tanh, and Relu are used for non-linear transformation, so that the output of the fully connected layer is differentiable, which can ensure that the output of the fully connected layer can be differentiated, thereby ensuring that the main neural network can use the gradient descent method for iterative training.

[0053] Step 204: Based on the parameter combination pattern of the fully connected layer for the input information, update the fully connected layer of the main neural network, and process the input information based on the main neural network after updating the fully connected layer to obtain the output information corresponding to the input information.

[0054] The parameter combination pattern of the fully connected layer of the main neural network determined in step 203 can be used as the policy pattern of the updated fully connected layer. The output of the feature extraction layer is processed using the policy pattern of the updated fully connected layer to obtain the integrated result of the features. Then, the integrated result of the features can be classified in the main neural network, and the probabilities of belonging to each predetermined category are calculated to obtain the output information corresponding to the input information; or regression can be performed on the integrated result of the features in the main neural network to obtain the output information corresponding to the input information.

[0055] Since the input information obtained is utilized when determining the probability distributions of the various parameters in the parameter memory, the above-described embodiments of the present disclosure can generate different probability distributions of the fully connected layer parameters based on different input information, thereby achieving dynamic updating of the parameters of the main neural network as the input information changes. By dynamically adjusting the parameters of the main neural network based on the input information, the main neural network can be made more adaptable to changes in the input information, improving the processing effect of the main neural network in processing the input information. Moreover, the above information processing method for dynamically adjusting parameters according to the input information enables the main neural network to provide an information processing mode with multiple policy logics, thus solving the problem of poor model representation ability of a single policy and achieving a good processing effect for complex tasks.

[0056] In some alternative implementation manners of this embodiment, the above main neural network can be pre-trained based on sample information of the same type as the input information. First, the task type executed by the main neural network can be determined, such as speech recognition or image recognition, and then the data type of the corresponding input information and the data type of the output information can be determined according to the task type. For example, in a speech recognition task, the input information is speech data and the output information is text data, then sample data of the speech type can be obtained to construct a sample set for training the main neural network. It is also possible to obtain text-type annotation information corresponding to the sample data of the speech type as the annotation information of the sample set. In this way, after training is completed, the main neural network can perform information processing on input information of the same type as the sample information input during the training process.

[0057] In some other alternative implementation manners of this embodiment, the above main neural network can also be pre-trained using sample information of a different type from the input information. In this case, the parameters of the main neural network can be dynamically adjusted through the above-described information processing method based on a neural network to make it adaptable to the processing of input information of a type different from the type of sample information used during the training process.

[0058] Continue to refer to Figure 3 , which shows a flowchart of another embodiment of the information processing method based on a neural network according to the present disclosure. As Figure 3 shown, the process 300 of the information processing method based on a neural network in this embodiment includes the following steps:

[0059] Step 301, obtain input information.

[0060] Step 301 of this embodiment is the same as step 201 of the foregoing embodiment. The specific implementation manner of step 301 can refer to the description of the foregoing step 201 and will not be elaborated here.

[0061] Step 302: Use the feature extraction layer of the main neural network to extract features from the input information to obtain an abstract representation of the input information.

[0062] Before using the meta neural network to predict the probability distribution of each parameter in the parameter memory, the input information can be first subjected to feature extraction to convert the input information into a corresponding abstract representation.

[0063] The main neural network can be, for example, a convolutional neural network, where the feature extraction layer can include at least one of a convolutional layer, a deconvolutional layer, and a pooling layer. In practice, the feature extraction layer usually includes multiple layers, each layer including multiple nodes. The outputs of each layer in the feature extraction layer can be combined as the abstract representation of the input information, or the output of the last layer in the feature extraction layer can be used as the abstract representation of the input information.

[0064] Step 303: Based on the abstract representation of the input information, use the meta neural network to predict the probability distribution of each parameter in the parameter memory.

[0065] In this embodiment, the meta neural network can predict the probability distribution of each parameter in the parameter memory based on the abstract representation of the input information obtained in step 302. Among them, the parameter memory can be pre-constructed based on the parameters of the fully connected layer of the main neural network. The parameter memory includes all the parameters of the fully connected layer of the main neural network. In some alternative implementation manners, the parameter memory can be a two-dimensional parameter matrix, and each element in the matrix corresponds to a parameter of the fully connected layer.

[0066] Optionally, the abstract representation of the input information can be used as the dynamic parameter of the meta neural network, and the meta neural network including the dynamic parameter is used to predict the probability distribution of each parameter in the parameter memory.

[0067] Represent the input information as x i , the abstract representation of the input information extracted by the feature extraction layer is F(x i ), and the fully connected layer Z i in the main neural network can be expressed as:

[0068] Z i = W T F(x i ) + b (1)

[0069] where W ∈ R H×C represents the parameters of the fully connected layer, H×C represents the size of the parameter space to which the parameters of the fully connected layer belong, and b represents the bias of the fully connected layer.

[0070] The meta neural network G W (F((x i ), θ G) to predict the probability distribution of the parameters in the fully connected layer of the main neural network, θ G represents the parameters of the meta neural network, and the meta neural network G W (F(x i ), θ G ) The predicted probability distribution can be used as the parameter value of each corresponding parameter in the fully connected layer. Then, formula (1) is rewritten as:

[0071] Z i = G W (F(x i ), θ G ) T F(x i ) + b (2)

[0072] The parameter memory can be represented as a two-dimensional matrix M ∈ R K× (H × C), where K×(H×C) represents the size of the two-dimensional matrix. Then, the meta neural network can be represented as:

[0073] G W (F(x i ), θ G ) = ψ(M T Q(F(x i ), θ Q )) (3)

[0074] where ψ is a non-linear transformation function, such as activation functions identity, Relu, sigmoid, tanh, etc.; Q(F(x i ), θ Q ) is a neural network model, where θ Q represents the parameters of the neural network model, and the neural network model can be a multi-layer perceptron model, a convolutional neural network model, or a recurrent neural network model.

[0075] The probability distribution of each parameter in the parameter matrix M can be calculated using the above formula (3).

[0076] Step 304, based on the probability distribution of each parameter in the parameter memory, determine the parameter combination pattern of the fully connected layer of the main neural network with respect to the input information.

[0077] Step 304 of this embodiment is the same as step 203 of the foregoing embodiment. The specific implementation manner of step 304 can refer to the description of step 203 above and will not be elaborated here.

[0078] Step 305: Update the fully connected layer of the main neural network based on the parameter combination pattern of the fully connected layer with respect to the input information, and process the abstract representation of the input information based on the updated fully connected layer to obtain the output information corresponding to the input information.

[0079] After determining the parameter combination pattern of the fully connected layer of the main neural network with respect to the input information, this parameter combination pattern can be used as the parameter combination pattern of the updated fully connected layer, or in other words, as the strategy pattern for the updated fully connected layer to calculate the output of the feature extraction layer. Use the strategy pattern of the updated fully connected layer to calculate the abstract representation of the input information obtained in step 302. Specifically, the probability distribution of the parameters obtained in step 303 can be used as the value of the corresponding weight parameters of the fully connected layer. After performing weighted summation on the abstract representation of the input information based on the weight parameters, perform a non-linear transformation, such as using functions like Relu, softmax, sigmoid, etc. to perform the non-linear transformation, and classify or regress the result of the non-linear transformation to obtain the processing result of the input information, that is, obtain the output information corresponding to the input information.

[0080] In this embodiment, feature extraction is performed on the input information to obtain the abstract representation of the input information, which can extract the effective features of the input information, enabling the meta-neural network to efficiently obtain relatively accurate prediction results when using the abstract representation of the input information to predict the probability distribution of the parameters in the parameter memory. Since the meta-neural network does not need to perform feature extraction on the input information, it effectively improves the prediction speed of the meta-neural network, thereby improving the dynamic update speed of the parameters in the fully connected layer of the main neural network, and further improving the information processing efficiency.

[0081] Continue to refer to Figure 4 which shows a schematic diagram of the algorithm principle used in the information processing method based on a neural network according to the present disclosure. As Figure 4 shown, in step (1), first convert the input information x into the corresponding abstract representation F(x) in the main neural network; then in step (2), transmit the abstract representation F(x) of the input information x to the meta-neural network G, use the multi-layer perceptron Q to predict the probability distribution of each parameter in the parameter matrix M, and integrate it with the parameter matrix M to obtain the parameter combination pattern in the main neural network, where the multi-layer perceptron Q uses the abstract representation F(x) of the input information x as the dynamic parameter; then in step (3), apply the parameter combination pattern to the fully connected layer of the main neural network to obtain the parameter W(x) of the fully connected layer; in step (4), further process the abstraction F(x) of the input information based on the fully connected layer to obtain the output information L y .

[0082] From Figure 4It can be seen that in the method of the above embodiments of the present disclosure, a meta-neural network and a parameter memory are added on the basis of the main neural network to achieve dynamic adjustment of the parameters of the trained main neural network. The implementation method is relatively simple and will not increase excessive computational complexity. This method can dynamically set various logical modes according to input information, and can solve the problem that a single model cannot express the logical modes of complex systems.

[0083] Further referring to Figure 5 , as an implementation of the above neural network-based information processing method, the present disclosure provides an embodiment of a neural network-based information processing device. This device embodiment corresponds to the Figure 2 and Figure 3 shown method embodiments, and this device can be specifically applied to various electronic devices.

[0084] As Figure 5 shown, the neural network-based information processing device 500 of this embodiment includes: an acquisition unit 501, a prediction unit 502, a determination unit 503, and a processing unit 504. Among them, the acquisition unit 501 is configured to acquire input information; the prediction unit 502 is configured to predict the probability distribution of each parameter in the parameter memory based on the input information, where the parameter memory is pre-constructed based on the fully connected layer of the main neural network; the determination unit 503 is configured to determine the parameter combination mode of the fully connected layer of the main neural network with respect to the input information based on the probability distribution of each parameter in the parameter memory; the processing unit 504 is configured to update the fully connected layer of the main neural network based on the parameter combination mode of the fully connected layer with respect to the input information, and process the input information based on the main neural network after updating the fully connected layer to obtain the output information corresponding to the input information.

[0085] In some embodiments, the above device further includes: an extraction unit configured to extract features of the input information using the feature extraction layer of the main neural network to obtain an abstract representation of the input information; and the above prediction unit 502 is further configured to: predict the probability distribution of each parameter in the parameter memory based on the abstract representation of the input information; the above processing unit 504 is further configured to: process the abstract representation of the input information based on the updated fully connected layer.

[0086] In some embodiments, the above prediction unit 502 is further configured to: use the abstract representation of the input information as the dynamic parameter of the meta-neural network, and predict the probability distribution of each parameter in the parameter memory using the meta-neural network including the dynamic parameter.

[0087] In some embodiments, the determining unit 503 is further configured to determine the parameter combination pattern of the fully connected layer of the main neural network with respect to the input information in the following manner: using the probability values in the probability distribution as the weight coefficients represented by the corresponding parameters in the fully connected layer, performing a weighted sum on each node in the previous layer of the fully connected layer of the main neural network based on the weight coefficients and then performing a non-linear transformation to obtain the parameter combination pattern of the fully connected layer of the main neural network with respect to the input information.

[0088] In some embodiments, the main neural network is pre-trained based on sample information of the same type as the input information.

[0089] It should be understood that the various units described in the apparatus 500 correspond to the respective steps in the method described with reference Figure 2 and Figure 3 Thus, the operations and features described above for the neural network-based information processing method also apply to the apparatus 500 and the units included therein, and will not be repeated here.

[0090] The neural network-based information processing apparatus 500 according to the above embodiments of the present disclosure obtains input information through an obtaining unit, and then a prediction unit predicts the probability distribution of each parameter in a parameter memory based on the input information, where the parameter memory is pre-constructed based on the fully connected layer of the main neural network. Then, a determining unit determines the parameter combination pattern of the fully connected layer of the main neural network with respect to the input information based on the probability distribution of each parameter in the parameter memory, adjusts the parameters of the fully connected layer of the main neural network according to the parameter combination pattern, and then a processing unit processes the input information based on the main neural network after adjusting the parameters of the fully connected layer to obtain output information corresponding to the input information, achieving dynamic adjustment of the parameters of the main neural network based on the input information, and can use different parameter combination patterns for different inputs to perform the information processing task of the neural network, solving the problem of poor model expression ability of a single strategy, and achieving good processing effects for complex tasks.

[0091] Next, reference is made to Figure 6 , which shows a schematic structural diagram of an electronic device (such as Figure 1 the server shown) 600 suitable for implementing the embodiments of the present disclosure. Figure 6 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0092] As Figure 6As shown, the electronic device 600 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage device 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.

[0093] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 6 the electronic device 600 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had. Figure 6 Each block shown in it may represent a device or, as needed, multiple devices.

[0094] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present disclosure include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by a processing device 601, the above-described functions defined in the methods of the embodiments of the present disclosure are performed. It should be noted that the computer-readable medium described in the embodiments of the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the embodiments of the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the embodiments of the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0095] The above computer-readable medium may be included in the above electronic device; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to: obtain input information; based on the input information, adopt a meta-neural network to predict the probability distribution of each parameter in the parameter memory, where the parameter memory is pre-constructed based on the fully-connected layer of the main neural network; based on the probability distribution of each parameter in the parameter memory, determine the parameter combination pattern of the fully-connected layer of the main neural network with respect to the input information; based on the parameter combination pattern of the fully-connected layer with respect to the input information, update the fully-connected layer of the main neural network, and based on the main neural network after updating the fully-connected layer, process the input information to obtain output information corresponding to the input information.

[0096] Computer program code for performing the operations of the embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, execute as a stand-alone software package, execute partially on the user's computer and partially on a remote computer, or execute entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0097] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0098] The units involved in the embodiments described in this disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: A processor includes an acquisition unit, a prediction unit, a determination unit, and a processing unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases. For example, the acquisition unit can also be described as "the unit for acquiring input information".

[0099] The above description is only a preferred embodiment of this disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in this disclosure is not limited to the technical solution formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in this application that have similar functions.

Claims

1. An information processing method based on a neural network, comprising: Obtaining input information, where the input information is information to be input into a main neural network for performing a processing task using the main neural network. The input information includes: text, image, or audio - video. The processing task includes: speech recognition, text classification, image recognition. The main neural network is pre - trained based on sample information of the same type as the input information, and the main neural network is used to process input information of the same type as the sample information input during the training process; Based on the input information, using a meta - neural network to predict the values of the parameters of each full - connection layer of the main neural network corresponding to the input information, obtaining the probability distribution of each parameter in the parameter memory, where the parameter memory is pre - constructed based on the full - connection layer of the main neural network; the meta - neural network is a network pre - trained outside the main neural network. When training the meta - neural network, the full - connection layer parameters of the main neural network dynamically adjusted for different input information such as text, image, and audio - video are obtained to construct a sample set; Based on the probability distribution of each parameter in the parameter memory, determining the parameter combination pattern of the full - connection layer of the main neural network with respect to the input information; Based on the parameter combination pattern of the full - connection layer with respect to the input information, updating the full - connection layer of the main neural network to provide an information processing mode with multi - strategy logic for different processing tasks such as speech recognition, text classification, and image recognition for the main neural network; And based on the main neural network after updating the full - connection layer, processing the input information to obtain output information corresponding to the input information, so that the main neural network after updating the full - connection layer adapts to the change of the input information.

2. The method according to claim 1, wherein, The method further includes: Using the feature extraction layer of the main neural network to extract features from the input information to obtain an abstract representation of the input information; and The step of using a meta - neural network to predict the probability distribution of each parameter in the parameter memory based on the input information includes: Based on the abstract representation of the input information, using a meta - neural network to predict the probability distribution of each parameter in the parameter memory; The step of processing the input information based on the main neural network after updating the full - connection layer includes: Based on the updated full - connection layer, processing the abstract representation of the input information.

3. The method according to claim 2, wherein, The step of using a meta - neural network to predict the probability distribution of each parameter in the parameter memory based on the abstract representation of the input information includes: Taking the abstract representation of the input information as the dynamic parameter of the meta - neural network, and using the meta - neural network containing the dynamic parameter to predict the probability distribution of each parameter in the parameter memory.

4. The method according to claim 1, wherein The step of determining the parameter combination pattern of the full - connection layer of the main neural network with respect to the input information based on the probability distribution of each parameter in the parameter memory includes: Use the probability values in the probability distribution as the weight coefficients represented by the corresponding parameters in the fully connected layer. After performing weighted summation on each node in the layer above the fully connected layer of the main neural network based on the weight coefficients and then performing a non-linear transformation, obtain the parameter combination pattern of the fully connected layer of the main neural network with respect to the input information.

5. An information processing device based on a neural network, comprising: An acquisition unit configured to acquire input information, where the input information is information to be input into a main neural network for performing a processing task using the main neural network. The input information includes: text, image, or audio-video, and the processing task includes: speech recognition, text classification, image recognition; the main neural network is pre-trained based on sample information of the same type as the input information, and the main neural network is used to perform information processing on input information of the same type as the sample information input during the training process. A prediction unit configured to, based on the input information, use a meta neural network to predict the values of the respective parameters of the fully connected layer of the main neural network corresponding to the input information, obtaining a probability distribution of each parameter in a parameter memory, where the parameter memory is pre-constructed based on the fully connected layer of the main neural network; the meta neural network is a network other than the main neural network that has been pre-trained, and during the training of the meta neural network, the parameter values of the fully connected layer of the main neural network dynamically adjusted for different input information such as text, image, and audio-video are used to construct a sample set. A determination unit configured to, based on the probability distribution of each parameter in the parameter memory, determine the parameter combination pattern of the fully connected layer of the main neural network with respect to the input information. A processing unit configured to, based on the parameter combination pattern of the fully connected layer with respect to the input information, update the fully connected layer of the main neural network to provide an information processing mode with multi-strategy logic for different processing tasks such as speech recognition, text classification, and image recognition for the main neural network, and based on the main neural network after updating the fully connected layer, process the input information to obtain output information corresponding to the input information, so that the main neural network after updating the fully connected layer adapts to the changes in the input information.

6. The device according to claim 5, wherein, The device further includes: An extraction unit configured to use the feature extraction layer of the main neural network to extract features from the input information, obtaining an abstract representation of the input information; and The prediction unit is further configured to: Based on the abstract representation of the input information, use a meta neural network to predict the probability distribution of each parameter in the parameter memory. The processing unit is further configured to: Based on the updated fully connected layer, process the abstract representation of the input information.

7. The device according to claim 6, wherein The prediction unit is further configured to: Use the abstract representation of the input information as the dynamic parameter of the meta neural network, and use the meta neural network including the dynamic parameter to predict the probability distribution of each parameter in the parameter memory.

8. The device according to claim 5, wherein, The determination unit is further configured to determine the parameter combination pattern of the fully connected layer of the main neural network with respect to the input information in the following manner: Take the probability value in the probability distribution as the weight coefficient represented by the corresponding parameter in the fully connected layer, and perform a non-linear transformation after weighted summation of each node in the layer above the fully connected layer of the main neural network based on the weight coefficient, so as to obtain the parameter combination pattern of the fully connected layer of the main neural network with respect to the input information.

9. An electronic device, comprising: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-4.

10. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, the method according to any one of claims 1-4 is implemented.

Citation Information

Patent Citations

  • A video smoke recognition method based on a multi-task depth convolutional neural network

    CN108985192A

  • Information processing method and device applied to convolutional neural network

    CN109165736A