Method and device for improving KANs network applicability, equipment and storage medium

By obtaining the learningable prompt set and the meta-learner to generate low-dimensional weights, the problem of KANs network parameters is solved, and efficient training and applicability are achieved in large-scale network structures.

CN120449951APending Publication Date: 2025-08-08XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510512562.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

KANs network parameters are relatively redundant, have high requirements for computing resources, and are not easy to apply in large-scale network structures.

Method used

By obtaining the set of learnable prompts and the meta-learner, the meta-learner generates lower-dimensional learnable prompts, reducing the activation function weight of the KANs network, and iterative training is combined with the stochastic gradient descent algorithm to reduce the parameter scale.

Benefits of technology

The computing resource requirements of KANs networks are reduced, and their applicability and training efficiency in large-scale network structures are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120449951A_ABST
    Figure CN120449951A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for improving KANs network applicability, equipment and a storage medium, and relates to the technical field of artificial intelligence. According to the specific implementation scheme, the method comprises the steps of obtaining a learnable prompt set and a meta-learner; wherein the learnable prompt set comprises a plurality of first learnable prompts in one-to-one correspondence with a plurality of activation functions of the to-be-trained KANs network, and the dimensions of the first learnable prompts are smaller than the weight definition dimensions corresponding to the activation functions; according to the learnable prompt set, obtaining a plurality of first weights corresponding to the plurality of activation functions through a meta-learner; and performing iterative training on the to-be-trained KANs network based on the first weight according to the training set to obtain a target KANs network. The problems that in the prior art, KANs network parameter redundancy is large, the requirement for computing resources is high, and practical application in a large-scale network structure is not easy can be solved, the parameter scale of the KANs network can be reduced, the requirement for the computing resources is lowered, and the applicability in the large-scale network structure is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and storage medium for improving the applicability of a KANs network. Background Art

[0002] Kolmogorov-Arnold Networks (KANs) replace the fixed weights of traditional multilayer perceptrons (MLPs) with learnable single-variable functions, thereby improving model interpretability. They can be used in scenarios such as image classification, symbolic regression, and solving partial differential equations. However, KANs have a massive parameter size, G+k times that of traditional MLPs (where G is the number of B-spline segments and k is the polynomial order). This results in KANs requiring a significant amount of memory and computing resources.

[0003] At present, polynomials or wavelet functions are generally used to replace the B-spline basis functions in KANs networks (such as FastKAN network and WavKAN network) to optimize computational efficiency.

[0004] However, existing improvements still have the problem of parameter redundancy, and the number of parameters is still much higher than that of MLP, which makes it difficult to complete efficient training in an environment with limited computing resources, limiting its practical application in large-scale network structures. Summary of the Invention

[0005] The embodiments of the present application provide a method, apparatus, device, and storage medium for improving the applicability of KANs networks, thereby solving the problems of the prior art KANs networks, namely, large parameter redundancy, high requirements for computing resources, and difficulty in practical application in large-scale network structures. These methods can reduce the parameter scale of the KANs network, lower the computing resource requirements of the KANs network, and improve the applicability of the KANs network in large-scale network structures.

[0006] In a first aspect, an embodiment of the present application provides a method for improving the applicability of a KANs network, including:

[0007] Obtain a learnable prompt set and a meta-learner; wherein the learnable prompt set includes multiple first learnable prompts, the multiple first learnable prompts correspond one-to-one to multiple activation functions of the KANs network to be trained, the dimension of the first learnable prompt is smaller than the weight definition dimension corresponding to the activation function of the KANs network to be trained, and the meta-learner is used to output the weight of the activation function corresponding to the learnable prompt according to the input learnable prompt; according to the learnable prompt set, obtain multiple first weights corresponding to the multiple activation functions of the KANs network to be trained through the meta-learner; wherein the multiple activation functions of the KANs network to be trained correspond one-to-one to the multiple first weights; according to the training set, iteratively train the KANs network to be trained based on the first weights to obtain a target KANs network; wherein the training set includes multiple training samples, the training samples include images to be classified and classification labels, and the target KANs network is used to output corresponding image categories according to the input images.

[0008] Furthermore, a set of prompts can be learned Among them, L represents the total number of layers of the KANs network to be trained, I l Represents the set of activation function indices of the I layer of the KANs network to be trained, with I l =[n l ]×[n l+1 ],n l Indicates the number of nodes in the lth layer of the KANs network to be trained, n l+1 Indicates the number of nodes in the l+1th layer of the KANs network to be trained, It represents the first learnable hint corresponding to the αth activation function of the lth layer in the KANs network to be trained, and d represents the dimension of the first learnable hint.

[0009] Furthermore, d is set as a hyperparameter, and d=1 is taken.

[0010] Furthermore, the meta-learner is a two-layer MLP network structure with one hidden layer.

[0011] Furthermore, the KANs network to be trained includes a standard KANs network, a ConvKAN network, a FastKAN network, and a WavKAN network.

[0012] Furthermore, the meta-learners include multiple, the KANs network to be trained includes multiple groups, each group includes multiple continuous layer structures, and the multiple meta-learners correspond one-to-one to the multiple groups; according to the learnable prompt set, the first weights of the multiple activation functions of the KANs network to be trained are obtained through the meta-learners, including: according to the learnable prompt set, the multiple first weights of the multiple activation functions corresponding to the multiple groups in the KANs network to be trained are obtained through the multiple meta-learners.

[0013] Furthermore, the method further includes: updating the meta-learner and the first learnable hint by a stochastic gradient descent algorithm during iterative training of the KANs network to be trained based on the first weight.

[0014] In a second aspect, an embodiment of the present application provides a device for improving the applicability of a KANs network, including:

[0015] An acquisition module is configured to acquire a set of learnable prompts and a meta-learner; wherein the set of learnable prompts includes a plurality of first learnable prompts, the plurality of first learnable prompts correspond one-to-one to a plurality of activation functions of the KANs network to be trained, the dimension of the first learnable prompts being smaller than the dimension of weights defined by the activation functions of the KANs network to be trained, and the meta-learner is configured to output weights of the activation functions corresponding to the learnable prompts based on the input learnable prompts.

[0016] A weight module is configured to obtain, through a meta-learner, a plurality of first weights corresponding to a plurality of activation functions of a KANs network to be trained based on a set of learnable prompts; wherein the plurality of activation functions of the KANs network to be trained corresponds one-to-one to the plurality of first weights.

[0017] The training module is used to train the KANs network to be trained based on the first weight according to the training set to obtain a target KANs network; wherein the training set includes multiple training samples, the training samples include images to be classified and classification labels, and the target KANs network is used to output the corresponding image category according to the input image.

[0018] In a third aspect, an embodiment of the present application provides a device comprising: a processor; a memory for storing processor-executable instructions; and a method for implementing the first aspect or any possible implementation of the first aspect when the processor executes the executable instructions.

[0019] In a fourth aspect, an embodiment of the present application provides a non-volatile computer-readable storage medium, which includes a device for storing a computer program or instruction, and when the computer program or instruction is executed, the method of the first aspect or any possible implementation method of the first aspect is implemented.

[0020] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:

[0021] The embodiment of the present disclosure obtains a set of learnable prompts and a meta-learner, and obtains a plurality of first weights corresponding to a plurality of activation functions of a KANs network to be trained by the meta-learner according to the set of learnable prompts.

[0022] According to the training set, the KANs network to be trained based on the first weight is iteratively trained to obtain the target KANs network. The first weight that can reduce the parameter scale is obtained through a lower-dimensional learnable hint and a lightweight meta-learner. By introducing the first weight to the KANs network to be trained, the parameter scale in the KANs network to be trained can be reduced, so that the target KANs network obtained by training has a smaller parameter scale, thereby reducing the computing resource requirements of the KANs network and improving the applicability of the KANs network in large-scale network structures. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments of the present application or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0024] Figure 1 A flowchart of a method for improving the applicability of a KANs network provided in an embodiment of the present application;

[0025] Figure 2 A diagram showing the visualization of the classification accuracy and parameter count of the KANs network under different test sets after training on the SVHN dataset using traditional methods and the method of this application in image classification tasks.

[0026] Figure 3 This is the T-SNE dimensionality reduction visualization result of the features extracted from the test set after the KANs network was trained on the SVHN dataset using the method of this application;

[0027] Figure 4 In order to use the traditional method and the method of this application in low-dimensional and high-dimensional function fitting task scenarios, the KANs network is used in the function Schematic diagram of the visualization results of the number of parameters and fitting MSE under different input dimensions;

[0028] Figure 5 A schematic diagram of the composition of an apparatus for improving the applicability of a KANs network provided in an embodiment of the present application. DETAILED DESCRIPTION

[0029] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0030] The following description of some of the technologies involved in the embodiments of this application is provided to facilitate understanding and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for the sake of clarity and conciseness, some descriptions of well-known functions and structures are omitted from the following description.

[0031] Kolmogorov-Arnold Networks (KANs) replace the fixed weights of traditional multilayer perceptrons (MLPs) with learnable single-variable functions, thereby improving model interpretability. They can be used in scenarios such as image classification, symbolic regression, and solving partial differential equations. However, KANs have a large parameter size, G+k times that of traditional MLPs (where G is the number of B-spline segments and k is the polynomial order). This results in a high memory usage during training and high computational resource requirements.

[0032] At present, polynomials or wavelet functions are generally used to replace the B-spline basis functions in KANs networks (such as FastKAN network and WavKAN network) to optimize computational efficiency.

[0033] However, existing improvements still have the problem of parameter redundancy during model training, and the number of parameters is still much higher than that of MLP, which makes it difficult to complete efficient training in an environment with limited computing resources, limiting its practical application in large-scale network structures.

[0034] Against this background technology, the present disclosure provides a method for improving the applicability of KANs networks, which can solve the problems of existing KANs networks with large parameter redundancy, high requirements for computing resources, and difficulty in practical application in large-scale network structures. It can reduce the parameter scale of the KANs network, reduce the computing resource requirements of the KANs network, and improve the applicability of the KANs network in large-scale network structures.

[0035] The execution subject of the method for improving the applicability of KANs networks provided in the embodiments of the present disclosure may be a computer or server, or other electronic devices with data processing capabilities; alternatively, the execution subject of the method may be a processor (e.g., a central processing unit (CPU)) in the above electronic device; alternatively, the execution subject of the method may be an application (APP) installed in the above electronic device that can implement the functions of the method; alternatively, the execution subject of the method may be a functional module or unit in the above electronic device that has the functions of the method. The execution subject of the method is not limited herein.

[0036] The method for improving the applicability of the KANs network is exemplarily described below with reference to the accompanying drawings.

[0037] Figure 1 Schematic diagram of a method for improving the applicability of a KANs network provided in an embodiment of the present application.

[0038] in, Figure 1 This is only an execution order shown in the embodiment of the present application, and does not represent the only execution order of the method for improving the applicability of the KANs network. If the final result can be achieved, Figure 1 The steps shown can be performed in parallel or in reverse. Figure 1 As shown, the method may include:

[0039] S101. Obtain a set of learnable prompts and a meta-learner.

[0040] The set of learnable prompts includes multiple first learnable prompts, and the multiple first learnable prompts correspond one-to-one to multiple activation functions of the KANs network to be trained. The dimension of the first learnable prompts is smaller than the weight definition dimension corresponding to the activation function of the KANs network to be trained. The meta-learner is used to output the weight of the activation function corresponding to the learnable prompt based on the input learnable prompt.

[0041] For example, in a KANs network, a learnable hint serves as a dynamically adjustable parameter of the activation function. The quality of its initialization directly affects the training stability and convergence speed of the model. Kaiming initialization can be used to obtain multiple first learnable hints, thereby obtaining a set of learnable hints. It is understood that the first learnable hint can serve as an identifier for the activation function in the KANs network to be trained.

[0042] A specific, learnable set of prompts Among them, L represents the total number of layers of the KANs network to be trained, I l Represents the set of activation function indices of the lth layer of the KANs network to be trained, with Il =[n l ]×[n l+1 ],n l Indicates the number of nodes in the lth layer of the KANs network to be trained, n l+1 Indicates the number of nodes in the lth layer of the KANs network to be trained, It represents the first learnable hint corresponding to the αth activation function in the l+1th layer of the KANs network to be trained, and d represents the dimension of the first learnable hint.

[0043] Furthermore, d can be set as a hyperparameter, taking d = 1, to minimize the number of trainable parameters and improve storage efficiency while maintaining the expressiveness of the model.

[0044] For example, the meta-learner can learn the potential shared function family behind different activation functions in the KANs network, and thus learn the shared activation function weight setting rules.

[0045] Specifically, the meta-learner is a two-layer MLP network structure with one hidden layer.

[0046] For example, the parameters of the meta-learner can be obtained through xavier initialization (Sergey Gonzarov initialization).

[0047] For example, when the input of the meta-learner M is When , the output of the meta-learner M under the meta-learner parameter θ can be expressed as

[0048] It should be noted that the input of the meta-learner is the first learnable prompt, and the output is the weight of the activation function corresponding to the first learnable prompt. The output dimension is consistent with the weight definition dimension corresponding to the activation function of the KANs network to be trained, which can be expressed as It represents the weight of the αth activation function in the lth layer of the KANs network to be trained, and dim(*) represents the dimension.

[0049] For example, the size of the hidden layer of the meta-learner is adjustable, and can be selected as 32, 64, or 128 to adapt to KANs networks of different scales and tasks.

[0050] S102: According to the learnable prompt set, a meta-learner is used to obtain a plurality of first weights corresponding to a plurality of activation functions of the KANs network to be trained.

[0051] Among them, the multiple activation functions of the KANs network to be trained correspond one-to-one to the multiple first weights.

[0052] For example, any first learnable prompt in the set of learnable prompts is input into the meta-learner, and the meta-learner can output the first weight of the activation function corresponding to the first learnable prompt. For multiple first learnable prompts input into the meta-learner, the first weight of the activation function corresponding to each first learnable prompt can be obtained.

[0053] For example, taking the learnable prompt set including the first learnable prompts A, B, and C, and the activation functions corresponding to the first learnable prompts A, B, and C are activation functions a, b, and c respectively, the first learnable prompts A, B, and C are input into the meta-learner, and the weights of the activation functions a, b, and c can be obtained respectively. The weights of the activation functions a, b, and c are the first weights.

[0054] For example, taking the traditional KANs network as an example, the traditional KANs network uses the B-spline activation function, and the calculation method of each activation function can be expressed as Where B(x)=[SiLU(x),B1(x),…,B G+k (x)] T is a set of bases, B(x) represents the B-spline activation function, x represents the feature, SiLU(*) represents the SiLU activation function, B G+k (*) represents the G+kth B-spline activation function, G is the number of B-spline segments, k is the polynomial order, and w is the weight. and the meta-learner M under parameters θ θ , the calculation of each activation function can be expressed as have

[0055] S103 : According to the training set, the KANs network to be trained based on the first weight is trained to obtain a target KANs network.

[0056] The training set includes multiple training samples, and the training samples include images to be classified and classification labels. The target KANs network is used to output the corresponding image category based on the input image.

[0057] Illustratively, the KANs network to be trained based on the first weight may be a KANs network to be trained after the weights of each activation function in the network are respectively set to the first weight corresponding to each activation function.

[0058] For example, the image to be classified can be used as the input of the KANs network to be trained, and the classification label can be used as the output of the KANs network to be trained. The KANs network to be trained is trained to obtain a target KANs network.

[0059] It should be noted that the content of the training samples in the training set can be modified based on the target KANs network of the desired function, thereby obtaining the target KANs network of the desired function. For example, in the application scenarios of symbolic regression and solving partial differential equations, the training samples can be numerical values and numerical labels. The obtained target KANs network is used to obtain the corresponding dependent variable value based on the input independent variable value.

[0060] For example, training can be performed using a forward propagation method, and based on the output and classification labels of the KANs network to be trained, a loss value is calculated according to a loss function (such as a cross entropy loss function, etc.), and the network parameters of the KANs network to be trained are updated using a stochastic gradient descent method according to the loss value until the iteration termination condition is met (such as reaching the maximum number of iterations or the loss value converges), thereby obtaining the target KANs network.

[0061] In this way, the traditional KANs network with a scale of ∑((n l ×n l+1 )×(G+k+1)) parameter quantity is reduced to ∑((n l ×n l+1 )+(d hidden +1)(G+k+1)), where G is the number of B-spline segments, k is the polynomial order, and d hidden represents the size of the hidden layer in the meta-learner,

[0062] The embodiment of the present disclosure obtains a set of learnable prompts and a meta-learner, and obtains multiple first weights corresponding to multiple activation functions of the KANs network to be trained through the meta-learner based on the set of learnable prompts. According to the training set, the KANs network to be trained based on the first weights is iteratively trained to obtain a target KANs network. By introducing the first weights, the parameter scale of the KANs network can be reduced, the computing resource requirements required for the KANs network during training can be reduced, the training efficiency can be improved, and the applicability of the KANs network in large-scale network structures can be improved.

[0063] In some possible embodiments, the KANs network to be trained includes a standard KANs network, a ConvKAN network, a FastKAN network, and a WavKAN network.

[0064] For example, the standard KANs network is the original KANs network constructed using the Kolmogorov-Arnold theorem in the paper "KAN: Kolmogorov-Arnold Networks" by Ziming Liu et al. The standard KANs network (generally referred to as the KANs network) can be referred to in the aforementioned embodiments and will not be described in detail here.

[0065] Exemplarily, for the WavKAN network, the output of the meta-learner can be mapped to a parameter triple of the wavelet function, and the parameters of each wavelet activation function are dynamically generated by the meta-learner, encoding the function characteristics with a scalar prompt of the first learnable prompt.

[0066] For example, for the FastKAN network, the output of the meta-learner can be mapped to a weight vector of a radial basis function (RBF), and the parameters of the RBF can be generated by the meta-learner.

[0067] For example, for the ConvKAN network, learnable hints can be assigned to the activation function within the convolution kernel according to the input / output channel index and the convolution kernel position, and the corresponding parameters can be generated by the meta-learner.

[0068] It should be noted that other KANs networks may also be included, and there is no limitation to this.

[0069] In some possible implementations, the meta-learners include multiple, the KANs network to be trained includes multiple groups, each group includes multiple consecutive layer structures, and the multiple meta-learners correspond one-to-one to the multiple groups.

[0070] According to the set of learnable hints, the first weights of multiple activation functions of the KANs network to be trained are obtained through the meta-learner, including:

[0071] According to the learnable prompt set, multiple first weights of multiple activation functions corresponding to multiple groups in the KANs network to be trained are obtained through multiple meta-learners.

[0072] For example, when the KANs network to be trained is a deep structure, the entire layer structure can be divided into multiple groups by layer clustering based on the number of input channels and the number of output channels.

[0073] For example, taking the layer structure of the KANs network to be trained as 100 layers, the layer structure of the KANs network to be trained can be clustered according to the number of channels, the cluster center is set to C, and the 100 layers are divided into C groups using the K-means clustering algorithm.

[0074] It should be noted that the number of layers in each group can be the same or different, and there is no restriction on this. For example, if the group is divided into four groups, the first group may include layers 1 to 20 of the KANs network to be trained, the second group may include layers 21 to 60, the third group may include layers 61 to 75, and the fourth group may include layers 76 to 100.

[0075] Continuing with the above example, when divided into 4 groups, the number of meta-learners is also 4, and the 4 meta-learners correspond one-to-one to the 4 groups.

[0076] Exemplarily, among multiple meta-learners, the input of any meta-learner is the first learnable hint corresponding to the activation function of the layer structure of its corresponding group, and the output is the first weight of the activation function of the layer structure of its corresponding group.

[0077] Continuing with the above example, the input of the meta-learner corresponding to the first group is the first learnable cue corresponding to the activation functions of the 1st to 20th layers of the KANs network to be trained, and the output is the first weight of the activation functions of the 1st to 20th layers of the KANs network to be trained; the input of the meta-learner corresponding to the second group is the first learnable cue corresponding to the activation functions of the 21st to 60th layers of the KANs network to be trained, and the output is the first weight of the activation functions of the 21st to 60th layers of the KANs network to be trained; the input of the meta-learner corresponding to the third group is the first learnable cue corresponding to the activation functions of the 61st to 75th layers of the KANs network to be trained, and the output is the first weight of the activation functions of the 61st to 75th layers of the KANs network to be trained; the input of the meta-learner corresponding to the fourth group is the first learnable cue corresponding to the activation functions of the 76th to 100th layers of the KANs network to be trained, and the output is the first weight of the activation functions of the 76th to 100th layers of the KANs network to be trained.

[0078] It can be understood that the shallow structure can also set multiple meta-learners to output corresponding first weights for different partial layer structures of the KANs network to be trained.

[0079] This embodiment provides a corresponding meta-learner for the layer structure of the KANs network to be trained, so that the method of the present application can be better adapted to the KANs network to be trained, and can reduce the computing resources required for the KANs network during training while ensuring the performance of the KANs network as much as possible.

[0080] Furthermore, the method may also include:

[0081] During the iterative training of the to-be-trained KANs network based on the first weights, the meta-learner and the first learnable hint are updated by a stochastic gradient descent algorithm.

[0082] For example, the parameters of the meta-learner and the first learnable hint can be updated by a stochastic gradient descent algorithm according to the loss value obtained during iterative training of the KANs network to be trained.

[0083] In this way, the updated meta-learner and the first learnable hint can be used for inference or fine-tuning of the already trained KANs network on new tasks, further improving the performance of the KANs network.

[0084] Verification experiment

[0085] Figure 2 This is a diagram showing the visualization of the classification accuracy and number of parameters of the KANs network (where G = 5, k = 3) trained on the SVHN dataset in image classification tasks using traditional methods and the method of this application. Figure 2 , the dots represent the traditional method, the diamond dots represent the method of the present application, and different colors represent different test sets, among which light blue represents the SVHN test set (the fourth pair from top to bottom in the figure), blue represents the FMNIST test set (the second pair from top to bottom in the figure), light green represents the KMNIST test set (the third pair from top to bottom in the figure), green represents the MNIST test set (the first pair from top to bottom in the figure), pink represents the CIFAR-10 test set (the fifth pair from top to bottom in the figure), and red represents the CIFAR-100 test set (the sixth pair from top to bottom in the figure). It can be seen that on different test sets, compared with the traditional method, the method of the present application reduces the number of network parameters and improves the network classification accuracy.

[0086] Figure 3 This is the T-SNE dimensionality reduction visualization result of the features extracted from the test set after the KANs network was trained on the SVHN dataset using the method of this application. Figure 3 It can be seen that the method of this application can extract features of different categories with more obvious boundaries, thereby achieving more accurate classification prediction.

[0087] Figure 4 In order to use the traditional method and the method of this application in low-dimensional and high-dimensional function fitting task scenarios, the KANs network is used in the function Schematic diagram of the visualization results of the number of parameters (Parameter Count) and the mean squared error (MSE) of the fitting under different input dimensions (Problem Dimension). Figure 4 The blue dotted solid line represents the fitting MSE result of the traditional method, the blue dotted line represents the fitting MSE result of the method of the present application, the orange dotted solid line represents the parameter number result of the traditional method, and the orange dotted line represents the parameter number result of the method of the present application. It can be seen that compared with the traditional method, no matter how the input dimension changes, the method of the present application has fewer network parameters. At the same time, when the input dimension is low (approximately less than 20), the fitting MSE of the method of the present application is roughly the same as that of the traditional method. When the input dimension is high (approximately greater than 65), the fitting MSE of the method of the present application is lower.

[0088] Although the present application provides method operation steps such as embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative work. The order of steps listed in this embodiment is only one way of executing the order of many steps and does not represent the only execution order. When an actual device or client product is executed, it can be executed in the order of the method shown in this embodiment or the accompanying drawings or in parallel (for example, in a parallel processor or multi-threaded processing environment).

[0089] like Figure 5 As shown, the embodiment of the present application also provides a device for improving the applicability of a KANs network. The device includes:

[0090] Acquisition module 501 is configured to acquire a set of learnable cues and a meta-learner; wherein the set of learnable cues includes a plurality of first learnable cues, each of the plurality of first learnable cues corresponding one-to-one to a plurality of activation functions of the KANs network to be trained, wherein the dimensions of the first learnable cues are smaller than the dimensions defining the weights corresponding to the activation functions of the KANs network to be trained, and the meta-learner is configured to output weights of the activation functions corresponding to the learnable cues based on the input learnable cues;

[0091] A weight module 502 is configured to obtain, through a meta-learner, a plurality of first weights corresponding to a plurality of activation functions of the KANs network to be trained based on the set of learnable cues; wherein the plurality of activation functions of the KANs network to be trained corresponds one-to-one to the plurality of first weights;

[0092] The training module 503 is used to train the KANs network to be trained based on the first weight according to the training set to obtain a target KANs network; wherein the training set includes multiple training samples, and the training samples include images to be classified and classification labels. The target KANs network is used to output the corresponding image category based on the input image.

[0093] The beneficial effects and specific implementation methods of the present device embodiment can be referred to the aforementioned method embodiment, and will not be described in detail here.

[0094] Some modules in the apparatus described herein may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0095] The devices or modules described in the above application embodiments can be implemented by computer chips or physical devices, or by products with certain functions. For ease of description, the above devices are described separately by function in various modules. When implementing the embodiments of this application, the functions of each module can be implemented in the same or multiple software and / or hardware. Of course, a module that implements a certain function can also be implemented by combining multiple sub-modules or sub-units.

[0096] The methods, devices, or modules described in this application can be implemented in the form of computer-readable program code. The controller can be implemented in any appropriate manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (such as software or firmware) that can be executed by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to implement the same function of the controller in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the means for implementing various functions may be considered to be both a software module for implementing the method and a structure within a hardware component.

[0097] An embodiment of the present application further provides a device, comprising: a processor; a memory for storing processor-executable instructions; and when the processor executes the executable instructions, the method described in the embodiment of the present application is implemented.

[0098] The embodiments of the present application also provide a non-volatile computer-readable storage medium having a computer program or instruction stored thereon. When the computer program or instruction is executed, the method described in the embodiments of the present application is implemented.

[0099] In addition, each functional module in each embodiment of the present invention may be integrated into one processing module, or each module may exist independently, or two or more modules may be integrated into one module.

[0100] The above-mentioned storage medium includes, but is not limited to, random access memory (RAM), read-only memory (ROM), cache, hard disk drive (HDD), or memory card. The memory can be used to store computer program instructions.

[0101] It can be seen from the description of the above implementation methods that those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary hardware. Based on this understanding, the technical solution of the present application can essentially or the part that contributes to the prior art can be embodied in the form of a software product, or it can be embodied through the implementation process of data migration. The computer software product can be stored in a storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a mobile terminal, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application or certain parts of the embodiments.

[0102] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referenced to each other. Each embodiment focuses on the differences from other embodiments. All or part of this application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, mobile communication terminals, multi-processor systems, microprocessor-based systems, programmable electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc.

[0103] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit the present application. Although the present application has been described in detail with reference to the aforementioned embodiments, a person of ordinary skill in the art should understand that the technical solutions described in the aforementioned embodiments can still be modified, or some or all of the technical features therein can be replaced by equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the present application.

Claims

1. A method for improving the applicability of a KANs network, characterized in that: include: Obtain a set of learnable prompts and a meta-learner; The set of learnable prompts includes a plurality of first learnable prompts, each of the plurality of first learnable prompts corresponds one-to-one to a plurality of activation functions of the KANs network to be trained, the dimension of the first learnable prompts being smaller than the dimension of weight definition corresponding to the activation function of the KANs network to be trained, and the meta-learner is configured to output weights of the activation functions corresponding to the learnable prompts based on the input learnable prompts; According to the set of learnable prompts, obtaining, by the meta-learner, a plurality of first weights corresponding to a plurality of activation functions of the KANs network to be trained; Wherein, the multiple activation functions of the KANs network to be trained correspond one-to-one to the multiple first weights; Iteratively training the KANs network to be trained based on the first weights according to the training set to obtain a target KANs network; The training set includes multiple training samples, and the training samples include images to be classified and classification labels. The target KANs network is used to output corresponding image categories based on the input images.

2. The method according to claim 1, characterized in that The set of learnable hints Wherein, L represents the total number of layers of the KANs network to be trained, I l Represents the set of activation function indexes of the lth layer of the KANs network to be trained, with I l =[n l ]×[n l+1 ],n l Indicates the number of nodes in the first layer of the KANs network to be trained, n l+1 represents the number of nodes in the l+1th layer of the KANs network to be trained, represents the first learnable hint corresponding to the αth activation function of the lth layer in the KANs network to be trained, and d represents the dimension of the first learnable hint.

3. The method according to claim 2, characterized in that Set d as a hyperparameter and take d=1.

4. The method according to claim 1, wherein The meta-learner is a two-layer MLP network structure including one hidden layer.

5. The method according to claim 1, wherein The KANs network to be trained includes a standard KANs network, a ConvKAN network, a FastKAN network, and a WavKAN network.

6. The method according to claim 1, characterized in that The meta-learners include a plurality of meta-learners, the KANs network to be trained includes a plurality of groups, each group includes a plurality of continuous layer structures, and the plurality of meta-learners correspond one-to-one to the plurality of groups; The method of obtaining first weights of multiple activation functions of the KANs network to be trained by a meta-learner based on the learnable prompt set includes: According to the learnable prompt set, multiple first weights of multiple activation functions corresponding to multiple groups in the KANs network to be trained are obtained through multiple meta-learners.

7. The method according to claim 1, characterized in that The method further comprises: During the iterative training of the KANs network to be trained based on the first weights, the meta-learner and the first learnable hint are updated by a stochastic gradient descent algorithm.

8. A device for improving the applicability of a KANs network, characterized in that: include: an acquisition module configured to acquire a set of learnable prompts and a meta-learner; wherein the set of learnable prompts includes a plurality of first learnable prompts, each of the plurality of first learnable prompts corresponding one-to-one to a plurality of activation functions of the KANs network to be trained, wherein the dimension of the first learnable prompts is smaller than the dimension defining the weights corresponding to the activation functions of the KANs network to be trained, and the meta-learner is configured to output weights of the activation functions corresponding to the learnable prompts based on the input learnable prompts; a weight module, configured to obtain, by the meta-learner, a plurality of first weights corresponding to a plurality of activation functions of the KANs network to be trained based on the set of learnable prompts; wherein the plurality of activation functions of the KANs network to be trained corresponds one-to-one to the plurality of first weights; A training module is configured to train the KANs network to be trained based on the first weights according to a training set to obtain a target KANs network; wherein the training set includes multiple training samples, the training samples include images to be classified and classification labels, and the target KANs network is configured to output a corresponding image category based on an input image.

9. A device for executing a method for improving the applicability of a KANs network, characterized in that: include: processor; a memory for storing processor-executable instructions; When the processor executes the executable instructions, the method according to any one of claims 1 to 7 is implemented.

10. A non-volatile computer-readable storage medium, characterized in that: The device comprises a computer program or an instruction for storing the computer program or the instruction, which, when executed, enables the method according to any one of claims 1 to 7 to be implemented.