DNA damage repair mechanism inspired self-development neural network design method and storage medium

By dynamically adjusting the scale of neural network models and growing new neurons during the training process, the problem of traditional neural networks being forgotten and learning ability declined in sequence tasks is solved, and higher robustness and adaptability are achieved, which promotes the intelligent development of neural networks in image classification tasks.

CN119990224AActive Publication Date: 2025-05-13SOUTHEAST UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510062035.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

Traditional neural network image classification models have problems of forgetting and learning ability deterioration when processing sequence tasks, and the model structure cannot be dynamically adjusted during training, which limits its adaptability and efficiency in practical applications.

Method used

By dynamically adjusting the size of neural network models during training, new neurons are grown to cope with changing tasks and environments, a self-developmental neural network design method inspired by DNA damage repair mechanisms is adopted.

Benefits of technology

The intelligent development of neural networks in image classification tasks is realized, the robustness and adaptability of the model is improved, the training cost and resource consumption are reduced, and the problem of lack of learning ability in deep learning models in sequence tasks is alleviated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990224A_ABST
    Figure CN119990224A_ABST
Patent Text Reader

Abstract

The invention discloses a DNA damage repair mechanism inspired self-development neural network design method and a storage medium, and the method comprises the steps: selecting a neural network model, defining the hyper-parameters of the neural network model, obtaining an image data set, and dividing the image data set into a training set, a verification set and a test set to form a sequence task; training a neural network model based on the first task data in the training set; testing the trained neural network model based on the first task data in the verification set; carrying out self-development growth on the neural network model in width and depth directions, and increasing the scale of the neural network model; repeating the training and testing processes, judging whether the scale of the neural network model needs to be further increased or not, and if the scale is not increased, obtaining the classification precision of the model based on the task data in the test set; and repeating the process until all task data on the test set are tested. According to the method, the scale of the neural network model is dynamically adjusted in the training process, and new neurons are grown to cope with constantly changing tasks and environments, so that the intelligent development of the neural network in an image classification task and the practical application of the neural network in the image classification task are promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of neural network image processing, and in particular to a self-developing neural network design method and storage medium inspired by a DNA damage repair mechanism. Background Art

[0002] In traditional neural network image classification models, the model structure is determined at the very beginning and cannot be changed during the training and testing process, which limits its application in practice. In fact, data and tasks are often presented in a chronological order. For simple image classification tasks, only a small-scale model structure is needed, while for complex image classification tasks, a larger-scale model structure is required. In order to improve the utilization rate and maximize the efficiency of the model, different network scales should be used for different image classification tasks.

[0003] Secondly, when dealing with sequence tasks, the neural network image classification model not only forgets the old tasks, but also loses its learning ability, that is, the plasticity of the model is decreasing. The same model can achieve good performance when dealing with non-sequential tasks, but when dealing with sequence tasks, the overall performance on all data drops significantly, which limits the adaptability of the neural network image classification model to new tasks and new environments. Summary of the invention

[0004] Purpose of the invention: The purpose of the present invention is to provide a self-developing neural network design method and storage medium inspired by the DNA damage repair mechanism, including dynamically adjusting the scale of the neural network model during the training process and growing new neurons to cope with changing tasks and environments, thereby promoting the intelligent development of neural networks in image classification tasks and their practical applications.

[0005] Technical solution: To achieve the above purpose, the method for designing a self-developing neural network inspired by a DNA damage repair mechanism described in the present invention comprises the following steps:

[0006] S1: Initialize the parameter weight θ of the neural network model, define the maximum number of neural network model growth N, performance improvement threshold ε, model learning rate α, model training round number Epoch, and the number of learning categories for each task class_task;

[0007] S2: Obtain an image dataset containing image classification labels, and divide the image dataset into a training set, a validation set, and a test set based on the class_task value to form a sequence task;

[0008] S3: Train the neural network model based on the first task data in the training set and the learning rate α and model training epoch number defined in S1;

[0009] S4: Based on the first task data in the validation set, test the neural network model trained in S3 to obtain the classification accuracy C of the model on the task data;

[0010] S5: Perform self-development growth on the neural network model defined in S1 in the width and depth directions to increase the scale of the neural network model;

[0011] S6: Repeat S3 and S4 to obtain the classification accuracy C' of the new model on the task data. If C'-C>ε, repeat S3-S5 to continue to increase the size of the neural network model. If C'-C<ε and the number of times the neural network model grows in the depth direction is greater than the defined maximum number of growths N, stop the growth of the neural network model on the current task, and test based on the first task data of the test set to obtain the classification accuracy T1 of the model on the task data.

[0012] S7: Based on the second task data of the training set, repeat S3-S6 to obtain the classification accuracy T2 of the model on the task data;

[0013] S8: Repeat S7 until all image classifications of all tasks are tested, and all classification accuracies of the self-developed growth neural network model on the test set are obtained, and the performance of the self-developed growth neural network model is evaluated based on the classification accuracy.

[0014] Among them, the architecture of the neural network model described in S1 is a convolutional neural network CNN or a linear layer Linear.

[0015] The sequence task division in S2 is determined based on the total number of categories in the image dataset class_num and the number of learning categories for each task class_task. Specifically, the training set, validation set, and test set are each divided into class_num / class_task tasks, and the number of categories for each task increases successively, that is, the first task has class_task categories, the second task has 2*class_task categories, and the last task has class_num categories.

[0016] Among them, the neural network model in S5 performs self-developmental growth, first growing in the width direction and then growing in the depth direction; the growth in the width direction is to increase the number of new neurons or the number of new convolution kernel channels in the penultimate layer of the neural network model, and the growth in the depth direction is to insert a new linear layer or convolution layer in the penultimate layer of the neural network model.

[0017] Among them, for the neural network model of CNN architecture, growth in width direction refers to increasing the number of channels of the convolution kernel, and growth in depth direction refers to increasing the number of convolution layers of the entire neural network model. The number of channels of the newly added convolution layer is set to half of the number of channels of the convolution kernel of the previous layer of the model. The last layer of the neural network model of CNN architecture is the classification head. The size of the classification head is fixed and only the weight value is learned.

[0018] Among them, for the neural network model with Linear architecture, growth in width direction refers to increasing the number of neurons in the same layer, and growth in depth direction refers to increasing the number of Linear layers of the entire neural network model. The number of neurons in the newly added Linear layer is set to half the number of neurons in the previous layer of the model.

[0019] The self-development growth process of the neural network model is divided into two steps: the first step is to grow a new network structure, and the second step is to merge the original network structure and the newly added network structure; for the neural network model of the CNN architecture, the pre-growth parameter is μ 11 =[out ch ,in ch ,kernel size ,kernel size ], first increase the number of new convolution channels: μ 12 =[new ch ,in ch ,kernel size ,kernel size ], and then concatenate to get the new weight μ 1 =[out ch +new ch ,in ch ,kernel size ,kernel size ];

[0020] For the Linear architecture, the pre-growth parameter is μ 21 =[out fea ,in fea ,], first increase the number of new neurons: μ 22 =[new fea ,in fea ,], and then concatenate to get the new weight μ 2 =[out fea +new fea ,in fea ].

[0021] The present invention provides a computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions, which, when executed by a computing device, enable the computing device to execute the self-developing neural network design method inspired by the DNA damage repair mechanism described above.

[0022] Beneficial effects: The present invention has the following advantages: 1. Inspired by the DNA damage repair mechanism, the present invention dynamically adjusts the scale of the neural network model during the training process to achieve the effect of adaptively changing the model scale according to the complexity of the task. At the same time, it grows new numbers of neurons to cope with the ever-changing tasks and environments, thereby promoting the accumulation and inheritance of knowledge by the neural network model, thereby improving the robustness and adaptability of the model in the image classification process, further reducing training costs and resource consumption, and promoting the intelligent development of neural networks in image classification tasks and their practical applications;

[0023] 2. The present invention uses smaller model parameters to generate complex neural networks, alleviates the problem of lack of plasticity in deep learning models when learning sequence tasks, and improves the model's ability to adapt to and learn new tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a schematic diagram of the process of the present invention;

[0025] Figure 2 It is a flow chart of the growth of self-developing neural network;

[0026] Figure 3 The results of self-developmental growth training based on CNN architecture on CIFAR100 dataset;

[0027] Figure 4 The results of self-developmental growth training based on CNN architecture on CIFAR10 dataset;

[0028] Figure 5 It is the feature similarity matrix between different layers of the network grown based on the Linear architecture for the CIFAR10 dataset. DETAILED DESCRIPTION

[0029] The technical solution of the present invention is described in detail below in conjunction with the embodiments and drawings.

[0030] like Figure 1 As shown, the method for designing a self-developing neural network inspired by a DNA damage repair mechanism of the present invention comprises the following contents:

[0031] S1: Select the neural network architecture as CNN or Linear, initialize the model parameter weight θ, define the maximum number of network growth N, performance improvement threshold ε, model learning rate α, model training round number Epoch, number of learning categories for each task class_task, and initialize the depth growth number num_depth to 0.

[0032] S2: According to the class_task value in S1, the data set is divided into training set, validation set and test set to form a sequence task. It is determined based on the total number of categories in the data set class_num and the number of learning categories of each task class_task; specifically, the training set, validation set and test set are each divided into class_num / class_task tasks, and the number of categories of each task increases in sequence, that is, the first task has class_task categories, the second task has 2*class_task categories, and the last task has class_num categories.

[0033] S3: Based on the first task data of the training set in S2, use the learning rate α and model training round number Epoch in S1 to train the network model defined in S1. Each time the network grows, it is trained Epoch times.

[0034] S4: Using the model trained in S3, test its classification accuracy C on the first task of the validation set in S2.

[0035] S5: The network defined in S1 is grown in a self-developing manner in terms of width and depth to increase the scale of the model. The specific growth method is as follows: grow in width first, and then in depth. For the CNN architecture, growth in width refers to increasing the number of channels of the convolution kernel, and the number of channels added is set to 5. Growth in depth refers to increasing the number of convolutional layers of the network, and the number of channels of the new convolutional layer is half of the number of channels of the previous layer of the network. For the Linear architecture, growth in width refers to increasing the number of neurons in the same layer, and the number of neurons added is set to 50. Growth in depth refers to increasing the number of Linear layers of the network, and the number of neurons in the new Linear layer is half of the number of neurons in the previous layer of the network.

[0036] Growth in the width direction means adding new neurons or new convolution kernel channels to the penultimate layer of the model. Growth in the depth direction means inserting new Linear layers or convolution layers into the penultimate layer of the model. The last layer of the model is the classification head, which has a fixed size and only learns its weight values.

[0037] The self-development growth process of the neural network model is divided into two steps: the first step is to grow a new network structure, and the second step is to merge the original network structure and the newly added network structure; for the CNN architecture, its pre-growth parameter is μ 1 =[out ch ,in ch ,kernel size ,kernel size ], first increase the number of new convolution channels, μ 2 =[new ch ,in ch ,kernel size ,kernel size ], and then concatenate to get the new weight μ = [out ch +new ch ,in ch ,kernel size ,kernel size ], for the Linear architecture, its pre-growth parameter is μ 1 =[out fea ,in fea ,], first increase the number of new neurons, μ 2 =[new fea ,in fea ,], and then concatenate to get the new weight μ = [out fea +new fea ,in fea ].

[0038] S6: Repeat steps S3 and S4 to obtain a new classification accuracy of C' for the task; if C'-C>ε, repeat steps S3-S5 to continue to increase the network size; if C'-C<ε and the number of depth-direction growths num_depth increases by one, if num_depth is greater than N defined in S1, stop growing the model on the current task and test it on the first task of the test set in S2 to obtain a classification accuracy of T1; Figure 2 The figure shows a schematic diagram of the growth process of a self-developing neural network.

[0039] S7: Based on the second task data of the training set in S2, repeat steps S3-S6 and obtain a classification accuracy of T2.

[0040] S8: Repeat step S7 until all tasks are tested and the classification accuracy T1, T2…TN of the test set is obtained.

[0041] In order to evaluate whether the final network structure is reasonable, after the model has learned all the data, the similarity calculation is performed on the features learned between different layers of the model and the features between different neurons in the same layer. The calculation method is:

[0042]

[0043] Among them, x i ,x j is the feature of different layers of the network or the feature of different neurons in the same layer of the network, H is the normalized matrix, which is calculated as follows:

[0044]

[0045] I n is the identity matrix of dimension n. i The specific calculation method is as follows: based on the model θ after training all tasks, 1000 samples are randomly selected in the test set, and forward propagation is performed to obtain the features of different network layers. The feature dimension of the Linear layer is (batch, fea_dim), where batch is the batch size and fea_dim is the feature, which is calculated directly through the CKA function; for the CNN layer, the feature dimension obtained is (batch, out_ch, height, width), which needs to be averaged to obtain a two-dimensional dimension, that is, taking the mean on the weight and width dimensions to obtain the feature of (batch, out_ch), and then calculating its similarity through the CKA function.

[0046] Taking the CIFAR100 dataset and CIFAR10 dataset as examples, the self-development growth training results of the CNN architecture based on the CIFAR100 dataset are as follows: Figure 3 As shown in Figure 2, based on the CIFAR10 dataset, the self-development growth training results of the CNN architecture are as follows: Figure 4 shown.

[0047] The CIFAR10 dataset has 10 categories in total, with 2 categories for each task, and a total of 5 tasks; the CIFAR100 dataset has 100 categories in total, with 5 categories for each task, and a total of 20 tasks. Among them, seq S-DNN is the test result of the self-developed neural network on sequence tasks, seq baseline is the test result of the traditional neural network on sequence tasks, baseline S-DNN is the test result of the self-developed neural network on non-sequence tasks, and baseline is the test result of the unified neural network on non-sequence tasks. The experimental results show that the self-developed neural network alleviates the problem of lack of plasticity of the neural network when facing sequence tasks, that is, it is not that the model's expressive ability is insufficient, but that the learning ability decreases when learning sequence tasks; both test performance will decrease when learning sequence tasks, but the self-developed neural network decreases more slowly, alleviating the disadvantage of lack of plasticity.

[0048] Figure 5 The self-developed network grown under the Linear architecture for the CIFAR10 dataset calculates the feature similarity matrix between different layers based on CKA. The results show that the similarity of features learned between different layers is very low, which verifies the rationality of the self-developed neural network.

Claims

1. A method for designing a self-developing neural network inspired by DNA damage repair mechanism, characterized in that: The following steps are involved: S1: Initialize the parameter weight θ of the neural network model, define the maximum number of neural network model growth N, performance improvement threshold ε, model learning rate α, model training round number Epoch, and the number of learning categories for each task class_task; S2: Obtain an image dataset containing image classification labels, and divide the image dataset into a training set, a validation set, and a test set based on the class_task value to form a sequence task; S3: Train the neural network model based on the first task data in the training set and the learning rate α and model training epoch number defined in S1; S4: Based on the first task data in the validation set, test the neural network model trained in S3 to obtain the classification accuracy C of the model on the task data; S5: Perform self-development growth on the neural network model defined in S1 in the width and depth directions to increase the scale of the neural network model; S6: Repeat S3 and S4 to obtain the classification accuracy C' of the new model on the task data. If C'-C>ε, repeat S3-S5 to continue to increase the size of the neural network model. If C'-C<ε and the number of times the neural network model grows in the depth direction is greater than the defined maximum number of growths N, stop the growth of the neural network model on the current task, and test based on the first task data of the test set to obtain the classification accuracy T1 of the model on the task data. S7: Based on the second task data of the training set, repeat S3-S6 to obtain the classification accuracy T2 of the model on the task data; S8: Repeat S7 until all image classifications of all tasks are tested, and all classification accuracies of the self-developed growth neural network model on the test set are obtained, and the performance of the self-developed growth neural network model is evaluated based on the classification accuracy.

2. The method for designing a self-developing neural network inspired by the DNA damage repair mechanism according to claim 1, characterized in that: The architecture of the neural network model described in S1 is a convolutional neural network CNN or a linear layer Linear.

3. The method for designing a self-developing neural network inspired by the DNA damage repair mechanism according to claim 1, characterized in that: The sequence task division in S2 is determined based on the total number of categories in the image dataset class_num and the number of learning categories for each task class_task. Specifically, the training set, validation set, and test set are each divided into class_num / class_task tasks, and the number of categories for each task increases successively, that is, the first task has class_task categories, the second task has 2*class_task categories, and the last task has class_num categories.

4. The method for designing a self-developing neural network inspired by the DNA damage repair mechanism according to claim 1, characterized in that: The neural network model in S5 performs self-developmental growth, first growing in the width direction and then in the depth direction; growth in the width direction is to increase the number of new neurons or the number of new convolution kernel channels in the penultimate layer of the neural network model, and growth in the depth direction is to insert a new linear layer or convolution layer in the penultimate layer of the neural network model.

5. The method for designing a self-developing neural network inspired by the DNA damage repair mechanism according to claim 4, characterized in that: For the neural network model of CNN architecture, growth in width direction refers to increasing the number of channels of the convolution kernel, and growth in depth direction refers to increasing the number of convolution layers of the entire neural network model. The number of channels of the newly added convolution layer is set to half of the number of channels of the convolution kernel of the previous layer of the model. The last layer of the neural network model of CNN architecture is the classification head. The size of the classification head is fixed and only the weight value is learned.

6. The method for designing a self-developing neural network inspired by the DNA damage repair mechanism according to claim 4, characterized in that: For a neural network model with a linear architecture, growth in width refers to increasing the number of neurons in the same layer, and growth in depth refers to increasing the number of linear layers of the entire neural network model. The number of neurons in the newly added linear layer is set to half of the number of neurons in the previous layer of the model.

7. The method for designing a self-developing neural network inspired by the DNA damage repair mechanism according to claim 4, characterized in that: The self-development growth process of the neural network model is divided into two steps: the first step is to grow a new network structure, and the second step is to merge the original network structure and the newly added network structure; for the neural network model of the CNN architecture, the pre-growth parameter is μ 11 =[out ch ,in ch ,kernel size ,kernel size ], first increase the number of new convolution channels: μ 12 =[new ch ,in ch ,kernel size ,kernel size ], and then concatenate to get the new weight μ1=[out ch +new ch ,in ch ,kernel size ,kernel size ]; For the Linear architecture, the pre-growth parameter is μ 21 =[out fea ,in fea ,], first increase the number of new neurons: μ 22 =[new fea ,in fea ,], and then concatenate to get the new weight μ2 = [out fea +new fea ,in fea ].

8. A computer-readable storage medium storing one or more programs, characterized in that: The one or more programs include instructions which, when executed by a computing device, cause the computing device to perform any one of the methods according to claims 1 to 7.

Citation Information

Patent Citations

  • Automatic classification method of convolutional neural network constructed based on incremental branch growth

    CN110969211A

  • Automatic network growth method based on Split LBI algorithm

    CN112926723A

  • Generating neural network models, classifying physiological data, and classifying patients into clinical classifications

    US20240303492A1