Machine learning device and machine learning method

WO2026204663A1PCT designated stage Publication Date: 2026-10-01JVC KENWOOD CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2026/010707
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-26
Filing Date
2026-03-18
Publication Date
2026-10-01

Smart Images

  • Figure JP2026010707_01102026_PF_FP_ABST
    Figure JP2026010707_01102026_PF_FP_ABST
Patent Text Reader

Abstract

A machine learning device (30) comprises: a training unit (31) that trains a feature quantity output unit (20) that outputs feature quantities of input data for training a new class; a first connection adjustment unit (32) that forms a node path for the new class by connecting or disconnecting nodes between adjacent layers in the feature quantity output unit (20); and a second connection adjustment unit (34) that, when the similarity between a prototype and an adjusted feature quantity calculated on the basis of the output of an intermediate layer in the feature quantity output unit (20) is equal to or greater than a threshold value, sets the weights of the nodes in layers located further to the output side than the intermediate layer along the path to 0.
Need to check novelty before this filing date? Find Prior Art

Description

Machine learning device and machine learning method

[0001] This invention relates to machine learning technology.

[0002] Humans can learn new knowledge through long-term experience and retain old knowledge. On the other hand, the knowledge of a Convolutional Neural Network (CNN) depends on the dataset used for training, and in order to adapt to changes in the data distribution, the CNN parameters need to be retrained for the entire dataset. In a CNN, as it learns about new tasks, the estimation accuracy for old tasks decreases. Thus, in a CNN, continuous learning inevitably leads to catastrophic forgetting, where the learning results for old tasks are forgotten while learning a new task.

[0003] As a more efficient and practical method, incremental learning (or continual learning) has been proposed, which reuses already acquired knowledge and learns new tasks without forgetting knowledge from past tasks. Incremental learning is a learning method in which, when a new task or new data arises, instead of learning a model from scratch, the currently trained model is improved upon. In deep learning, there is a phenomenon called fatal forgetting, in which previously acquired knowledge is lost at a large rate, and the ability to perform tasks decreases significantly, and this is a particularly problematic issue in incremental learning. In classification tasks, incremental learning is a method that allows the model to learn and classify new classes (novel classes) from a state where it can classify previously learned classes (base classes). The biggest challenge is to acquire classification performance for new classes while maintaining classification performance for base classes by avoiding fatal forgetting.

[0004] As one continual learning method for avoiding catastrophic forgetting, NISPA (Neuro-Inspired Stability-Plasticity Adaptation) has been proposed (see, for example, Non-Patent Document 1). NISPA is a method that imitates the memory mechanism of the human brain, and deletes or adds paths between nodes in adjacent layers of a neural network during continual learning. In NISPA, paths between nodes with high activation values (stable nodes) obtained through learning of base classes are maintained, while paths between other nodes are randomly pruned. Accordingly, in NISPA, among the memory paths obtained through learning of base classes, paths between stable nodes that are highly likely to be commonly usable for classification of other classes can be preferentially maintained.

[0005] Mustafa Burak Gurbuz & Constantine Dovrolis (2022). NISPA: Neuro-Inspired Stability-Plasticity Adaptation for Continual Learning in Sparse Networks. International Conference on Machine Learning 2022. arXiv:2206.09117.Semantic memory-based dynamic neural network using memristive ternary CIM and CAM for 2D and 3D vision.Prototypical Networks for Few-shot Learning.

[0006] In NISPA, the number of nodes available for classification of other classes decreases every time learning of a new class task is performed, and thus there is a problem that the number of classes that can be learned is limited.

[0007] In view of the above, an object of the present invention is to provide a technique that makes it easy to leave nodes available for classification of other classes even when learning of a new class task is performed.

[0008] To solve the above problems, a machine learning device in one aspect of this embodiment includes: a learning unit that trains a feature output unit that outputs feature quantities of input data for training a new class using a neural network; a first connection adjustment unit that forms a node path for the new class by connecting or disconnecting nodes between adjacent layers in the neural network of the feature output unit; and a second connection adjustment unit that sets the weights of nodes in layers on the output side of the intermediate layer in the path to 0 if the similarity between the adjustment feature, which represents the representative feature quantities of the test input data for the new class calculated based on the output of the intermediate layer in the neural network of the feature output unit, and the prototype, which represents the representative feature quantities of the training input data, is greater than or equal to a threshold.

[0009] Another aspect of this embodiment is a machine learning method. This machine learning method includes: training a neural network that outputs features of input data using training input data of a new class; forming a node path for the new class by connecting or disconnecting nodes between adjacent layers in the neural network; generating a prototype that shows representative features of the training input data; calculating adjustment features that show representative features of the test input data of the new class based on the output of the hidden layer in the neural network; calculating the similarity between the adjustment features and the prototype; and, if the similarity is above a threshold, setting the weights of the nodes in the path that are on the output side of the hidden layer to 0.

[0010] Furthermore, any combination of the above components, as well as conversions of the expressions of this embodiment between methods, apparatus, systems, recording media, computer programs, etc., are also valid as embodiments of this embodiment.

[0011] According to one aspect of this embodiment, even when learning a task for a new class, it is possible to retain nodes that can be used for classifying other classes.

[0012] This is a functional block diagram schematically showing the general configuration of the classification device according to the embodiment. This is a flowchart illustrating the classification process by the classification device in Figure 1. This is a functional block diagram schematically showing the general configuration of the learning device according to the embodiment. This is a flowchart illustrating the learning process by the machine learning device. This is a diagram showing an example of connection adjustment processing. This is a diagram showing an example of connection adjustment processing in the second embodiment.

[0013] Embodiments of the present invention will be described below with reference to the drawings. The specific numerical values ​​and other details shown in these embodiments are merely examples to facilitate understanding of the invention and do not limit the present invention unless otherwise specified. Elements not directly related to the present invention are omitted from the drawings.

[0014] Figure 1 of the first embodiment is a schematic functional block diagram showing the general configuration of the classification device 1 according to the first embodiment. As shown in Figure 1, the classification device 1 comprises an input unit 10, a feature output unit 20, a feature adjustment unit 25, a feature storage unit 26, a classification unit 40, and an output unit 50.

[0015] The input unit 10 receives input data to be classified by the classification device 1. The input data is, for example, image data of an object, and the object being photographed can be an animal, a vehicle, a person, etc.

[0016] The feature output unit 20 outputs the features of the input data received by the input unit 10. The feature output unit 20 is a trained neural network model. For example, the feature output unit 20 is a trained CNN model. For example, the feature output unit 20 is trained using a dataset of basic classes (classes trained in the past), which are big data. The feature output unit 20 has undergone prior training, including deleting or adding paths between nodes in adjacent layers of the neural network. In this embodiment, the feature output unit 20 has multiple paths formed by stable nodes of the neural network, as described later. Each of these multiple paths can process a different class of task. The feature output unit 20 is trained by the machine learning device 30. The feature output unit 20 may have completed training, or it may be updatable by further training. The method for training the feature output unit 20 by the machine learning device 30 will be described later.

[0017] The feature adjustment unit 25 generates adjusted features by performing calculations on the features output from the feature output unit 20 (for example, calculating the mean, median, and representative value). The feature adjustment unit 25 generates adjusted features that represent the representative features of the data belonging to each class by performing calculations on the features output from the feature output unit 20 for each class. These adjusted features that represent the representative features of the data belonging to each class are also called prototypes (see, for example, Non-Patent Documents 2 and 3). For example, if there are 15 paths in the neural network of the feature output unit 20, 15 prototypes corresponding to each path are generated. The adjusted features in this embodiment are the mean values ​​of the features.

[0018] The feature storage unit 26 stores the prototypes of each class generated by the feature adjustment unit 25. The prototypes of each class are assumed to have been pre-calculated using the basic class dataset and stored in the feature storage unit 26.

[0019] The classification unit 40 classifies the input data received by the input unit 10. The classification unit 40 calculates the similarity between the adjusted features of the input data and the prototypes of all classes, and classifies the input data by selecting the class with the highest similarity. The similarity here is, for example, cosine similarity.

[0020] The output unit 50 outputs the classification result from the classification unit 40. That is, the output unit 50 outputs information indicating which class the input data was classified into.

[0021] Figure 2 is a flowchart illustrating the classification process performed by the classification device 1 shown in Figure 1.

[0022] In S11, the feature output unit 20 calculates the feature quantities of the input data received by the input unit 10. The feature output unit 20 supplies the calculated feature quantities to the feature adjustment unit 25.

[0023] In S12, the feature adjustment unit 25 calculates adjusted features from the features of the input data. In this embodiment, the feature adjustment unit 25 calculates adjusted features of the input data by calculating the average value of the features of the input data. The feature adjustment unit 25 supplies the calculated adjusted features to the classification unit 40.

[0024] In S13, the classification unit 40 reads out the prototypes of all classes from the feature storage unit 26 and calculates the similarity between the input adjusted features and the prototypes of all classes.

[0025] In S14, the classification unit 40 selects the class with the highest similarity from the similarity scores calculated for all classes. This determines the class of the input data, and the classification result for the input data class is output from the classification unit 40. After S14, the classification process by the classification device 1 is completed.

[0026] Figure 3 is a schematic functional block diagram of the machine learning device 30. As shown in Figure 3, the machine learning device 30 comprises a learning unit 31, a first connection adjustment unit 32, a determination unit 33, and a second connection adjustment unit 34. In the example of Figure 1, the machine learning device 30 is installed inside the classification device 1.

[0027] The learning unit 31 starts a learning phase, accepts the input of a dataset for each class, and trains the feature output unit 20 using that dataset. Each class has one or more learning phases, and a dataset is used for each learning phase. Each dataset contains a large number of samples. An example of a sample is an image, but it is not limited to this. If the sample is an image, one class may be about classifying images of dogs, and another class may be about classifying images of cats. After the learning unit 31 has trained the feature output unit 20 for one class, it then has it train for another class.

[0028] The first connection adjustment unit 32 adjusts the connection state between nodes of the feature output unit 20, that is, connects or disconnects nodes, when a certain learning phase is completed while the feature output unit 20 is learning a predetermined class, thereby forming a node path for the new class. The method for updating the path of the feature output unit 20 executed by the first connection adjustment unit 32 may be, for example, the method based on NISPA described in Non-Patent Document 1. For example, in NISPA, the connection state between nodes is adjusted based on the activation value of the nodes of the feature output unit 20. The activation value of a node is determined based on the activation value of the parent node connected in the layer immediately preceding that node, i.e., the layer closer to the input layer, and the weight of the connection with that parent node. For example, in NISPA, paths between nodes with high activation values ​​are connected, while paths to nodes with low activation values ​​are randomly disconnected.

[0029] The determination unit 33 determines whether the similarity between the adjusted features of the new class of test input data, calculated based on the output of the hidden layer in the neural network of the feature output unit 20, and the prototype is above a threshold.

[0030] The second connection adjustment unit 34, when it is determined that the similarity is above a threshold, sets the weights of the nodes in the path that are on the output side of the intermediate layer related to that determination to 0.

[0031] Figure 4 is a flowchart illustrating the learning process performed by the machine learning device 30. Hereinafter, the feature output unit 20 will be described as having a 15-layer neural network as an example.

[0032] In S101, the learning unit 31 trains the feature output unit 20 using the input training dataset for the new class as a task for the new class. The learning unit 31 updates the activation values ​​of the nodes in each layer of the feature output unit 20 through learning.

[0033] In S102, the first connection adjustment unit 32 adjusts the connection state between nodes of the feature output unit 20 based on the activation values ​​of the nodes. For example, the first connection adjustment unit 32 connects paths between nodes with high activation values ​​and randomly disconnects paths to nodes with low activation values.

[0034] In S103, the feature adjustment unit 25 calculates a prototype based on the features output from the feature output unit 20 when the training dataset for the learned task is input to the feature output unit 20, and saves the calculated prototype to the feature storage unit 26. Here, the prototype is the average value of the features. For example, if 100 images are used as input data for training, the prototype will be the average value of the features of the 100 images.

[0035] In S104, the feature adjustment unit 25 calculates the adjusted features of the test data based on the output of the nodes in the hidden layer, using a test dataset for a new class (hereinafter referred to as test data). As will be described later, steps S104 to S106 are repeatedly executed for adjacent hidden layers in predetermined cases. For example, in the first S104, the feature adjustment unit 25 first calculates the adjusted features of the test data based on the output of the nodes in the fifth layer, which is an intermediate layer. However, it is not limited to this, and for example, the learning unit 31 may use any hidden layer to execute the first S104.

[0036] In S105, the classification unit 40 calculates the similarity between the adjusted features calculated based on the output of the intermediate layer nodes and the prototype. Here, the prototype stored in the feature storage unit 26 in S103 is read out, and the similarity between the adjusted features calculated based on the output of the intermediate layer nodes and the read out prototype is calculated.

[0037] In S106, the determination unit 33 determines whether the similarity calculated by the classification unit 40 in S105 is equal to or greater than a threshold. If the similarity is equal to or greater than the threshold, the output of the intermediate layer node being determined in S105 approximates the prototype of the new task class, and therefore, sufficient classification accuracy can be considered to be obtained from the output of that intermediate layer node. On the other hand, if the similarity is not equal to or greater than the threshold, it can be considered that sufficient classification accuracy cannot be obtained from the output of that intermediate layer node.

[0038] If the similarity is not above the threshold (N in S106), the process proceeds to S107. In S107, the intermediate layer targeted in S104 to S106 is set as the layer adjacent to the output side of that intermediate layer. After S107, the process returns to S104. That is, adjustment features of the test data are calculated based on the output of the node of the next intermediate layer (S104), the similarity between these adjustment features and the prototype is calculated (S105), and it is determined whether the similarity is above the threshold (S106). This loop of processing from S104 to S107 is repeatedly executed until it is determined in S105 that the similarity calculated by the classification unit 40 is above the threshold.

[0039] On the other hand, if the similarity is above a threshold (Y in S106), the process proceeds to S108. In S108, the second connection adjustment unit 34 determines whether the intermediate layer is the intermediate layer closest to the output. If it is not the intermediate layer closest to the output (N in S108), that is, in this example, if the intermediate layer is not the 14th layer, the process proceeds to S109.

[0040] In S109, the second connection adjustment unit 34 sets the weights of the nodes in the layer that is on the output side of the intermediate layer in the path to 0. For example, if the intermediate layer is the 14th layer, the second connection adjustment unit 34 sets the weights of all nodes in the 15th layer of the feature output unit 20 to 0 in the path corresponding to the new class. As a result, the 15th layer is effectively disabled.

[0041] Figure 5 shows an example of connection adjustment processing. Figure 5(a) is a diagram showing an example of connection adjustment processing by NISPA, and Figure 5(b) is a diagram showing an example of connection adjustment processing according to the first embodiment. Specifically, Figure 5(a) schematically shows each node and its connection status in layers 12 to 15 of the feature output unit 20 when a certain learning phase of a certain class has been completed. Figure 5(b) schematically shows each node and its connection status in layers 12 to 15 of the feature output unit 20 after adjustment processing by the second connection adjustment unit 34 of the first embodiment. In Figures 5(a) and 5(b), solid lines indicate that nodes in adjacent layers are connected to each other, and dashed lines indicate that nodes in adjacent layers are not connected to each other.

[0042] Figures 5(a) and 5(b) show stable nodes 60 and plastic nodes 62. Stable nodes 60 and plastic nodes 62 are classified based on the activity value of the node. Stable nodes 60 are nodes whose activity value is greater than a predetermined reference value. Plastic nodes 62 are nodes whose activity value is less than or equal to a predetermined reference value. As shown in Figure 5(a), the stable node 60a of the 12th layer and the stable node 60b of the 13th layer, the stable node 60a of the 12th layer and the stable node 60c of the 13th layer, the stable node 60b of the 13th layer and the stable node 60d of the 14th layer, and the stable node 60d of the 14th layer and the stable node 60e of the 15th layer are connected to each other. These stable nodes 60a to 60e of the 12th to 15th layers constitute the node path R1.

[0043] Here, it is assumed that the similarity calculated based on the output of the nodes in the 13th layer is not equal to or greater than the threshold, and the similarity calculated based on the output of the nodes in the 14th layer is determined to be equal to or greater than the threshold. In this case, since the similarity calculated based on the output of each node in the 14th layer is equal to or greater than the threshold, the weight of the node in the 15th layer, which is the next layer, is set to 0. As a result, as shown in FIG. 5(b), the stable node 60e in the 15th layer shown in FIG. 5(a) is changed to a plastic node 62, and the connection between the stable node 60d in the 14th layer and the stable node 60e in the 15th layer is disconnected. As a result, a route R2 terminating at the stable node 60d in the 14th layer is formed.

[0044] Returning to S108, if the current layer is the output-side intermediate layer (Y in S108), the process proceeds to S110.

[0045] In S110, the learning unit 31 determines whether there is a task for the next new class. If there is a task for the next new class (Y in S110), the process returns to S101. If there is no task for the next new class (N in S110), the process ends.

[0046] In the embodiment, the intermediate layer to be processed in steps S104 to S106 in S107 is set to a layer adjacent to the output side of the intermediate layer, but the present invention is not limited thereto. For example, the intermediate layer may be set to a layer adjacent to the input side. In this case, for example, when it is determined in S106 that the similarity is equal to or greater than the threshold, the process may proceed to S107, and S104 to S107 may be repeatedly executed until it is determined in S106 that the similarity is not equal to or greater than the threshold. For example, when S104 is started from the 14th layer, and it is determined in S106 that the similarity between the adjusted feature amount based on the output of the 14th layer and the prototype is not equal to or greater than the threshold, the weight of the node in the 15th layer of the route corresponding to the new class may be set to 0.

[0047] In the example of FIG. 1, the machine learning device 30 is provided inside the classification device 1, but may be provided outside the classification device 1.

[0048] The operational effects of the embodiment will be described.

[0049] The machine learning device 30 of the first embodiment includes a learning unit 31 that trains a feature output unit that outputs feature quantities of training input data for a new class using a neural network; a first connection adjustment unit 32 that forms a node path for the new class by connecting or disconnecting nodes between adjacent layers in the neural network of the feature output unit 20; and a second connection adjustment unit 34 that sets the weights of nodes in layers on the output side of the intermediate layer in the path to 0 if the similarity between the adjustment feature, which represents the representative feature quantities of the test input data for the new class calculated based on the output of the intermediate layer in the neural network of the feature output unit 20, and the prototype, which represents the representative feature quantities of the training input data, is above a threshold. In NISPA, it is necessary to construct a path using nodes from each layer from the input layer to the output layer. Therefore, if many tasks for the new class are trained, it is easy to run out of nodes necessary for training. In contrast, according to the first embodiment, it is possible to optimize the length of the memory path for each class in NISPA by dynamically adjusting the number of layers required for class classification. Therefore, it is easier to retain more nodes that can be used for class classification than in NISPA. Furthermore, because the memory paths for each class tend to be shorter than in NISPA, the computational complexity required for classification can be reduced. As a result, the processing speed of classification can be improved.

[0050] In the first embodiment, if the similarity is not above a threshold, the second connection adjustment unit 34 sets the weight of the node of the layer that is on the output side of the path compared to other intermediate layers to 0 if the similarity between the adjustment feature calculated based on the output of other intermediate layers adjacent to the intermediate layer and the prototype is above a threshold. With this configuration, if it is determined that the similarity is not above a threshold, the similarity of the next intermediate layer is sequentially compared with the threshold until it is determined that the similarity is above a threshold, and an intermediate layer with sufficient classification accuracy can be efficiently searched. As a result, it is possible to reduce the amount of computation required for setting the path by the second connection adjustment unit 34.

[0051] The second embodiment of the present invention will be described below. In the drawings and description of the second embodiment, the same or equivalent components and members as in the first embodiment will be denoted by the same reference numerals. Descriptions that overlap with the first embodiment will be omitted as appropriate, and the description will focus on the configurations that differ from the first embodiment.

[0052] Figure 6 illustrates the connection adjustment process of the second embodiment. In the second embodiment, the first connection adjustment unit 32 forms paths so that all nodes between adjacent layers are connected from the input layer to a predetermined layer. As shown in Figure 6, in the second embodiment, from the input layer of the feature output unit 20, i.e., from the first layer to the predetermined k-th layer (k is a natural number of 2 or more), all nodes between adjacent layers are connected. This connection state may remain unchanged in the processing described later. Here, all nodes between adjacent layers being connected means that a node has a path between it and all nodes in the adjacent layer. The lower-order layers from the input layer to the predetermined layer are considered to be layers that transmit basic information of the input data that does not depend on the class being learned, such as the contour and color of image data. Therefore, by making the nodes of the lower-order layers on the input side fully connected in the feature output unit 20, the computational amount required to set the paths can be reduced, and the classification performance can be improved.

[0053] In the second embodiment, the learning unit 31 performs learning by fixing the weights of all nodes from the input layer to a predetermined layer. The neural networks constituting each layer from the input layer to the predetermined layer may be neural networks that have been pre-trained using big data. This configuration reduces the computational load required for learning. Furthermore, since basic information that is independent of the class being learned, such as contours and colors, is retained, it becomes possible to acquire knowledge of new classes while suppressing fatal forgetting caused by learning tasks of new classes.

[0054] The various processes of the classification device 1 and machine learning device 22 described above can, of course, be implemented as devices using hardware such as a CPU and memory, but they can also be implemented by firmware stored in ROM (read-only memory) or flash memory, or by software on a computer. The firmware program and software program can be recorded on a recording medium readable by a computer and provided, transmitted and received with a server via a wired or wireless network, or transmitted and received as data broadcasting on terrestrial or satellite digital broadcasting.

[0055] The present invention has been described above based on embodiments. The embodiments are illustrative, and it will be understood by those skilled in the art that various modifications are possible in combinations of their components and processing processes, and that such modifications also fall within the scope of the present invention.

[0056] This invention relates to machine learning technology.

[0057] 1 Classification device, 20 Feature output unit, 30 Machine learning device, 31 Learning unit, 32 First connection adjustment unit, 33 Judgment unit, 34 Second connection adjustment unit, 40 Classification unit, 50 Output unit.

Claims

1. A machine learning device comprising: a learning unit that trains a feature output unit that outputs feature quantities of input data for training a new class using a neural network; a first connection adjustment unit that forms a node path for the new class by connecting or disconnecting nodes between adjacent layers in the neural network of the feature output unit; and a second connection adjustment unit that sets the weights of nodes in layers on the output side of the path beyond the intermediate layer to 0 if the similarity between the adjustment feature, which represents the representative feature quantities of the test input data for the new class calculated based on the output of the intermediate layer in the neural network of the feature output unit, and the prototype, which represents the representative feature quantities of the training input data, is greater than or equal to a threshold.

2. The machine learning apparatus according to claim 1, wherein the second connection adjustment unit, if the similarity is not equal to or greater than the threshold, sets the weight of the node of the layer on the output side of the path than the other intermediate layer to 0 if the similarity between the adjustment feature calculated based on the output of other intermediate layers adjacent to the intermediate layer and the prototype is equal to or greater than the threshold.

3. The machine learning apparatus according to claim 1, wherein the first connection adjustment unit forms the path such that all nodes between adjacent layers are connected from the input layer to a predetermined layer.

4. The machine learning apparatus according to any one of claims 1 to 3, wherein the learning unit performs the learning by fixing the weights of all nodes between the input layer and a predetermined layer.

5. A machine learning method comprising: training a neural network that outputs features of input data using training input data of a new class; forming a node path for the new class by connecting or disconnecting nodes between adjacent layers in the neural network; generating a prototype that shows representative features of the training input data; calculating adjustment features that show representative features of the test input data of the new class based on the output of the hidden layer in the neural network; calculating the similarity between the adjustment features and the prototype; and, if the similarity is above a threshold, setting the weights of the nodes in the path that are on the output side of the hidden layer to 0.