Learning device
The learning device efficiently selects features to preserve order relationships in labels by imposing spatial constraints, enhancing model learning and estimation accuracy in applications such as OSNR estimation and healthcare.
Patent Information
- Application Number
- JP2024022934
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-19
- Publication Date
- 2025-08-29
AI Technical Summary
Existing machine learning methods struggle to efficiently extract features that preserve order relationships in labels, and selecting quadruplets for learning becomes computationally intensive, leading to inefficient learning processes.
A learning device and method that selects a set of features by imposing restrictions based on the distance between labels or features in a feature space, allowing for efficient learning that captures order relationships by training models using quadruplets.
The solution enables efficient and appropriate model learning that preserves order relationships in labels, reducing computational burden and improving estimation accuracy in applications like OSNR estimation and healthcare diagnostics.
Smart Images

Figure 2025126612000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a learning device, a learning method, and a program. [Background technology]
[0002] Techniques used in machine learning using features extracted from time-series data are known.
[0003] For example, Patent Document 1 discloses an apparatus for training a model for extracting features that preserve similarity relationships in time-series data for training. For example, according to Patent Document 1, the apparatus selects a triplet consisting of an anchor sample, a positive sample of the same class as the anchor sample, and a negative sample of a different class from the anchor sample. The apparatus then trains the model so that samples of the same class are close to each other in the feature space and samples of different classes are far apart in the feature space.
[0004] Another related technology is Patent Document 2, which discloses a device that performs learning so as to preserve the order relationship of labels. According to Patent Document 2, the device selects a quadruple consisting of an anchor sample, a positive sample, and negative sample 1 and negative sample 2, which have labels different from those of the anchor sample. The device then uses the selected quadruple to learn a model so as to minimize a predetermined loss. For example, this configuration makes it possible to extract features that capture the order when an order relationship exists between labels. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] International Publication No. 2020 / 049666 [Patent Document 2] Special Publication No. 2023-538190 Summary of the Invention [Problem to be solved by the invention]
[0006] When using the technology described in Patent Document 1, it is difficult to extract features that preserve the order relationships present in the labels. Therefore, when preserving the order relationships, it is desirable to use the technology described in Patent Document 2 rather than Patent Document 1. On the other hand, when attempting to select pairs such as quadruplets as described in Patent Document 2, the number of combinations of pairs to be selected during learning becomes enormous. As a result, there is a problem that learning takes a long time, and it may be difficult to perform appropriate learning when there is a predetermined relationship, such as an order relationship, among the labels.
[0007] Therefore, one of the objects of the present invention is to provide a learning method, a learning device, and a program that can solve the above-mentioned problems. [Means for solving the problem]
[0008] In order to achieve this object, a learning device according to one embodiment of the present disclosure includes: an extraction unit that extracts features according to input data; a selection unit that selects a set of features to be used when training a model from a feature group consisting of the plurality of feature values extracted by the extraction unit; a learning unit that learns a model using the set of features selected by the selection unit; and The selection unit selects a set of features from the group of features such that at least one of the distance between labels assigned to the features or the distance between the features in a feature space satisfies a predetermined condition. The structure is as follows.
[0009] Furthermore, a learning method according to another aspect of the present disclosure includes: The information processing device Extract features based on the input data, A set of features to be used when training a model is selected from a group of extracted features. Train a model using the selected feature set; When selecting a set of feature quantities, a set of feature quantities is selected from the group of feature quantities such that at least one of the distance between the labels assigned to the feature quantities or the distance between the feature quantities in the feature space satisfies a predetermined condition. The structure is as follows.
[0010] Furthermore, a program according to another aspect of the present disclosure includes: In the information processing device, Extract features based on the input data, A set of features to be used when training a model is selected from a group of extracted features. Train a model using the selected feature set; When selecting a set of feature quantities, a set of feature quantities is selected from the group of feature quantities such that at least one of the distance between the labels assigned to the feature quantities or the distance between the feature quantities in the feature space satisfies a predetermined condition. It is a program for realizing the processing. [Effects of the Invention]
[0011] According to the above-mentioned configurations, the above-mentioned problems can be solved. [Brief explanation of the drawings]
[0012] [Figure 1] 1 is a diagram illustrating an overview of a learning device according to a first embodiment of the present disclosure. [Figure 2] FIG. 1 is a diagram illustrating an overview of a learning device. [Figure 3] FIG. 2 is a block diagram illustrating an example of the configuration of a learning device. [Figure 4] FIG. 10 is a diagram illustrating an example of processing by a feature amount selection unit. [Figure 5] FIG. 10 is a diagram illustrating another example of processing by the feature amount selection unit. [Figure 6] 10 is a flowchart illustrating an example of the operation of the learning device. [Figure 7] FIG. 10 is a block diagram showing another example configuration of the learning device. [Figure 8] FIG. 10 is a diagram illustrating an application example of a learning device. [Figure 9] FIG. 10 is a diagram illustrating an application example of a learning device. [Figure 10] FIG. 10 is a diagram illustrating an example of estimation when the learning device described in the present disclosure is not used. [Figure 11] FIG. 10 is a diagram illustrating an example of estimation when using the learning device described in the present disclosure. [Figure 12] FIG. 10 is a diagram illustrating another application example of the learning device. [Figure 13] FIG. 10 is a diagram illustrating an example of a hardware configuration of a learning device according to a second embodiment of the present disclosure. [Figure 14] FIG. 2 is a block diagram illustrating an example of the configuration of a learning device. [Figure 15] 10 is a flowchart illustrating an example of the operation of the learning device. DETAILED DESCRIPTION OF THE INVENTION
[0013] [First embodiment] A first embodiment of the present invention will be described with reference to FIGS. 1 to 12. FIGS. 1 and 2 are diagrams for explaining an overview of a learning device 100. FIG. 3 is a block diagram showing an example of the configuration of the learning device 100. FIGS. 4 and 5 are diagrams showing an example of processing by the feature selection unit 154. FIG. 6 is a flowchart showing an example of operation of the learning device 100. FIG. 7 is a block diagram showing another example of the configuration of the learning device 100. FIGS. 8 and 9 are diagrams showing an example of application of the learning device 100. FIG. 10 is a diagram showing an example of estimation when the learning device 100 is not used. FIG. 11 is a diagram showing an example of estimation when the learning device 100 is used. FIG. 12 is a diagram showing another example of application of the learning device 100. Note that in the present disclosure, the drawings may be associated with one or more embodiments.
[0014] In a first embodiment of the present disclosure, a learning device 100 capable of performing learning that captures an order relationship when labels have an order relationship will be described. For example, there may be an order relationship between labels, such as "bad," "good," and "best" as shown in FIG. 1 . In such a case, the learning device 100 selects a quadruple consisting of an anchor sample, a positive sample with the same label as the anchor sample, and negative sample 1 and negative sample 2 with labels different from those of the anchor sample. The learning device 100 then uses the selected quadruple to train a model so as to minimize a predetermined loss. For example, the learning device 100 trains a model so that negative sample 1, which has a different label from the anchor sample but is closer to the positive sample than negative sample 2, is located farther away in the feature space. Furthermore, the learning device 100 trains a model so that negative sample 2, whose label is farther away from the positive sample than negative sample 1, is located farther away in the feature space than negative sample 1. For example, by performing such training, the learning device 100 performs learning that can preserve the order relationship between labels.
[0015] Furthermore, when selecting the above-described quadruples, the learning device 100 according to the present disclosure selects from among targets that satisfy a predetermined condition. For example, the learning device 100 selects quadruples, such as negative sample 1 and negative sample 2, from among feature quantities whose label difference from the anchor sample is within a predetermined value d. In other words, the learning device 100 can select quadruples after imposing a restriction based on the distance between labels. Furthermore, instead of the above-described selection based on the distance between labels, the learning device 100 may select quadruples, such as negative sample 1 and negative sample 2, from among feature quantities whose distance from the positive sample in the feature space is within a predetermined value δ. In other words, the learning device 100 can select quadruples after imposing a restriction based on the distance between features. For example, by performing selection based on the above-described restrictions based on the distance between labels or the distance between features, the learning device 100 can efficiently select quadruples by narrowing down the selection targets. Note that the learning device 100 may be configured to determine the method used to select quadruples according to predetermined conditions. For example, the learning device 100 can be configured to determine a method for selecting a quadruple depending on whether or not there is a predetermined relationship between the label and the input data, such as whether the data is similar when the labels are similar.
[0016] After selecting quadruples with the above-described method and imposing predetermined restrictions, the learning device 100 can use any method to train so that negative sample 1 is farther away and negative sample 2 is farther away in the feature space (see FIG. 2). For example, the learning device 100 may perform training using a method such as that described in Patent Document 2. As an example, the learning device 100 can train a model so that the loss shown in Equation 1 is minimized.
number
number
[0017] A more specific configuration example of the learning device 100 will be described below. As described above, the learning device 100 is an information processing device capable of learning a model by using a supervised metric learning method. For example, the learning device 100 acquires a time-series dataset by acquiring time-series data from various sensors and other devices. The learning device 100 also learns or updates a model based on a time-series dataset consisting of one or more pieces of time-series data.
[0018] Fig. 3 shows an example configuration of learning device 100. Referring to Fig. 3, learning device 100 has, as main components, for example, an operation input unit 110, a screen display unit 120, a communication I / F unit 130, a memory unit 140, and an arithmetic processing unit 150.
[0019] 3 illustrates an example in which the functions of learning device 100 are realized using a single information processing device. However, at least some of the functions of learning device 100 may be realized using multiple information processing devices, for example, on the cloud. Furthermore, learning device 100 may not include some of the components illustrated above, such as not having operation input unit 110 or screen display unit 120, or may have components other than those illustrated above.
[0020] Operation input unit 110 is made up of operation input devices such as a keyboard, a mouse, etc. Operation input unit 110 detects operations by the operator operating learning device 100 and outputs the operations to calculation processing unit 150.
[0021] The screen display unit 120 is composed of a screen display device such as a liquid crystal display, an organic EL (electro-luminescence) display, etc. The screen display unit 120 can display various information stored in the storage unit 140 on the screen in response to instructions from the arithmetic processing unit 150.
[0022] The communication I / F unit 130 is composed of a data communication circuit, etc. The communication I / F unit 130 performs data communication with various sensors and other external devices connected via communication lines.
[0023] The storage unit 140 is a storage device such as a hard disk or a memory. The storage unit 140 stores processing information and a program 144 required for various processes in the arithmetic processing unit 150. The program 144 is read into the arithmetic processing unit 150 and executed to realize various processing units. The program 144 is read in advance from an external device or a recording medium via a data input / output function such as the communication I / F unit 130, and is stored in the storage unit 140. Main information stored in the storage unit 140 includes, for example, model information 141, feature amount information 142, and time-series data information 143.
[0024] The model information 141 includes information about a model that extracts and outputs features that preserve local distance relationships for input time-series segments. For example, the model information 141 may include weight parameters included in the trained model described above. For example, the model included in the model information 141 is trained in advance using training time-series segments inside or outside the training device 100 and stored in the storage unit 140. The model information 141 is updated through training by the model training unit 155, which will be described later.
[0025] The feature information 142 includes information corresponding to features extracted from the time-series segments. For example, the feature information 142 includes binary codes obtained by converting features using any method. In the feature information 142, the binary codes may be associated with labels to be classified. For example, the binary codes included in the feature information 142 are obtained by converting the features extracted by the feature extraction unit 152 using the binary code conversion unit 153, and are stored in the storage unit 140.
[0026] The feature amount information included in the feature amount information 142 can be used when performing a search process, which will be described later. For example, the results of the search process may be used when determining whether or not an abnormality exists or to determine the label to which the abnormality belongs. The results of the search process may also be used for purposes other than those exemplified above.
[0027] The time-series data information 143 includes a time-series data set consisting of one or more pieces of time-series data. Here, the time-series data may be data in which numerical data such as observation data measured by a sensor at a predetermined interval is arranged in order of measurement time. The time-series data information 143 is updated as the time-series data acquisition unit 151 (described later) acquires time-series data from sensors and other devices.
[0028] The arithmetic processing unit 150 has an arithmetic device such as a CPU (Central Processing Unit) and its peripheral circuits. The arithmetic processing unit 150 reads and executes a program 144 from the storage unit 140, thereby causing the above hardware and the program 144 to work together to realize various processing units. Major processing units realized by the arithmetic processing unit 150 include, for example, a time-series data acquisition unit 151, a feature extraction unit 152, a binary code conversion unit 153, a feature selection unit 154, and a model learning unit 155.
[0029] In addition, the arithmetic processing unit 150 may have a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), an MPU (Micro Processing Unit), an FPU (Floating point number Processing Unit), a PPU (Physics Processing Unit), a TPU (Tensor Processing Unit), a quantum processor, a microcontroller, or a combination thereof, instead of the above-mentioned CPU.
[0030] The time-series data acquiring unit 151 acquires a time-series data set by acquiring time-series data from various sensors and other devices. For example, the time-series data acquiring unit 151 can acquire time-series data from an optical transponder, an optical performance monitor that measures an optical signal-to-noise ratio (OSNR), a plurality of sensors installed in a plant or a data center, various healthcare-related sensors for measuring blood pressure, heart rate, etc., and any other device.
[0031] Furthermore, the time series data acquisition unit 151 stores the acquired time series data set in the storage unit 140 as time series data information 143 .
[0032] The feature extraction unit 152 extracts features from the time series segments by inputting the time series segments to a model stored as model information 141. For example, the feature extraction unit 152 divides the time series data set acquired by the time series data acquisition unit 151 into multiple time series segments using a time window of a certain period. Then, the feature extraction unit 152 inputs each of the divided time series segments to a trained model to extract features. Note that the feature extraction unit 152 may divide the time series data into the time series segments using any method. For example, the size of the time window may be set arbitrarily. Furthermore, the feature extraction unit 152 may divide the time series data into multiple time series segments so that they overlap for an arbitrary period, or may divide the time series data into multiple time series segments so that the time series segments do not overlap.
[0033] The binary code conversion unit 153 converts the features extracted by the feature extraction unit 152 into binary code, which is information corresponding to the features. In the present disclosure, the method of conversion into binary code is not particularly limited. The binary code conversion unit 153 may convert the features extracted by the feature extraction unit 152 into binary code using any method. Furthermore, the binary code conversion unit 153 can store the converted binary code in the storage unit 140 as feature information 142.
[0034] Furthermore, the feature extraction unit 152 or the binary code conversion unit 153 can assign labels to the features or binary codes to be classified using any method. For example, the feature extraction unit 152 or the binary code conversion unit 153 may assign labels to the features, etc., determined based on the Euclidean distance between the time-series segment from which the features are extracted and a pre-stored time-series segment to be compared. Information about the labels may be acquired based on an operation by the operator on the operation input unit 110, etc.
[0035] The feature selection unit 154 selects a set of features, such as a quadruple, to be used for learning from a group of features made up of multiple features extracted by the feature extraction unit 152. As described above, the feature selection unit 154 selects quadruples after imposing predetermined restrictions such as the distance between labels and the distance between features. The feature selection unit 154 can select multiple quadruples by repeating the above selection.
[0036] For example, as shown in FIG. 4, the feature selection unit 154 can select quadruples after imposing a restriction based on the distance between labels. Specifically, the feature selection unit 154 selects an anchor sample by randomly selecting a feature from a feature group including multiple feature groups extracted by the feature extraction unit 152, and selects a positive sample from feature groups to which the same label as the anchor sample is assigned. The feature selection unit 154 also selects negative sample 1 and negative sample 2 from feature groups whose label difference from the selected anchor sample is within a predetermined value d. That is, the feature selection unit 154 selects negative sample 1 and negative sample 2 that satisfy the formula shown in Equation 3. At this time, the feature selection unit 154 performs the above selection so that the label difference between the anchor sample and negative sample 2 is greater than that of negative sample 1. That is, the feature selection unit 154 selects negative sample 1 and negative sample 2 that satisfy the above restriction and are assigned different labels.
number
[0037] Furthermore, as shown in FIG. 5, the feature selection unit 154 may select quadruples after imposing a restriction based on the distance between features in the feature space. Specifically, the feature selection unit 154 selects an anchor sample by randomly selecting a feature from a feature group including multiple feature groups extracted by the feature extraction unit 152, and selects a positive sample from among the feature groups to which the same label as the anchor sample is assigned. The feature selection unit 154 also selects negative sample 1 and negative sample 2 from among the feature groups whose distance from the positive sample in the feature space is within a predetermined value δ. In other words, the feature selection unit 154 selects negative sample 1 and negative sample 2 that satisfy the formula shown in Equation 4. In this case, the feature selection unit 154 performs the above selection so that the label difference between the anchor sample and negative sample 1 is greater than the label of negative sample 2. In other words, the feature selection unit 154 selects negative sample 1 and negative sample 2 that satisfy the above restriction and are assigned different labels.
number
[0038] For example, the feature selection unit 154 can select quadruples using one or both of the above-described methods. The feature selection unit 154 may be configured to determine which of the above two methods to use depending on whether or not a predetermined relationship exists between the labels and the time-series data that is the input data. For example, when it is determined that a predetermined relationship exists between the labels and the input data, such as when the time-series data is close when the labels are close, the feature selection unit 154 can be configured to select quadruples after imposing a restriction based on the distance between the labels. Furthermore, when it is not determined that a predetermined relationship exists between the labels and the input data, the feature selection unit 154 can be configured to select quadruples after imposing a restriction based on the distance between the features in the feature space. Note that the feature selection unit 154 may determine the method to use when selecting quadruples according to conditions other than those exemplified above. Furthermore, the feature selection unit 154 may determine the method to use when selecting quadruples according to an operation using the operation input unit 110 or an instruction from an external device.
[0039] When performing a restriction based on the distance between labels, it is desirable that the feature selection unit 154 selects negative sample 1 and negative sample 2 whose labels are distant in the same direction from the anchor sample. For example, when there are labels such as "1," "2," "3," "4," and "5," and the label of the anchor sample is "3," it is desirable that the feature selection unit 154 selects negative sample 1 and negative sample 2 from among feature quantities assigned with labels whose values are smaller than the label "3," or selects negative sample 1 and negative sample 2 from among feature quantities assigned with labels whose values are larger than the label "3."
[0040] The model learning unit 155 performs learning using the quadruples selected by the feature selection unit 154, thereby updating the model stored as the model information 141. As described above, the model learning unit 155 may learn the model by using a supervised metric learning method such as that described in Patent Document 2.
[0041] The above is an example of the configuration of the learning device 100.
[0042] The model learned by the model learning unit 155 may be any model that can handle time-series data. For example, the model may be any of a 1D-CNN (1 Dimensional-Convolutional Neural Network), a GRU (Gated Recurrent Unit), an LSTM (Long Short-Term Memory), a Transformer, etc.
[0043] Furthermore, the features used to select quadruples after imposing restrictions based on the distance between features do not necessarily have to be features extracted from the latest current model.
[0044] Next, an example of the operation of the learning device 100 will be described with reference to FIG.
[0045] Fig. 6 shows an example of the operation of the learning device 100. Referring to Fig. 6, the time-series data acquisition unit 151 acquires time-series data from various sensors and other devices included in the learning device 100, thereby acquiring a time-series data set (step S101).
[0046] The feature extraction unit 152 extracts features based on the time series segments by inputting the time series segments to the model stored as the model information 141 (step S102). For example, the feature extraction unit 152 divides the time series data set acquired by the time series data acquisition unit 151 into multiple time series segments using a time window of a certain period. Then, the feature extraction unit 152 inputs each divided time series segment to a trained model to extract features.
[0047] The feature selection unit 154 selects quadruples to be used for learning from a group of features made up of multiple feature amounts extracted by the feature extraction unit 152 (step S103). As described above, the feature selection unit 154 selects quadruples after imposing predetermined restrictions such as the distance between labels and the distance between feature amounts.
[0048] The model learning unit 155 performs learning using the quadruples selected by the feature selection unit 154, thereby updating the model stored as the model information 141 (step S104). As described above, the model learning unit 155 may learn the model by using a supervised metric learning method such as that described in Patent Document 2.
[0049] As described above, the learning device 100 includes a feature selection unit 154 and a model learning unit 155. With this configuration, the model learning unit 155 can learn a model using quadruples efficiently selected by the feature selection unit 154. As a result, more efficient and appropriate model learning can be achieved. Furthermore, as the number of labels increases, there is a higher possibility that quadruples that do not contribute to learning will be selected, for example, when the labels of both negative samples are significantly different from the label of the anchor sample and the labels of the negative samples are also significantly different from each other. Even in such cases, the feature selection unit 154 can be used to make more appropriate selections and achieve learning that captures the order of the labels.
[0050] The configuration of the learning device 100 is not limited to the configuration exemplified with reference to Fig. 3. For example, the learning device 100 may be configured with a portion of the configuration exemplified in Fig. 3, such as omitting the binary code conversion unit 153. Fig. 7 shows another example configuration of the learning device 100. Referring to Fig. 7, for example, the arithmetic processing unit 150 can implement a search unit 156 and an output unit 157 in addition to the configuration exemplified in Fig. 3 by reading and executing the program 144 from the storage unit 140.
[0051] The search unit 156 performs search processing based on the feature amount extracted from the time-series segment to be searched.
[0052] For example, when the time-series data acquisition unit 151 acquires a time-series data set to be searched, the feature extraction unit 152 divides the time-series data set to be searched into multiple time-series segments and extracts features. The binary code conversion unit 153 converts the features extracted by the feature extraction unit 152 into binary codes, which are information corresponding to the features. The search unit 156 acquires the binary codes converted as described above. The search unit 156 then searches the feature information 142 for binary codes similar to the acquired binary code. For example, the search unit 156 calculates the distance between the acquired binary code and each binary code included in the feature information 142. The search unit 156 then acquires a binary code that satisfies an arbitrary condition, such as a binary code with the smallest calculated distance, from among the binary codes included in the feature information 142, as a binary code similar to the acquired binary code.
[0053] The search unit 156 can also perform various processes according to the search results. For example, the search unit 156 can infer the label of the time-series segment to be searched by identifying a label associated with the binary code acquired by the search. The search unit 156 can also detect an anomaly according to the search results. For example, a binary code corresponding to time-series data at the time of an anomaly is stored in advance in the feature information 142. The search unit 156 calculates an anomaly score according to the distance between the binary code to be searched and the searched binary code. The search unit 156 can then detect an anomaly based on the calculated anomaly score. For example, the search unit 156 can detect an anomaly based on a comparison result between the calculated anomaly score and a predetermined threshold. The search unit 156 can also perform processes according to search results other than those exemplified above.
[0054] The output unit 157 outputs the search results obtained by the search unit 156. For example, the output unit 157 may output the estimated label, anomaly score, or the like. The output unit 157 can display the search results on the screen display unit 120 or transmit the search results to an external device via the communication I / F unit. Note that the output unit 157 may be configured to output information stored in the storage unit 140, such as the feature amount information 142, in addition to the search results.
[0055] As described above, the learning device 100 can extract features and select quadruples from time-series data acquired from optical transponders, optical performance monitors that measure optical signal-to-noise ratios, etc., multiple sensors installed in plants and data centers, various healthcare-related sensors for measuring blood pressure, heart rate, etc., and any other devices. For example, as shown in FIG. 8, the learning device 100 may be configured to acquire time-series data from optical transponders and optical performance monitors included in an optical network 200. As an example, as shown in FIG. 9, the learning device 100 can be communicatively connected to an optical transponder 210 that receives an optical signal via a connected optical fiber and converts it into an electrical signal. In this case, the time-series data acquisition unit 151 of the learning device 100 can acquire time-series data corresponding to the intensity of the optical signal acquired by a sensor included in the optical transponder 210. In addition to the example shown in FIG. 9, the optical transponder 21 itself may have the function of the learning device 100 illustrated in FIG. 3, etc.
[0056] 10 and 11 show example estimation results when OSNR, which is a label, is estimated based on time-series data acquired from the optical transponder 210. FIG. 10 shows an example estimation result when the learning device 100 is not applied (when no selection is made by the feature selection unit 154), and FIG. 11 shows an example estimation result when the learning device 100 is applied (when selection is made by the feature selection unit 154). Referring to FIGS. 10 and 11, it can be seen that application of the learning device 100 improves estimation performance, particularly in regions of 30 dB or higher. In other words, use of the learning device 100 can improve performance when making ordered estimations such as OSNR estimation, thereby realizing more appropriate learning for making more appropriate estimations.
[0057] 12, the learning device 100 may be applied to the field of healthcare. For example, the learning device 100 can be configured to acquire time-series data from various healthcare-related biosensors 300, such as a blood pressure monitor or a heart rate monitor, owned by a subject such as a patient. In this case, the search unit 156 can be configured to infer labels having an order relationship according to the degree of health, the possibility of illness, and the like. By using the learning device 100 to make such inferences, more accurate inferences can be realized, and more appropriate support can be provided for decision-making by doctors and patients.
[0058] As described above, the learning device 100 can be applied to various situations. For example, the learning device 100 may be applied to situations other than the above-described examples, such as a plant.
[0059] In this disclosure, the case of selecting a quadruple has been exemplified. However, the present invention may be applied to cases other than selecting a quadruple. For example, the present invention may be applied to a learning device that selects triples as disclosed in Patent Document 1, or a learning device that selects any combination of quintuples or more. Furthermore, the present invention is not limited to cases where there is an order relationship between labels, and can also be applied to cases where there is some relationship between labels, such as when there is a similarity relationship between labels.
[0060] Furthermore, in this disclosure, the case where features are extracted from time-series data has been described. However, the present invention may also be applied to the case where features are extracted from any data other than time-series data. In other words, the learning device 100 may be configured to extract features in response to input of any data other than time-series data.
[0061] [Second embodiment] Next, a second embodiment of the present disclosure will be described with reference to Fig. 13 to Fig. 15. Fig. 13 is a diagram illustrating an example of the hardware configuration of a learning device 400. Fig. 14 is a block diagram illustrating an example of the configuration of the learning device 400. Fig. 15 is a flowchart illustrating an example of the operation of the learning device 400.
[0062] In the second embodiment of the present disclosure, a learning device 400 will be described, which is an information processing device that performs machine learning by selecting combinations such as quadruplets from a group of features extracted from input data. Fig. 13 shows an example of the hardware configuration of the learning device 400. Referring to Fig. 13, the learning device 400 has, as an example, the following hardware configuration. ·CPU(Central Processing Unit)401(Arithmetic unit) ROM (Read Only Memory) 402 (storage device) RAM (Random Access Memory) 403 (storage device) Programs 404 loaded into RAM 403 A storage device 405 for storing the program group 404 A drive device 406 that reads and writes data from a recording medium 410 outside the information processing device A communication interface 407 for connecting to a communication network 411 outside the information processing device Input / output interface 408 for inputting and outputting data Bus 409 connecting each component
[0063] 14 by the CPU 401 acquiring and executing the program group 404. The program group 404 is stored in advance in the storage device 405 or the ROM 402, for example, and is loaded into the RAM 403 or the like by the CPU 401 for execution as needed. The program group 404 may be supplied to the CPU 401 via the communication network 411, or may be stored in advance in the recording medium 410, with the drive device 406 reading out the programs and supplying them to the CPU 401.
[0064] 13 shows an example of the hardware configuration of the learning device 400. The hardware configuration of the learning device 400 is not limited to the above. For example, the learning device 400 may be configured with only a part of the above configuration, such as excluding the drive device 406. Furthermore, the CPU 401 may be a GPU, as exemplified in the first embodiment.
[0065] The extraction unit 421 extracts features according to input data. For example, the extraction unit 421 can extract features by inputting input data such as time-series data acquired by various sensors into a trained model.
[0066] The selection unit 422 selects a set of features to be used when training a model from a group of features made up of multiple features extracted by the extraction unit 421. For example, the selection unit 422 can select a quadruple made up of an anchor sample, a positive sample, a first negative sample, and a second negative sample.
[0067] The learning unit 423 learns a model using the set of feature amounts selected by the selection unit 422. For example, the learning unit 423 may learn a model so as to minimize loss as described in Patent Document 2.
[0068] The above is an example of the configuration of the learning device 400. Next, an example of the operation of the learning device 400 will be described with reference to FIG.
[0069] Fig. 15 is a flowchart showing an example of the operation of the learning device 400. Referring to Fig. 15, the extraction unit 421 extracts a feature amount according to input data (step S201).
[0070] The selection unit 422 selects a set of features to be used when training a model from a group of features made up of multiple features extracted by the extraction unit 421 (step S202). For example, the selection unit 422 can select a quadruple made up of an anchor sample, a positive sample, a first negative sample, and a second negative sample.
[0071] The learning unit 423 learns a model using the set of features selected by the selection unit 422 (step S203).
[0072] The above is an example of the operation of the learning device 400.
[0073] As described above, the learning device 400 includes the selection unit 422 and the learning unit 423. With this configuration, the learning unit 423 can learn a model using the set of features selected by the selection unit 422. As a result, more efficient and appropriate model learning can be achieved.
[0074] The above-described learning device 400 can be realized by incorporating a predetermined program into an information processing device such as the learning device 400. Specifically, a program according to another aspect of the present invention is a program for causing an information processing device such as the learning device 400 to perform the following processing: extracting features according to input data, selecting a set of features to be used when training a model from a feature group consisting of the extracted features, training the model using the selected set of features, and selecting a set of features from the feature group such that at least one of the distance between labels assigned to the features or the distance between the features in the feature space satisfies a predetermined condition.
[0075] Furthermore, a learning method executed by an information processing device such as the learning device 400 described above is a method of extracting features according to input data, selecting a set of features to be used when training a model from a group of features consisting of the extracted multiple features, training the model using the selected set of features, and selecting the set of features from the group of features such that at least one of the distance between labels assigned to the features or the distance between features in the feature space satisfies a predetermined condition.
[0076] Even if the invention is a program having the above-mentioned configuration, or a computer-readable recording medium having the program recorded thereon, or a learning method, it can achieve the same functions and effects as the above-mentioned learning device 400, and therefore can achieve the above-mentioned objective of the present disclosure.
[0077] <Additional Notes> A part or all of the above-described embodiments can be described as follows: The learning device and the like according to the present invention will be outlined below. However, the present invention is not limited to the following configuration.
[0078] (Appendix 1) an extraction unit that extracts features according to input data; a selection unit that selects a set of features to be used when training a model from a feature group consisting of the plurality of feature values extracted by the extraction unit; a learning unit that learns a model using the set of features selected by the selection unit; and The selection unit selects a set of features from the group of features such that at least one of the distance between labels assigned to the features or the distance between the features in a feature space satisfies a predetermined condition. Learning device. (Appendix 2) 10. The learning device according to claim 1, The selection unit selects a set of feature amounts from the group of feature amounts by applying a restriction based on a distance between labels assigned to each feature amount. Learning device. (Appendix 3) 3. The learning device according to claim 2, The selection unit selects an anchor sample from the feature group, and also selects at least one negative sample from the feature group to which a label has been assigned whose difference with the label assigned to the anchor sample is within a predetermined value and which has a different label from that of the selected anchor sample. Learning device. (Appendix 4) 4. The learning device according to claim 3, The selection unit selects a first negative sample and a second negative sample that are assigned labels whose difference from the labels assigned to the anchor samples is within a predetermined value from among the feature amounts included in the feature group and that have labels different from those of the selected anchor samples. Learning device. (Appendix 5) 5. The learning device according to claim 4, The selection unit selects a first negative sample and a second negative sample whose labels are spaced apart in the same direction from the anchor sample. Learning device. (Appendix 6) 10. The learning device according to claim 1, wherein: The selection unit determines whether to select a set of feature quantities in which the distance between the labels satisfies a condition or a set of feature quantities in which the distance between the feature quantities in the feature space satisfies a condition, depending on the relationship between the input data and the label. Learning device. (Appendix 7) 7. The learning device according to claim 6, The selection unit determines to select the set of feature quantities in which a distance between labels satisfies a condition when there is a predetermined relationship between the input data and the labels. Learning device. (Appendix 8) 10. The learning device according to claim 1, wherein: The selection unit selects, from the feature group, a quadruple consisting of an anchor sample, a positive sample assigned the same label as the anchor sample, and a first negative sample and a second negative sample assigned a label different from that of the anchor sample, where at least one of the distance between the labels or the distance between the features in a feature space satisfies a predetermined condition. Learning device. (Appendix 9) The information processing device Extract features based on the input data, A set of features to be used when training a model is selected from a group of extracted features. Train a model using the selected feature set; When selecting a set of feature quantities, a set of feature quantities is selected from the group of feature quantities such that at least one of the distance between the labels assigned to the feature quantities or the distance between the feature quantities in the feature space satisfies a predetermined condition. Learning device. (Appendix 10) In the information processing device, Extract features based on the input data, A set of features to be used when training a model is selected from a group of extracted features. Train a model using the selected feature set To realize the processing, When selecting a set of feature quantities, a set of feature quantities is selected from the group of feature quantities such that at least one of the distance between the labels assigned to the feature quantities or the distance between the feature quantities in the feature space satisfies a predetermined condition. program.
[0079] Note that some or all of the configurations described in Supplementary Notes 2 to 8 that are dependent on the learning device described in Supplementary Note 1 may also be dependent in a similar dependent relationship on the learning method described in Supplementary Note 9 and the program described in Supplementary Note 10. Furthermore, not limited to Supplementary Notes 9 and 10, some or all of the configurations described as Supplements may also be dependent on various hardware, software, various recording means for recording software, or systems within the scope of the above-mentioned embodiments.
[0080] The programs described in the above embodiments and appendices may be stored in a storage device or a computer-readable recording medium, such as a portable medium such as a flexible disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0081] Although the present invention has been described above with reference to the above-mentioned embodiments, the present invention is not limited to the above-mentioned embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present invention within the scope of the present invention. [Explanation of symbols]
[0082] 100 Learning Device 110 Operation input section 120 Screen display section 130 Communication I / F section 140 Storage section 141 Model Information 142 Feature Information 143 Time Series Data Information 144 Programs 150 Processing unit 151 Time series data acquisition unit 152 Feature Extraction Unit 153 Binary Code Conversion Unit 154 Feature Selection Unit 155 Model Learning Department 156 Search Department 157 Output section 200 Optical Network 210 Optical Transponder 300 Biometric Sensor 400 Learning Device 401 CPU 402 ROM 403 RAM 404 Programs 405 Storage device 406 Drive Unit 407 Communication Interface 408 Input / Output Interface 409 Bus 410 Recording Media 411 Communication Network 421 Extraction part 422 Selection Section 423 Learning Department
Claims
1. an extraction unit that extracts features according to input data; a selection unit that selects a set of features to be used when training a model from a feature group consisting of the plurality of feature values extracted by the extraction unit; a learning unit that learns a model using the set of features selected by the selection unit; and The selection unit selects a set of features from the group of features such that at least one of the distance between labels assigned to the features or the distance between the features in a feature space satisfies a predetermined condition. Learning device.
2. The learning device according to claim 1 , The selection unit selects a set of feature amounts from the group of feature amounts by applying a restriction based on a distance between labels assigned to each feature amount. Learning device.
3. The learning device according to claim 2, The selection unit selects an anchor sample from the feature group, and also selects at least one negative sample from the feature group to which a label is assigned whose difference with the label assigned to the anchor sample is within a predetermined value and which has a different label from that of the selected anchor sample. Learning device.
4. The learning device according to claim 3, The selection unit selects a first negative sample and a second negative sample that are assigned labels whose difference with the labels assigned to the anchor samples is within a predetermined value from among the feature amounts included in the feature group and that have labels different from those of the selected anchor samples. Learning device.
5. The learning device according to claim 4, The selection unit selects a first negative sample and a second negative sample whose labels are spaced apart in the same direction from the anchor sample. Learning device.
6. The learning device according to claim 1 , The selection unit determines whether to select a set of feature quantities in which the distance between the labels satisfies a condition or a set of feature quantities in which the distance between the feature quantities in the feature space satisfies a condition, depending on the relationship between the input data and the label. Learning device.
7. The learning device according to claim 6, The selection unit determines to select the set of feature quantities in which a distance between labels satisfies a condition when there is a predetermined relationship between the input data and the labels. Learning device.
8. The learning device according to claim 1 , The selection unit selects, from the group of features, a quadruple consisting of an anchor sample, a positive sample assigned the same label as the anchor sample, and a first negative sample and a second negative sample assigned a label different from that of the anchor sample, where at least one of the distance between the labels or the distance between the features in a feature space satisfies a predetermined condition. Learning device.
9. The information processing device Extract features based on the input data, A set of features to be used when training a model is selected from a group of extracted features. Train a model using the selected feature set; When selecting a set of feature quantities, a set of feature quantities is selected from the group of feature quantities such that at least one of the distance between the labels assigned to the feature quantities or the distance between the feature quantities in the feature space satisfies a predetermined condition. Learning device.
10. In the information processing device, Extract features based on the input data, A set of features to be used when training a model is selected from a group of extracted features. Train a model using the selected feature set To realize the processing, When selecting a set of feature quantities, a set of feature quantities is selected from the group of feature quantities such that at least one of the distance between the labels assigned to the feature quantities or the distance between the feature quantities in the feature space satisfies a predetermined condition. program.
Citation Information
Patent Citations
Classification of ordinal time series with missing information
JP2023538190A
Time-series data processing device
WO2020049666A1