Learning data generation system, learning system, learning data generation method, learning method, and program
The training data generation system addresses the challenge of matching unknown wireless devices by clustering and generating a relationship matrix to enhance the accuracy of device recognition through improved training data creation.
Patent Information
- Application Number
- JP2021194387
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-11-30
- Publication Date
- 2026-01-21
- Estimated Expiration
- 2041-11-30
AI Technical Summary
Existing supervised learning methods struggle to improve matching accuracy for unknown wireless transmitting devices, as they rely on labeled training data which is lacking for such devices, and existing technologies fail to distinguish and accurately match signals from multiple unknown devices.
A training data generation system that inputs unknown signals or radio wave features into a supervised learning model, performs clustering, and generates a relationship matrix to improve matching accuracy by associating these with estimated information, allowing for the creation of improved training data.
Enhances the accuracy of matching unknown transmitting devices by generating training data that improves the extraction and identification of sample features from unknown signals, leading to better device recognition.
Smart Images

Figure 0007803100000001 
Figure 0007803100000002 
Figure 0007803100000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a training data generation system, a training system, a training data generation method, a training method, and a program. [Background technology]
[0002] A technology has been proposed that identifies wireless terminal devices, such as mobile terminal devices, by matching them based on received signals received from the wireless terminal devices. To achieve this, it is necessary to register the characteristics of the received signals in advance in association with the wireless terminal devices. Such pre-registration can be achieved using supervised machine learning. Received signals typically vary depending on factors such as the reception environment, even when the same wireless terminal device is the radio wave source, and may also be received together with unknown signals. However, by taking these factors into consideration during training, it is possible to improve the accuracy of matching known wireless terminal devices.
[0003] Patent Documents 1 and 2 describe radio wave specification learning devices that learn to determine that an unknown target (radio wave source) is unknown.
[0004] The radio wave parameter learning device described in Patent Document 1 includes a radio wave parameter input unit, a target information input unit, and a learning unit. The radio wave parameter input unit receives, for each target, a plurality of parameters derived by A / D converting electromagnetic waves arriving from the target, or a radio wave parameter having the plurality of parameters simulating the target. The target information input unit receives, for each radio wave parameter, target information that is information about the target that is determined to match the radio wave parameter. The learning unit learns the range of the radio wave parameter that is determined to match the target based on the radio wave parameter input from the radio wave parameter input unit and the target information input from the target information input unit. The learning unit learns to determine that identification target parameters of an identification target that fall outside the learned range are unknown parameters whose radio wave parameters are unknown.
[0005] The radio wave parameter learning device described in Patent Document 2 includes a radio wave parameter input unit, a coincidence information input unit, and a learning unit. The radio wave parameter input unit receives, for each target, a plurality of parameters derived by A / D converting electromagnetic waves arriving from the target, or a radio wave parameter having the plurality of parameters simulating the target. The coincidence information input unit receives, for each radio wave parameter, coincidence information defining a parameter range of identification target parameters of an identification target that are determined to match the radio wave parameters. The learning unit learns the range of the identification target parameters that are determined to match the radio wave parameters based on the radio wave parameters input from the radio wave parameter input unit and the coincidence information input from the coincidence information input unit. The learning unit learns to determine that an identification target parameter that falls outside the learned range is an unknown parameter whose radio wave parameter is unknown.
[0006] Furthermore, Patent Document 3 describes a transmission device matching device that can improve the efficiency of operator work by automating template registration of unregistered transmission devices. The transmission device matching device described in Patent Document 3 includes a receiving unit that receives a signal wirelessly transmitted from a transmission device, a matching unit, and a template feature registration unit. The matching unit calculates the similarity between a sample feature generated from the received signal received by the receiving unit and a pre-registered template feature, and matches the transmission device by comparing the similarity with a matching threshold. The template feature registration unit generates a template feature from a sample feature that failed to be matched by the matching unit. [Prior art documents] [Patent documents]
[0007] [Patent Document 1] Japanese Patent Publication No. 2020-173171 [Patent Document 2] Japanese Patent Publication No. 2020-173172 [Patent Document 3] International Publication No. 2021 / 070248 Summary of the Invention [Problem to be solved by the invention]
[0008] As described above, supervised machine learning (supervised learning) is sometimes used as a method for matching wireless terminal devices, which are transmitters that transmit wireless signals. Supervised learning requires training data, i.e., data with correct labels. For transmitters whose signals are known, in other words, for known signals, data augmentation and domain adaptation techniques can be used to improve matching accuracy. Data augmentation is a method of increasing training data, such as by adding artificial noise or environmental information to received signals received from the same transmitter. Domain adaptation is a type of transfer learning and can refer to a learning method performed with a small amount of data. Domain adaptation is a method in which, for example, a large number of signals are acquired in environment A and a trained learning model is used to train and update only a portion of the learning model using a small number of signals acquired in environment B.
[0009] However, in supervised learning, it may be difficult to improve the accuracy of matching an unknown transmitting device when the signal is an unknown transmitting device, in other words, when the signal is an unknown signal. Specifically, even for an unknown transmitting device, learning about other transmitting devices of the same model can be expected to improve the accuracy of matching the unknown transmitting device, but if no other transmitting devices of the same model have been learned, it is difficult to improve the matching accuracy.
[0010] As such, it is desirable to improve the accuracy of matching unknown transmission devices, that is, to perform learning so as to improve the accuracy of matching unknown transmission devices. Note that the technologies described in Patent Documents 1 and 2 merely classify unknown signals as known signals, and are not technologies for matching unknown transmission devices, as can be seen from the fact that they cannot distinguish and match signals from multiple unknown transmission devices, for example. Furthermore, while the technology described in Patent Document 3 can register template features for unregistered transmission devices, there is room for improvement in terms of improving the accuracy of the registered template features, that is, in terms of improving the matching accuracy for unregistered transmission devices.
[0011] In view of the above-mentioned problems, the present disclosure aims to provide a training data generation system, a training system, a method, and a program that enable improvement in matching accuracy in the process of matching an unknown transmitting device based on a signal wirelessly transmitted from the transmitting device. [Means for solving the problem]
[0012] A training data generation system according to a first aspect of the present disclosure includes an input unit that inputs n pieces of first information, where N and n are positive integers, and the n pieces of first information are either n unknown signals that are signals wirelessly transmitted from N unknown transmission devices or n unknown radio wave features that are radio wave features generated from the n unknown signals; an extraction unit that inputs the n pieces of first information to a supervised learning model generated from training data including second information, the n pieces of first information being either known signals that are signals wirelessly transmitted from transmission devices whose transmission sources are known or known radio wave features that are radio wave features generated from the known signals, and a correct answer label associated with the second information, and extracts n sample features corresponding to each of the n pieces of first information; The estimation device includes a first clustering unit that executes a first clustering process on the n sample features that are the extracted results; an estimation information acquisition unit that inputs m pieces of estimation target information to be estimated by an estimation device that executes a process different from the first clustering process on the n unknown signals or the n unknown radio wave features, where M and m are positive integers, and acquires M pieces of estimation information associated with any one of the n unknown signals or any one of the n unknown radio wave features; and a generation unit that generates a relationship matrix that indicates the relationship between the first clustering result, which indicates the results classified into K groups by the first clustering process, and the M pieces of estimation information, where K is a positive integer.
[0013] A training data generation method according to a second aspect of the present disclosure includes: executing an input process to input n pieces of first information, where N and n are positive integers, and the n pieces of first information are either n unknown signals that are signals wirelessly transmitted from N unknown transmission devices or n unknown radio wave features that are radio wave features generated from the n unknown signals; inputting the n pieces of first information into a supervised learning model generated from training data including second information, the second information being either known signals that are signals wirelessly transmitted from a transmission device whose source is known or known radio wave features that are radio wave features generated from the known signals, and a correct label linked to the second information; and extracting n sample features corresponding to each of the n pieces of first information. a first clustering process is performed on the n sample features that are the extracted results, where M and m are positive integers, and m pieces of estimation target information to be estimated are input to an estimation device that performs a process different from the first clustering process on the n unknown signals or the n unknown radio wave features; an estimation information acquisition process is performed to obtain M pieces of estimation information associated with any one of the n unknown signals or any one of the n unknown radio wave features; and a generation process is performed to generate a relationship matrix that shows the relationship between the first clustering result, which shows the results classified into K groups by the first clustering process, and the M pieces of estimation information, where K is a positive integer.
[0014] A program according to a third aspect of the present disclosure is a program for causing a computer to execute a training data generation process, wherein the training data generation process includes: executing an input process for inputting n pieces of first information, where N and n are positive integers, and the n pieces of first information are either n unknown signals that are signals wirelessly transmitted from N unknown transmitting devices or n unknown radio wave features that are radio wave features generated from the n unknown signals; inputting the n pieces of first information into a supervised learning model generated from training data including second information, the n pieces of first information being either known signals that are signals wirelessly transmitted from transmitting devices whose transmission sources are known or known radio wave features that are radio wave features generated from the known signals, and a correct label associated with the second information; and executing an extraction process for extracting n sample features corresponding to each of the n pieces of first information; a first clustering process is performed on the n sample features obtained as a result of the first clustering process, where M and m are positive integers, and m pieces of estimation target information to be estimated are input to an estimation device that performs a process different from the first clustering process on the n unknown signals or the n unknown radio wave features; an estimation information acquisition process is performed to obtain M pieces of estimation information associated with any one of the n unknown signals or any one of the n unknown radio wave features; and a generation process is performed to generate a relationship matrix that shows the relationship between the first clustering result, which shows the results classified into K groups by the first clustering process, and the M pieces of estimation information, where K is a positive integer. [Effects of the Invention]
[0015] The present disclosure provides a training data generation system, a training system, a training data generation method, a training method, and a program that enable improvement in matching accuracy in a process of matching an unknown transmitting device based on a signal wirelessly transmitted from the transmitting device. Note that the present disclosure may achieve other effects instead of or in addition to the above effect. [Brief explanation of the drawings]
[0016] [Figure 1]1 is a block diagram showing an example of the configuration of a training data generation system according to a first embodiment. [Figure 2] 2 is a block diagram showing an example of the configuration of a learning system including the learning data generation system of FIG. 1. FIG. [Figure 3] 2 is a block diagram showing an example of the configuration of a transmission device verification system including the training data generation system of FIG. 1. FIG. [Figure 4] FIG. 10 is a diagram illustrating an overview of a transmission device verification system according to a second embodiment. [Figure 5] FIG. 10 is a block diagram showing an example of a functional configuration of a transmission device verification system according to a second embodiment. [Figure 6] FIG. 10 is a diagram illustrating an example of the layout of a transmission device verification system according to a second embodiment. [Figure 7] FIG. 6 is a diagram showing an example of clusters output by a clustering unit in FIG. 5. [Figure 8] 6 is a diagram showing an example of a relationship matrix generated by a matrix generation unit and visualized by a visualization unit in FIG. 5. FIG. [Figure 9] 6 is a schematic diagram for explaining an overview of a learning data generation process performed for an unknown transmission device in the transmission device verification system of FIG. 5. FIG. [Figure 10] FIG. 10 is a schematic diagram showing an example of a result of changing a clustering threshold in the learning data generation process of FIG. 9. [Figure 11] 10 is a schematic diagram for explaining an example of the effect of re-learning using additional training data generated as a result of the training data generation process of FIG. 9. FIG. [Figure 12] 6 is a flowchart illustrating an example of processing in the transmission device verification system of FIG. 5. FIG. [Figure 13] FIG. 10 is a block diagram showing an example of a functional configuration of a transmission device verification system according to a third embodiment. [Figure 14] FIG. 10 is a block diagram showing an example of a functional configuration of a transmission device verification system according to a fourth embodiment. [Figure 15] FIG. 2 illustrates an example of a hardware configuration included in the device. DETAILED DESCRIPTION OF THE INVENTION
[0017] Hereinafter, embodiments will be described with reference to the drawings. In this specification and the drawings, elements that can be similarly described will be assigned the same reference numerals, and redundant description will be omitted. Furthermore, although some of the drawings described below depict unidirectional arrows, these arrows simply indicate the direction of flow of a certain signal (data) and do not exclude bidirectionality.
[0018] First Embodiment To explain briefly, in the first embodiment and other embodiments described later, a method can be adopted in which a training data set is added to training data using labels obtained by estimating unknown transmission devices using other means. However, if this method is adopted as is, it cannot be used if the accuracy (reliability) of the labels obtained by the above estimation is low, and performance may be degraded instead. Therefore, in each embodiment, a method for solving such a problem will be described.
[0019] The first embodiment will be described with reference to Figures 1 to 3. Figure 1 is a block diagram showing an example of the configuration of a training data generation system according to the first embodiment.
[0020] 1, a training data generation system according to this embodiment (hereinafter referred to as the present system) 1 includes an input unit 1a, an extraction unit 1b, a first clustering unit 1c, an estimated information acquisition unit 1d, and a generation unit 1e. The present system 1 can be configured as a single device, or can be configured as a distributed system in which functions are distributed across multiple devices.
[0021] The input unit 1a inputs n pieces of first information, which are either n unknown signals (unregistered signals or unlearned signals) or n unknown radio wave features wirelessly transmitted from N unknown transmitting devices (unregistered or unlearned transmitting devices). The unknown radio wave features refer to unregistered radio wave features or unlearned radio wave features. Here, N and n are positive integers, and the n unknown radio wave features are radio wave features generated from the n unknown signals, respectively. The number of dimensions (number of types of information) of each unknown radio wave feature does not matter. The unregistered transmitting device refers to a transmitting device for which template features for matching transmitting devices by matching sample features extracted by the extraction unit 1b, described later, are not registered. The unregistered signal refers to a signal transmitted from such an unregistered transmitting device. The unlearned signal and transmitting device refer to a signal and transmitting device, regardless of whether they are registered or unregistered, that are not the subject of learning in the supervised learning model, described later.
[0022] When n unknown signals are input, the unknown signals can be separated in advance based on radio wave features such as signal strength and frequency band. However, n can be unknown at the time of input, and the n unknown signals can be defined as being classified into n as extraction results by extraction unit 1b (described later). In this way, n and N can be unknown at the time of input by input unit 1a, and the number of original first information from which n sample features are extracted by extraction unit 1b (described later) is expressed as n. Furthermore, data on multiple received signals and radio wave features is basically acquired from one transmitting device, and the total number n of first information is the sum of the number of data from each of the N transmitting devices.
[0023] As can be seen from these examples, for the sake of simplicity, the following explanation will be based on the assumption that one sample feature is extracted per piece of first information, but the number of dimensions (number of types of features) of the sample feature extracted from one piece of first information is not important.
[0024] The input unit 1a can input n pieces of first information at once, but can also input some or all of the first information individually. For example, when the first information is an unknown signal, the input unit 1a can be a radio receiving unit that receives the unknown signal and can receive multiple unknown signals simultaneously. When the first information is an unknown radio wave feature, the input unit 1a can input the unknown radio wave feature individually each time it is generated from each unknown signal.
[0025] The extraction unit 1b inputs n pieces of first information into a supervised learning model generated from learning data, which will be described below, and extracts n sample features corresponding to the n pieces of first information, respectively. The extraction unit 1b can be configured to include such a learning model or to be able to access such a learning model. The number of dimensions of each sample feature is not important, but can be lower than the number of dimensions of the unknown radio wave feature when the first information is an unknown radio wave feature.
[0026] The training data includes second information, which is either a known signal or a known radio wave feature generated from the known signal, and a correct label for the second information. Here, the known signal included in the training data is a signal whose wireless transmission source is known, i.e., a signal whose transmission source is known, regardless of whether it is registered. The correct label during training is a label associated with the known signal or a label associated with the radio wave feature of the known signal. In this way, the correct label can be a label associated with the second information. Note that the correct label may also be a label associated with a sample feature of the known signal during registration. In this way, the known signal and the known radio wave feature refer to a signal and a radio wave feature whose transmission source is known, respectively.
[0027] As described above, the learning model is a model that has undergone machine learning to obtain correct sample features from second information, which is a known signal or known radio wave features, and to output correct labels indicating the sample features. Alternatively, a model can be adopted in which a portion of the model is extracted as a sample feature extractor and machine-learned to ensure that sample features with the same correct label have high similarity and correlation, and sample features with different correct labels have low similarity and correlation. Any machine learning algorithm can be used, and for example, a convolutional neural network (CNN) can be used, which inputs the waveform of a received signal or its radio wave features and outputs correct labels indicating correct features. However, the machine learning algorithm is not limited to this, and various algorithms can be applied, such as other deep learning algorithms that input the values of received signals or radio wave features and output correct labels indicating correct features.
[0028] Furthermore, both the unknown transmitting device and the known transmitting device can be wireless terminal devices (wireless terminals) capable of wireless communication, but may also include devices that unintentionally emit radio waves (such as LED (light-emitting diode) devices that emit noise or wireless devices with faulty amplifiers). Hereinafter, regardless of whether they are known or unknown, transmitting devices may be referred to as "transmitting terminals" or simply "terminals."
[0029] The first clustering unit 1c executes a first clustering process on the n sample features extracted by the extraction unit 1b. Note that the term "first" here is merely used to distinguish it from another clustering unit (second clustering unit) or another clustering process (second clustering process) described later in the third embodiment.
[0030] The first clustering unit 1c outputs the results of the first clustering process (hereinafter referred to as the first clustering result) classified into K groups. Here, K is a positive integer and can be a predetermined number of clusters. However, even in clustering with a predetermined number of clusters, there may be cases where a cluster does not contain a single sample feature. Therefore, K is simply the number of groups classified as the first clustering result. The first clustering result may include, for example, an identification number or identifier such as a cluster number assigned to each sample feature.
[0031] Any algorithm may be used for the first clustering process, and for example, clustering based on similarity comparison or machine learning methods such as k-means, x-means, vbgmm, etc. Furthermore, as the algorithm for the first clustering process, a hierarchical clustering method using a group average method, Ward's method, minimum distance method, maximum distance method, etc. may also be used.
[0032] The estimation information acquisition unit 1d inputs m pieces of estimation target information to be estimated by an estimation device that executes a process different from the first clustering process for n unknown signals or n unknown radio wave features. Here, m is a positive integer. The estimation target information input by the estimation information acquisition unit 1d is an unknown signal transmitted from the same transmitting terminal as the first information input by the input unit 1a or an unknown radio wave feature generated from the unknown signal, and this information is the target of estimation in the estimation device. However, the unknown signal or unknown radio wave feature that is the source of estimation (the target of estimation) is not limited to being input to the estimation device as data in the same format as the first information input by the input unit 1a. To give a simple example, for example, if the first information is an unknown signal, the unknown radio wave feature can be input to the estimation device, and if the first information is an unknown radio wave feature, the unknown signal can be input to the estimation device. Furthermore, even if both the first information and the second information are unknown radio wave features, they may be generated by any method, and therefore may have different data formats. Furthermore, since the number of unknown signals or unknown radio wave features to be estimated may differ from the number n of first information, the number of estimation target information input to the estimation information acquisition unit 1d is described as m.
[0033] Then, the estimated information acquisition unit 1d acquires M pieces of estimated information associated with any one of the n unknown signals or any one of the n unknown radio wave features, where M is a positive integer. Note that the order of the processing by the estimated information acquisition unit 1d and the first clustering processing by the first clustering unit 1c does not matter, and they can also be executed in parallel.
[0034] Of course, the above-described estimation device can also apply an estimation process including a clustering process, but the clustering process differs from the first clustering process in at least one of the algorithm, input data format, and clustering threshold. Here, the above-described estimation device does not necessarily output N pieces of estimated information even if there are the same number of estimation devices as the actual number N of transmitting devices, so the number of pieces of estimated information is described as M rather than N. Furthermore, the above-described estimation device performs a process different from the first clustering process, so the number of pieces of estimated information output does not necessarily equal K pieces of clustering results obtained by the first clustering unit 1d, so the number of pieces of estimated information is described as M rather than K. For convenience, this estimation device is described as not being included in the present system, but the present system can be configured by including this estimation device.
[0035] The estimation information acquisition unit 1d acquires M pieces of estimation information from the m pieces of estimation target information input, but the method for doing so is not limited. For example, as described in the second embodiment, the m pieces of estimation target information to be input may be m radio wave features or signals, and M pieces of estimation information associated with the m radio wave features or signals may be input together. Including this case, the estimation information acquisition unit 1d may already associate the m pieces of estimation target information with any one of the n pieces of unknown information or any one of the n unknown radio wave features at the time of inputting the m pieces of estimation target information. Alternatively, the estimation information acquisition unit 1d may perform such association upon inputting the m pieces of estimation target information. In the latter case, the estimation information acquisition unit 1d may associate the n pieces of unknown information or unknown radio wave features with each of the finally acquired m pieces of estimation target information, for example, by calculating the similarity between each of the n unknown signals or unknown radio wave features and the n pieces of unknown information or unknown radio wave features.
[0036] The generation unit 1e generates a relationship matrix indicating the relationship between the first clustering results classified into K groups by the first clustering process and the M pieces of estimated information acquired by the estimated information acquisition unit 1d. When generating the relationship matrix, the first clustering results and the estimated results (estimated information) can be associated based on the unknown signals or unknown radio wave features, respectively, that is, with the unknown signals or unknown radio wave features as common terms. The generation unit 1e can generate the relationship matrix by arranging statistical values such as the frequencies of the unknown signals or unknown radio wave features in a K×M matrix.
[0037] The present system is capable of generating such a relationship matrix, and the generated relationship matrix can be used to generate the following training data. That is, the present system can generate training data that can improve the extraction accuracy of extracting sample features from signals wirelessly transmitted from unknown transmitting terminals or from the radio wave features of those signals. As a result, the present system can generate training data that can improve the matching accuracy in the process of matching unknown transmitting terminals based on signals wirelessly transmitted from those terminals (including based on the radio wave features of those signals).
[0038] Next, a learning system including the learning data generation system 1 will be described with reference to Fig. 2. Fig. 2 is a block diagram showing an example of the configuration of a learning system including the learning data generation system 1 of Fig. 1.
[0039] The learning system 2 shown in Figure 2 includes the present system 1 and learning unit 2a shown in Figure 1, and further includes a label setting unit 1f and a data generation unit 1g in the present system 1. The learning system 2 can be configured as a single device, but can also be configured as a distributed system in which functions are distributed across multiple devices.
[0040] The label setting unit 1f sets correct labels to at least some of the intersections (K×M intersections) of the generated relationship matrix. The intersections can also be called cells. The label setting unit 1f can include an operation unit that accepts input or selection input for specifying correct labels from an operator who generates learning data.
[0041] Although not shown, the learning system 2 preferably includes a display unit (display device) that displays the generated relationship matrix. This allows the operator to set the correct label while visually checking the visualized relationship matrix. The display unit can also display the first clustering result. This display unit can also be included in the learning data generation system 1 of FIG. 1.
[0042] However, the label setting unit 1f can also be configured to automatically set correct labels without operator input. In this case, a learning model for such setting can be used. For example, settings based on operator input can be performed at the beginning of operation, and matching accuracy can be obtained using the settings and the settings with improved matching accuracy as training data to generate the learning model. The algorithm for this learning model is not limited. Alternatively, correct labels can be automatically set by threshold judgment or the like based on information regarding the relationships between K clusters and information such as the reliability of M pieces of estimated information. In the description of this embodiment and the second and subsequent embodiments, it is assumed that correct labels are set by an operator. However, correct labels can also be set automatically as described above. In this case, the operator's judgment can be explained by replacing it with automatic setting based on the judgment (result output) of the learning model or threshold judgment or the like.
[0043] The data generation unit 1g generates learning data for updating the learning model used by the extraction unit 1b based on the unknown signals or unknown radio wave features associated with each of the at least some of the intersections and the correct labels assigned to each of the at least some of the intersections. The K×M intersections included in the relationship matrix include intersections associated with at least one of the n unknown signals or unknown radio wave features and intersections not associated with any of the n unknown signals or unknown radio wave features. While the description here is based on the assumption that n≦m, if n>m, it can be said that the K×M intersections include intersections associated with at least one of the m pieces of estimation target information and intersections not associated with any of the m pieces of estimation target information. If n=m, all n unknown signals or unknown radio wave features are associated with any of the m pieces of estimation target information (radio wave features or signals) and are associated with any of the intersections. On the other hand, if there is an unknown signal or unknown radio wave feature among the n unknown signals or unknown radio wave feature quantities that cannot be associated with any of the m pieces of estimation target information, such an unknown signal or unknown radio wave feature quantity can be excluded from the relationship matrix. Similarly, if there is information among the m pieces of estimation target information that cannot be associated with any of the n unknown signals or unknown radio wave feature quantities, such information can be excluded from the relationship matrix.
[0044] The data generating unit 1g can generate one set of training data for each correct label. This set can include data (raw data sets, or statistical data such as averages and medians) obtained from one or more unknown signals or unknown radio wave features corresponding to one or more intersections to which correct labels are set, and the correct labels.
[0045] The learning unit 2a performs machine learning based on additional learning data, which is learning data generated by the system 1, and original learning data, which is learning data including second information that was the learning target data in the original learning model and correct labels linked to sample features of known signals. The learning unit 2a updates the learning model through this machine learning. Depending on the algorithm of the learning model, this update can also be performed using only the additional learning data.
[0046] In the learning system 2, by updating the learning model in this way, it is possible to improve the extraction accuracy of extracting sample features from signals wirelessly transmitted from unknown transmitting terminals or from the radio wave features of those signals, thereby improving the matching accuracy in the process of matching unknown transmitting terminals based on signals wirelessly transmitted from those terminals (including based on the radio wave features of those signals).
[0047] Next, a transmission device verification system including the training data generation system 1 will be described with reference to Fig. 3. Fig. 3 is a block diagram showing an example of the configuration of a transmission device verification system including the training data generation system 1 of Fig. 1.
[0048] The transmission device verification system 3 shown in Fig. 3 is a system including the present system 1 shown in Fig. 1 and a verification unit 3a, in other words, a system having a verification function in a training data generation system. In this way, the transmission device verification system 3 can also be called a training data generation system because it generates training data.
[0049] 1 and 2, it has been assumed that a known signal (i.e., a signal that is known or registered by the transmitting device that is the sender) is not received, or that radio wave features corresponding to known signals are not input, or that input signals are ignored. In contrast, in the transmitting device matching system 3, the input unit 1a is configured to input these signals as well. In particular, when the information input to the input unit 1a is a signal, known signals may be mixed in, and it is inevitable that known signals and unknown signals may be received. Therefore, it is beneficial to configure the system to handle such cases as well.
[0050] An input unit 1a in the transmission device matching system 3 inputs third information, which is either signals wirelessly transmitted from any of a plurality of transmission terminals or radio wave features generated from the signals, as information including the first information. An extraction unit 1b in the transmission device matching system 3 inputs the third information to a learning model and extracts sample features for the third information.
[0051] The matching unit 3a then matches the sample features of the extracted third information with pre-registered template features. The template features may be stored in a storage device in advance. It is desirable that the template features of a certain transmitting terminal are representative of many sample features. Successful matching of the sample features means that matching of the transmitting terminal has been successful. The matching method used by the matching unit 3a is not critical, and various methods can be applied, such as a method of similarity comparison. A detailed description of an example of matching by calculating similarity will be omitted, but the technology described in Patent Document 3, for example, can be applied. Note that "matching" can also be rephrased as "identifying," "specifying," "determining," etc.
[0052] On the other hand, a sample feature determined to be unknown or unregistered in the matching is a sample feature extracted from an unknown signal wirelessly transmitted from an unregistered transmitting terminal or an unknown radio wave feature generated from the unknown signal. The unknown signal or unknown radio wave feature can be used as a target for generating learning data in the system 1. Note that, when matching is performed by calculating a similarity, a sample feature determined to be unknown or unregistered in the matching can refer to a sample feature whose similarity is lower than a matching threshold.
[0053] Therefore, the first clustering unit 1c in the transmission device matching system 3 performs a first clustering process on the n sample features that have been determined to be unknown or unregistered in the matching by the matching unit 3a. The processes by the estimated information acquisition unit 1d and the generation unit 1e are as described for the present system 1, and estimated information is acquired and a relationship matrix is generated. However, the first clustering unit 1c can also target not only the sample features that have been determined to be unknown or unregistered by the matching unit 3a, but also other sample features that were used for matching by the matching unit 3a for the first clustering process.
[0054] The transmission device matching system 3 may further include a template feature registration unit (not shown) that generates and registers template features from sample features determined to be unknown or unregistered in matching by the matching unit 3a. A method for generating template features from sample features determined to be unknown or unregistered in matching may be, but is not limited to, the method described in Patent Document 3, for example. However, the template feature registration unit can generate template features not only from sample features determined to be unknown or unregistered by the matching unit 3a, but also from other sample features used for matching by the matching unit 3a.
[0055] Although not shown in the drawings or explained, the transmitting device verification system 3 can also be configured to include a label setting unit 1f, a data generating unit 1g, and a learning unit 2a. The transmitting device verification system 3 can be configured as a single device, but can also be configured as a distributed system in which functions are distributed among multiple devices.
[0056] For example, the matching unit 3a can be placed in an edge device as a transmitting device matching device, and the other components can be placed in a high-performance device, or the learning unit 2a can be placed in a high-performance device, and the other components can be placed in an edge device. Generally, the required computing performance during machine learning learning is often significantly greater than that during operation (inference). In particular, when deep learning is used for model generation (generation of learning parameters), a high-performance graphics processing unit (GPU) can be considered for use during learning. Although GPUs offer high performance, they are expensive and consume a lot of power, making them unsuitable for applications in which a large number of transmitting device matching devices are deployed over a wide area as edge devices.
[0057] However, this problem can be solved by generating a learning model through learning on an external high-performance device, and then extracting sample features on the edge device using a low-cost, low-power hardware accelerator dedicated to inference.In other words, this configuration makes it possible to perform terminal matching and generate additional learning data at low cost and with low power consumption in applications where a large number of transmission and matching devices are deployed over a wide area as edge devices.
[0058] <Second embodiment> The second embodiment will be described with reference to Figures 4 to 12, focusing on the differences from the first embodiment, but the various examples described in the first embodiment can also be applied to the second embodiment. Of course, the various examples described in this embodiment can also be applied to the first embodiment.
[0059] First, an overview of a transmission device verification system that functions as a training data generation system according to a second embodiment will be described with reference to Fig. 4. Fig. 4 is a diagram for explaining the overview of the transmission device verification system according to the second embodiment. Note that the description of this overview is not intended to be limiting in any way.
[0060] As shown in FIG. 4, a transmission device verification system 10 according to this embodiment (hereinafter referred to as the present system 10) can include a transmission device verification unit 11, an estimation unit 12, and a training data processing unit 13.
[0061] The transmission device matching unit 11 is an example of the input unit 1a, extraction unit 1b, and matching unit 3a in FIG. 3. The input unit 1a in this example receives an unknown signal and performs generation 11a of radio wave features from the received unknown signal. The extraction unit 1b in this example inputs each generated radio wave feature to a learning model 11b that has been trained using a known signal, thereby generating sample features 11c. The matching unit 3a in this example then performs matching determination 11e by matching each generated sample feature 11c with a template (template feature) 11d registered in advance in a database (DB).
[0062] If the matching is successful, the matching decision 11e outputs a matching result indicating that the transmitting terminal is the corresponding template. If all the input unknown signals are signals transmitted from unregistered (unknown) transmitting terminals and the matching accuracy is high, there will basically be no sample features that result in successful matching, and all sample features to be matched will be determined to be unregistered (unknown). On the other hand, if the matching decision 11e determines that a sample is unknown or unregistered in the matching, it outputs the sample features determined as such to the learning data processing unit 13.
[0063] The received signal that is wirelessly transmitted and received and input to the transmitting device matching unit 11 is called an unknown signal because it is before matching and is unknown to the transmitting device matching unit 11. However, in a situation where there is a sample feature that is successful in matching, the received signal that is input to the transmitting device matching unit 11 will include not only unknown signals wirelessly transmitted from unregistered transmitting terminals but also known signals wirelessly transmitted from registered transmitting terminals.
[0064] The training data processing unit 13 receives the sample features output from the transmission device matching unit 11 and generates training data for the received sample features. The training data processing unit 13 includes a clustering unit 13a, a relationship visualization unit 13b, and a data generation unit 13c. The clustering unit 13a corresponds to the first clustering unit 1c in FIG. 2, the relationship visualization unit 13b corresponds to the generation unit 1e in FIG. 2, and the data generation unit 13c corresponds to the label setting unit 1f and the data generation unit 1g in FIG. 2. The clustering unit 13a performs a first clustering process on the received sample features, and as a result, can assign a cluster number to each sample feature.
[0065] Prior to describing the relationship visualization unit 13b and the data generation unit 13c, the estimation unit 12 will be described. The estimation unit 12 includes an estimation device that estimates transmitting terminals, as described in the first embodiment. This estimation device includes an estimated information acquisition unit 1d, and executes a process different from the first clustering process on n unknown signals or n unknown radio wave features, and outputs M pieces of estimated information as information indicating M transmitting terminals. The output M pieces of estimated information are information associated with any one of the n unknown signals or any one of the n unknown radio wave features.
[0066] However, in this embodiment, m radio wave features are input as m pieces of estimation target information, and M pieces of estimated information including M temporary labels associated with one or more of the m radio wave features are also input. That is, the M pieces of estimated information are included in the information input from the estimation device. Note that m signals can also be used instead of the m radio wave features. The M temporary labels are labels indicating each of the M transmitting devices. In other words, in this embodiment, M pieces of estimated information including M temporary labels indicating each of the M transmitting devices and m radio wave features (or signals) associated with any of the M temporary labels are input from the estimation device. The estimation unit 12 only needs to be able to generate M temporary labels by estimation from the same unknown signal (or unknown radio wave features generated by any method from the unknown signal) as the unknown signal input to the transmission device matching unit 11, that is, from the m signals or radio wave features, and the estimation method is not important.
[0067] Here, the tentative label represents information about the transmitting terminal, and can be, for example, a model name or an individual ID, but is not a reliable label such as a known signal, i.e., a correct label, but is a label inferred by the estimation device based on some other information. As will be described later, each piece of estimated information can include, in addition to the tentative label, reliability information indicating its reliability in association with it.
[0068] In this example, the estimation unit 12 has a position estimation function for estimating the position of the transmitting terminal, a bandwidth estimation function for estimating the bandwidth of the signal wirelessly transmitted by the transmitting terminal (hereinafter referred to as the wireless transmission signal), and a modulation scheme estimation function for estimating the modulation scheme used when generating the wireless transmission signal.
[0069] Additionally, the estimation unit 12 may have a frequency estimation function for estimating the frequency or frequency band of the wireless transmission signal, a transmission power estimation function for estimating the power value of the wireless transmission signal, and a transmission frequency estimation function for estimating the frequency at which the wireless transmission signal is transmitted (the frequency at which the transmitting terminal transmits a signal). The estimation unit 12 may also have a transmission time occupancy estimation function for estimating the occupancy rate of time during which the wireless transmission signal is transmitted, a packet length estimation function for estimating the transmission packet length of the wireless transmission signal, and a transmission data amount estimation function for estimating the amount of data transmitted as the wireless transmission signal. The estimation unit 12 may also have a frequency hopping estimation function for estimating a frequency switching pattern when the wireless transmission signal is transmitted using a frequency hopping method. The estimation unit 12 may also have a spectrogram estimation function for estimating the spectrogram of the wireless transmission signal and a spectrum estimation function for estimating the spectrum of the wireless transmission signal.
[0070] However, the estimation unit 12 can have at least one of the functions listed here, for example. That is, the estimated information can be information estimated by an estimation device, for example, at least one of the positions of each of the M transmitting terminals, the bands and frequencies of signals wirelessly transmitted by each of the M transmitting terminals, and the modulation methods used by each of the M transmitting terminals.
[0071] The relationship visualization unit 13b inputs the temporary labels estimated by the estimation device and the first clustering results that are the processing results of the clustering unit 13a, and generates a relationship matrix that indicates the relationship between the two. The relationship matrix is generated for reference when setting correct labels using the estimated information estimated by the estimation device. The temporary labels and the first clustering results can be associated with each other using the original unknown signals or unknown radio wave features as keys (common terms).
[0072] The relationship visualization unit 13b visualizes the generated relationship matrix by displaying it on a display unit. The data generation unit 13c sets correct labels to some of the K × M intersections of the relationship matrix, as described for the label setting unit 1f. In FIG. 4, each intersection indicates the occurrence frequency among the original n sample features, in other words, the occurrence frequency among the n unknown signal or unknown radio wave features. However, this example is not limiting, and instead of the occurrence frequency, the display may be in the form of an occurrence rate relative to the total. Furthermore, instead of the occurrence frequency or occurrence rate, various first clustering results, such as the sum of the distance from the center of gravity for each cluster or the distance from the center of the cluster area, may be displayed at the intersection. Furthermore, two or more types of values may be displayed at each intersection. Of course, the display format of the relationship matrix is not limited to the example. Note that in this embodiment, as in the first embodiment, description is given on the assumption that n≦m, and description of the case where n>m is omitted.
[0073] The setting of correct labels can be performed by an operator, and for intersections with high reliability shown in the relationship matrix, correct labels are set for the data linked to those intersections. Of course, if there are no intersections with high reliability, it is possible not to set correct labels at all. Reliability will be described later, but for example, the larger the value at an intersection (for example, a frequency value such as the number of data corresponding to that intersection or other statistical value), the higher the reliability of that intersection can be recognized by the operator.
[0074] Furthermore, the data generation unit 13c generates additional training data in which a correct label is added only to data with high reliability. Specifically, the data generation unit 13c generates training data based on the correct labels that the operator has determined to be highly reliable and set for each of some intersections, and the unknown signals or unknown radio wave features associated with each of these intersections. This training data is training data for additional training to update the learning model 11b, and can be referred to as additional training data. Here, one correct label can be set for multiple intersections.
[0075] For example, this additional training data may include data such as unknown signals associated with temporary labels or cluster numbers corresponding to one or more intersections determined to be highly reliable, and set correct labels (generated labels). The generation of additional training data may be performed automatically by, for example, associating data such as unknown signals with temporary labels and cluster numbers as described above, and by referring to text information. The text information is information listing combinations of temporary labels, cluster numbers, and correct labels for each set correct label, and may be stored in, for example, a table format.
[0076] As can be seen from this explanation, the temporary label can be information on a candidate to be used as the correct label of the learning model. However, the candidate correct label is not limited to this. For example, a label named to include a cluster number can be set as the correct label, or a correct label can be set using a completely unique serial number or the like.
[0077] In this way, the data generator 13c can generate one set of training data for each set correct label. This set can include data (raw data sets, or data of statistical values such as averages and medians) obtained from one or more unknown signals or unknown radio wave features corresponding to one or more intersections to which correct labels are set, and the correct labels.
[0078] The learning model 11b is updated by performing machine learning based on this additional learning data and the original learning data, which is learning data including the second information that was the learning target data in the original learning model and the correct label linked to the second information.
[0079] To improve accuracy, the correct labels are generally set by an operator, and the subsequent generation of additional training data can be performed automatically. However, since the reliability level can also be determined automatically, by incorporating a function to accurately determine this reliability, it is also possible to set the correct labels and generate additional training data without the intervention of an operator.
[0080] In this way, the present system 10 generates a relationship matrix showing the relationship between the results of estimation by another estimation device and the first clustering results, and visualizes the generated relationship matrix to show highly reliable intersections. Therefore, the present system 10 allows an operator to set a correct label for an unknown signal based on the reliability, that is, to perform highly reliable correct labeling, and as a result, it is possible to generate highly reliable additional training data for the unknown signal.
[0081] The system 10 can then retrain the learning model 11b using this additional training data, thereby enabling extraction of appropriate sample features. As a result, the system 10 can improve the matching accuracy of unknown transmitting terminals and the matching accuracy of known transmitting terminals.
[0082] A more specific configuration example of this embodiment will be described in detail below with reference to Figs. 5 to 12. First, an example of the configuration and arrangement of this system (transmission device matching system) 10 will be described with reference to Figs. 5 to 8. Fig. 5 is a block diagram showing an example of the functional configuration of this system 10, and Fig. 6 is a diagram showing an example of the arrangement of this system 10. Fig. 7 is a diagram showing an example of clusters (examples of sample features) output by the clustering unit in Fig. 5, and Fig. 8 is a diagram showing an example of a relationship matrix generated by the matrix generation unit and visualized by the visualization unit in Fig. 5.
[0083] The system 10 shown in FIG. 5 has the functions outlined in FIG. 4. That is, the system 10 compares a transmitting terminal (not shown) with a sample feature generated from a received signal wirelessly transmitted by the transmitting terminal using a learning model based on individual differences in the radio waves, and the template feature by performing a similarity calculation or the like. The template feature can be registered in an internal database in advance. Furthermore, for a received signal determined to be unknown or unregistered in the comparison, the system 10 generates additional learning data for updating the learning model used to generate sample feature data, and updates the learning model.
[0084] Here, we will explain individual differences in radio waves. Differences in the specifications of transmitting terminals, or even if the specifications are the same, variations in the characteristics of analog circuits implemented in transmitting terminals, can result in individual differences in the transmitted radio waves. The system 10 registers, for each transmitting terminal, the features of the radio waves transmitted by that transmitting terminal (e.g., statistical values of sample features generated from the received signal using a learning model) as template features in a database. When the system 10 receives radio waves, it generates sample features of the received signal using the learning model. The system 10 identifies the transmitting terminal that is the source of the received radio waves by comparing these sample features with the template features in the database. For example, when performing similarity calculations in the matching, if there is a terminal with a template feature greater than a predetermined threshold, the transmitting terminal that is the source of the received radio waves is identified. If there are multiple terminals with template features greater than a predetermined threshold, the largest one among them may be output as the matching result, or a predetermined number of candidates (two or more) may be output along with estimated probability and similarity information.
[0085] Transmitting terminal verification includes "individual identification" that identifies the individual transmitting terminal. Furthermore, transmitting terminal verification also includes "model identification" that identifies the model that transmitted the radio waves, although it does not specify which individual transmitted the radio waves. Furthermore, it can also include "attribute identification" that identifies the attribute or use of the transmitting terminal, such as a consumer terminal, commercial radio device, interference source, or specified low-power radio. In consideration of this situation, in the following explanation, "individual identification," "model identification," and "attribute identification" may be collectively referred to as "radio wave identification" or "terminal verification."
[0086] The system 10 only needs to be able to extract features of received radio waves (received wireless signals) using a learning model, and it is not necessary for the transmitting terminal to transmit radio waves to the system 10. The system 10 can be used (applied) for a variety of purposes, such as detecting and tracking suspicious individuals in urban areas and various facilities (airports, shopping malls, etc.), understanding customer movements in stores and commercial facilities, and managing entrance and exit to limited areas using radio waves. Of course, the uses are not limited to those exemplified here.
[0087] The present system 10 can determine the identity of a transmitting terminal using the characteristic quantities of radio waves. However, the present system 10 cannot directly identify the owner of the transmitting terminal based on the characteristic quantities. In this way, the characteristic quantities of radio waves used by the present system 10 are anonymous, and the present system 10 can perform processing that takes into consideration the privacy of individuals.
[0088] Each component of the system 10 shown in FIG. 5 will now be described. 5, the system 10 may include a receiving unit 111, a radio wave feature generating unit 112, a learning unit 113, and a matching unit 130. The system 10 may further include a feature clustering unit 140, a temporary label acquiring unit 150, a matrix generating unit 151, a visualization unit 152, a data generating unit 153, and a label setting unit 154.
[0089] The receiving unit 111 is an example of the input unit 1a in FIG. 3, and receives radio waves (wireless signals) from transmitting terminals, including a transmitting terminal that is the target of radio wave identification. The receiving unit 111 can be configured to include a radio wave sensor for receiving radio waves. The number of receiving units 111 included in the present system 10 may be one or more. In other words, the present system 10 may include at least one receiving unit 111.
[0090] Here, with reference to FIG. 6, an example of the arrangement of the present system 10 including the receiving unit 111 shown in FIG. 5 and transmitting terminals will be described. The example of FIG. 6 shows the present system 10 and transmitting terminals 900a and 900b arranged within a target area A1 for terminal matching by the present system 10. Note that the transmitting terminal 900a is a transmitting terminal to be matched by the present system 10, and the transmitting terminal 900b is a transmitting terminal that is not a target for matching by the present system 10, i.e., a transmitting terminal for which template features are not registered. In the present disclosure, unless there is a particular reason to distinguish between the transmitting terminal 900a and the transmitting terminal 900b, they will simply be referred to as "transmitting terminal 900." Note that while FIG. 6 illustrates one transmitting terminal 900a to be matched, in reality, multiple transmitting terminals 900a to be matched are included. In other words, at least one transmitting terminal 900a is usually present in the field (target area).
[0091] Examples of the transmitting terminal 900 include mobile terminal devices such as mobile phones (including those called smartphones), game consoles, and tablet terminals, as well as computers (personal computers, laptop computers). Alternatively, the transmitting terminal 900 may be an IoT (Internet of Things) terminal or an MTC (Machine Type Communication) terminal that emits radio waves. However, the transmitting terminal 900 (including the target of terminal matching by the present system 10) is not limited to the above examples. That is, in the present disclosure, any device that emits radio waves can be the target of terminal matching by the present system 10, or the target of generating additional training data for a training model generated by the training unit 113, which will be described later.
[0092] As described above, the radio waves transmitted by the transmitting terminal 900a do not need to be radio waves transmitted to the system 10 (to the receiving unit 111). For example, the receiving unit 111 may receive radio waves transmitted by the transmitting terminal 900 to a wireless communication base station or access point for a mobile phone or the like, or radio waves transmitted by the transmitting terminal 900 to search for a wireless communication base station or access point.
[0093] Furthermore, the present system 10 is assumed to be installed in an environment where a large number of unspecified transmitting terminals whose template features are not registered in the database may transmit data. Data on a large number of unspecified transmitting terminals whose template features are not registered in the database is not included in the training data used when the training unit 113 generates a training model (training model 11b in FIG. 4). Therefore, in such an installation environment, it is not expected that the training model will generate sample features with high accuracy for unknown signals or unknown features. This results in a decrease in the matching accuracy of known transmitting terminals and an increased possibility of mismatching an unknown transmitting terminal as a known transmitting terminal. To solve such problems, in this embodiment, additional training data for unregistered transmitting terminals is generated and the training model is updated.
[0094] Returning to the detailed description of each part in Fig. 5, the radio wave feature generation unit 112 generates a radio wave feature from the received signal received by the reception unit 111. The radio wave feature used by the present system 10 to verify the transmitting terminal that is the source of the radio wave can be various features that reveal individual differences between the transmitting terminals 900.
[0095] Examples of radio wave feature quantities include transients (rising and falling edges) of the received signal in the receiver 111, power spectral density of a reference signal portion such as a preamble, and error vector amplitude of the received signal. Other examples of radio wave feature quantities include IQ phase (in-phase / quadrature phase) error and IQ imbalance. Alternatively, features indicating one or more of frequency offset and symbol clock error may be used as radio wave feature quantities. However, the examples of radio wave feature quantities given here are not intended to limit the feature quantities used by the present system 10 to identify a transmitting terminal.
[0096] The matching unit 130 can include a sample feature extraction unit 132 , a threshold determination unit 133 , a first feature matching unit 134 , a template feature storage unit (template feature storage unit) 135 , and an output unit 136 .
[0097] The matching unit 130 has the function of the matching unit 3a in Fig. 3 and performs the matching decision 11e described in Fig. 4. That is, the matching unit 130 matches the transmitting terminal by matching the sample feature extracted from the radio wave feature generated by the radio wave feature generation unit 112 from the received signal received by the receiving unit 111 with the pre-registered template feature, and outputs the matching result. Matching can be performed, for example, by calculating the similarity between the sample feature and the template feature and comparing the similarity with a matching threshold Th1. The following assumes that matching is performed by similarity calculation, but other matching methods can also be used. The pre-registered template feature is the template feature stored as a database in the template feature storage unit 135.
[0098] In this way, the matching unit 130 performs terminal matching (individual identification, model identification, attribute identification) based on the generated features. This matching process is performed by the first feature matching unit 134. The output unit 136 is responsible for outputting the matching result by the first feature matching unit 134.
[0099] The first feature matching unit 134 calculates one-to-L similarities between sample features and pre-registered template features, and performs matching by comparing each of the calculated L similarities (e.g., similarity scores) with a matching threshold Th1. Note that L is a positive integer. If there is a terminal having template features with a similarity greater than the matching threshold Th1 (if the terminal data exists in the database), the first feature matching unit 134 outputs the matching result (e.g., the ID of the transmitting terminal that is the source of the identified radio waves) from the output unit 136. If there are multiple terminals with template features greater than the matching threshold Th1, the first feature matching unit 134 may output the largest of these as the matching result, or may output a predetermined number (2 or more) as candidates. However, the similarity calculation may also end the matching process by determining that matching is successful when a similarity greater than the matching threshold Th1 is found.
[0100] Various methods can be used to calculate the similarity between sample features and template features, such as cosine similarity, Euclidean score (Euclidean distance), Mahalanobis distance, Manhattan distance, and correlation coefficient. Of course, a combination of these methods can be used to calculate the similarity. Methods other than the methods for calculating the similarity exemplified here can also be employed. The similarity can be calculated as a similarity score, as shown in some examples. While an explanation of an example of similarity calculation is omitted, the technology described in Patent Document 3, for example, can be applied. The more similar the features are, the higher the similarity (closer to 1) will be output, and the more different they are, the lower the similarity (closer to 0) will be output. The following explanation will be based on this premise, but the present invention is not limited to this.
[0101] The threshold determination unit 133 can calculate the following curves based on the results of similarity calculation between sample features generated from the learning data (dataset with correct answer labels) used to train the learning model and template features already registered in the template feature storage unit 135. The curves calculated here can be a curve of false acceptance rate and a curve of false rejection rate, and the technology described in Patent Document 3, for example, can be applied.
[0102] Then, the threshold determination unit 133 can determine a matching threshold Th1 to be used in the first feature matching unit 134 to compare the similarity between the sample feature and the template feature based on, for example, a predetermined false acceptance rate and a target error rate for matching.
[0103] Furthermore, the threshold determination unit 133 can have a function of setting a clustering threshold Th2 to be used in the first clustering process in the clustering unit 148, which will be described later, in accordance with an operation input by an operator or automatically. That is, the threshold determination unit 133 can control the clustering unit 148 to change the clustering threshold Th2. In this way, by making the clustering threshold Th2 variable, the first clustering process can be executed using a plurality of clustering thresholds Th2. Here, since an appropriate clustering threshold Th2 may change each time depending on the input unknown signal or unknown radio wave feature, making the clustering threshold Th2 variable is also beneficial in this respect.
[0104] The threshold determination unit 133 outputs the matching threshold Th1 to the first feature amount matching unit 134 and the clustering threshold Th2 to the clustering unit 148. For convenience of explanation, Fig. 5 shows the threshold determination unit 133 as if it were not connected to the first feature amount matching unit 134 and the clustering unit 148, but they are actually connected to each other.
[0105] Furthermore, by combining multiple radio wave features, i.e., by increasing the dimensionality of the features, it is possible to expect an improvement in matching accuracy. However, this may increase the amount of matching calculations and the database size. Therefore, the sample feature extraction unit 132, which will be described later, extracts sample features by extracting lower-dimensional sample features from the high-dimensional radio wave features generated by the radio wave feature generation unit 112.
[0106] In particular, when the receiving unit 111 receives a signal, the sample feature extraction unit 132 can generate sample features from radio wave features using a learning model. This learning model is a model for extracting sample features from the radio wave features generated by the radio wave feature generation unit 112, and is generated by the learning unit 113. The sample feature extraction unit 132 can be provided with this learning model or configured to be able to access this learning model. Furthermore, this learning model can also be installed in the sample feature extraction unit 132, with learning parameters set for hardware serving as a feature extractor.
[0107] The learning unit 113 performs machine learning using training data in which correct labels are assigned to radio wave features to generate a learning model. The training data used for training can be training data for a transmitting terminal whose transmitting device is known (the correct label is determined), or training data for a transmitting terminal whose template features are registered. The learning unit 113 can use any machine learning or deep learning algorithm, such as a support vector machine, boosting, or neural network. Since well-known technologies can be used for the algorithms, such as the support vector machine, a description thereof will be omitted. The correct label represents the transmitting terminal (wireless terminal), and can be, for example, a model name, an individual ID, or a serial number. In other words, the correct label is information for identifying and specifying the transmitting terminal. In machine learning, when constructing a learning model, a combination of features and correct labels is provided as a training dataset.
[0108] Furthermore, it is desirable that the learning unit 113 trains the learning model using features to which appropriate correct labels have been assigned, which have been acquired in advance in an environment where only a sufficient number of specific transmitting terminals transmit radio waves. In other words, it is desirable to generate the learning model by learning in advance the relationship between the transmitting terminals and the radio wave features of the signals transmitted by the transmitting terminals in an ideal environment (an environment where there are no terminals other than the terminal to be trained). For example, in the example of FIG. 6, the learning model is generated using radio waves (signals) transmitted by transmitting terminal 900a in an environment where transmitting terminal 900b does not exist.
[0109] If the sample feature extraction unit 132 includes a feature extractor, the learning parameters of the generated learning model are set in the feature extractor. The learning parameters may be, for example, a network configuration, weights, biases, etc., as long as they represent the learning model.
[0110] The feature clustering unit 140 is an example of the first clustering unit 1c in FIG. 3, and can include a sample feature temporary storage unit 146, a second feature matching unit 147, and a clustering unit 148.
[0111] The sample feature temporary storage unit 146 is a temporary storage unit that temporarily stores sample features that are determined to be unknown or unregistered in the matching by the matching unit 130. In other words, the sample feature temporary storage unit 146 temporarily stores sample features that have failed matching by the first feature matching unit 134 (all of the similarities with registered template features are smaller than the matching threshold Th1).
[0112] The second feature matching unit 147 matches the temporarily stored sample features with each other and outputs the matching result to the clustering unit 148. For example, the second feature matching unit 147 calculates the similarity between the sample features held in the sample feature temporary holding unit 146 at predetermined intervals.
[0113] Furthermore, the similarity calculated by the second feature matching unit 147 can be calculated using a variety of methods, such as cosine similarity, Euclidean score, Mahalanobis distance, Manhattan distance, and correlation coefficient. Of course, similarity calculation can also be performed by combining two or more of these methods. Furthermore, methods other than the similarity calculation methods exemplified here can also be employed. This similarity can also be calculated as a similarity score. Similar to the first feature matching unit 134, the second feature matching unit 147 can also match sample features using methods other than similarity calculation.
[0114] The clustering unit 148 performs a first clustering process based on the similarities between the sample features output from the second feature matching unit 147, and groups those whose similarity is equal to or greater than a clustering threshold Th2. Of course, there may be multiple groups, i.e., multiple clusters. Specifically, in the first clustering process, clustering is performed so that sample features with a similarity equal to or greater than the clustering threshold Th2 are grouped in the same cluster, and sample features with a similarity equal to or less than the clustering threshold Th2 are grouped in different clusters. The clustering unit 148 then assigns a cluster number to each cluster, reads out the target sample features from the sample feature temporary storage unit 146, and outputs them to the matrix generation unit 151.
[0115] The clustering unit 148 obtains sample features (or statistical values of sample features) for each cluster in this way and outputs them to the matrix generation unit 151. Sample features for a certain cluster can be output together with intensity values for node numbers, as shown in the graph of FIG. 7, for example. Here, the node numbers are numbers (orders) that identify the sample features included in that cluster, and the intensities indicate the feature values for each node number for calculating similarity. In this example, the number of dimensions of the sample features is 16.
[0116] The number of dimensions of the sample features (number of types of information) is not important, as in the description of the number of dimensions of the unknown radio wave features in the first embodiment. The sample features can be extracted as, for example, 16-dimensional values obtained by dimensionally compressing the final fully connected layer from the learning model by supervised learning. Note that during supervised learning, learning is performed such that, for example, this 16-dimensional value is associated with the correct label in the final layer.
[0117] For convenience, in FIG. 7, statistical values (e.g., median, average, etc.) of the intensities of the multiple sample features included in the cluster are plotted as line graphs (thick solid lines). Furthermore, the clustering unit 148 may output, for example, an intensity value for each of the multiple sample features included in the cluster. For convenience, in FIG. 7, the graph is shown as a diagonally shaded area surrounded by a thin solid line, and this diagonally shaded area includes a group of graphs of the intensities of each sample feature for the number of occurrence frequencies. Visualizing the sample features of each cluster (diagonally shaded area) in this way can be used to estimate the reliability of the cluster, the distance between different clusters, etc.
[0118] The specific technique for the first clustering process is not limited to the clustering method based on similarity comparison as described above, and machine learning techniques such as k-means, x-means, vbgmm, etc. Also, the first clustering process may use a hierarchical clustering method using a group average method, Ward's method, minimum distance method, maximum distance method, etc.
[0119] 3, and acquires M temporary labels estimated by the estimation device. Each temporary label is associated with an unknown signal that is the same as the unknown signal received by the receiving unit 111 (or an unknown radio wave feature generated from the unknown signal by any method).
[0120] The matrix generation unit 151 is an example of the generation unit 1e in Fig. 3, the visualization unit 152 is an example of the display unit described in the first embodiment, and the label setting unit 154 is an example of the label setting unit 1f in Fig. 2. The matrix generation unit 151 and the visualization unit 152 correspond to the relationship visualization unit 13b in Fig. 4.
[0121] The matrix generation unit 151 receives the temporary labels estimated by the estimation device and the first clustering results that are the processing results of the clustering unit 148, and generates a K×M relationship matrix that indicates the relationship between the two. The temporary labels and the first clustering results can be associated with each other using the original unknown signals or unknown radio wave features as keys (common terms). The generated relationship matrix can be, for example, as shown in FIG. 8, in which the frequency values of the unknown signals or unknown radio wave features common to the temporary labels and the cluster numbers are input at the intersections between them.
[0122] The visualization unit 152 visualizes the generated relationship matrix by displaying it on a display unit. This allows the operator to visually recognize the relationship matrix. The visualization unit 152 can also have a function to display the relationship matrix and the first clustering results, or only the first clustering results, on the display unit. This allows the operator to visually recognize the degree of similarity between the sample features included in each cluster and confirm the reliability of the cluster. The first clustering results can be displayed, for example, as a distribution graph of each cluster or a graph showing the strength (degree of similarity) of each cluster, as shown in FIG. 7.
[0123] Furthermore, the feature amount clustering unit 140 can vary the clustering threshold Th2, which allows the matrix generation unit 151 to output multiple relationship matrices. That is, the feature amount clustering unit 140 can be configured to execute the first clustering process multiple times with different clustering thresholds Th2 to obtain multiple first clustering results. Then, the matrix generation unit 151 can be configured to generate a relationship matrix for each of the multiple first clustering results.
[0124] The upper limit of the number of clusters output as the first clustering result by feature amount clustering section 140, that is, the upper limit of the number of clusters classified in the first clustering process, can also be set in advance.
[0125] The label setting unit 154 corresponds to a part of the data generation unit 13c in FIG. 4 and is an example of the label setting unit 1f in FIG. 2. The label setting unit 154 includes an operation unit that accepts operation input from an operator, and sets correct labels to some of the K×M intersections of the relationship matrix in accordance with the operation input. The operator can set correct labels for highly reliable intersections shown in the relationship matrix, thereby setting correct labels to data linked to the intersections (which may be multiple intersections). For example, the operator can recognize that the larger the value at an intersection (for example, a frequency value such as the number of data corresponding to the intersection or other statistical value), the higher the reliability of the intersection.
[0126] The data generation unit 153 corresponds to a part of the data generation unit 13c in Fig. 4 and is an example of the data generation unit 1g in Fig. 2. The data generation unit 153 generates additional training data based on the correct labels that the operator has determined to be highly reliable and set for each of some of the intersections, and the unknown signals or unknown radio wave features associated with each of these intersections. This additional training data is data for updating the training model of the training unit 113. One correct label can also be set for multiple intersections.
[0127] For example, this additional training data may include data such as unknown signals associated with temporary labels or cluster numbers corresponding to one or more intersections determined to be highly reliable, and the set correct labels (generated labels). In this way, the data generation unit 153 can generate one set of training data for each set correct label. This set may include data (raw data sets, or data of statistical values such as averages and medians) obtained from one or more unknown signals or unknown radio wave features corresponding to one or more intersections for which correct labels have been set, and the correct labels.
[0128] The learning unit 113 updates the learning model by performing machine learning based on this additional learning data and original learning data, which is learning data including the second information that was the learning target data in the original learning model and the correct labels linked to the sample features of the known signal. Here, even if the original learning data is not input during this update, it may be possible to update the learning model by inputting the additional learning data into the learning model that essentially uses the original learning data.
[0129] Next, specific examples will be given with reference to Fig. 9 to Fig. 11. Fig. 9 is a schematic diagram for explaining an overview of the learning data generation process performed for an unknown transmitting terminal in the present system 10. Fig. 10 is a schematic diagram showing an example of the result of changing the clustering threshold Th2 in the learning data generation process of Fig. 9, and Fig. 11 is a schematic diagram for explaining an example of the effect of re-learning using additional learning data generated as a result of the learning data generation process of Fig. 9.
[0130] FIG. 9 illustrates an example in which the following processing is executed. That is, an example is given in which a group of signals to be processed is subjected to a transmission device matching process in the matching unit 130 of the present system 10 shown in FIG. 5, feature amount clustering by the feature amount clustering unit 140, and visualization processing by the matrix generation unit 151 and visualization unit 152. Here, the group of signals to be processed is not assigned with information on correct labels, and is an unknown signal in the sense that it is before matching. Furthermore, in order to generate a relationship matrix to be subjected to the visualization process, it is first necessary to input a temporary label. The temporary label can be generated by the matching device 500 and output to the present system 10, and can be acquired by the temporary label acquisition unit 150. The matching device 500 is a device that functions as the estimation device described above.
[0131] The matching unit 130 can perform matching processing on the power spectral density of the reference signal portion of the received signal as an example of the radio wave feature of the received signal, and Fig. 9 shows such an example. Looking at the spectrogram in Fig. 9, in which each power spectral density (each spectrum) is arranged in time series, the target signal group includes unregistered signals and registered signals. Of these, unknown signals from unregistered transmitting terminals that have not yet been learned are the targets of processing by the feature clustering unit 140. Here, an example is shown in which the feature clustering unit 140 classifies the unknown signals into four clusters (cluster numbers 00, 01, 02, and 03).
[0132] Furthermore, in the example of FIG. 9, the collation device 500 estimates that there are six signals a to f in the same target signal group, that is, there are six transmitting terminals corresponding to the six signals a to f.
[0133] Specifically, the verification device 500 performs frequency band estimation processing and position estimation processing, and estimates signal a with almost certainty (with a fairly high degree of reliability) as being a signal from one individual. Furthermore, the verification device 500 estimates signals b to d as signals from three individuals, but with low reliability, and estimates signals e and f as signals from two individuals with certainty (with a high degree of reliability). In this case, the verification device 500 assigns temporary labels a to f to signals a to f, respectively, and can output the temporary labels to the system 10 with temporary label reliability information indicating the reliability of each temporary label. Note that this example illustrates an example in which temporary labels are assigned as a result of estimating both the frequency band and the position. In this example, the same temporary label is assigned to the same signal or radio wave feature based on both the frequency band estimation result and the position estimation result, and if temporary label reliability information is to be added, it can be added to that temporary label. In other words, two or more different temporary labels are not assigned to a given signal or radio wave feature. An example of temporary label reliability information will be described later with reference to FIG. 10.
[0134] Here, the temporary labels a to f can also include temporary labels assigned to known signals transmitted by a transmitting device registered in the present system 10. Of the temporary labels a to f received by the present system 10 from the matching device 500, signals that have been registered on the present system 10 side can be excluded from the targets for generating the relationship matrix by comparing with corresponding signals, for example. However, even without such exclusion, the signals can be included as they are in the relationship matrix and the operator can make a decision.
[0135] The temporary label acquisition unit 150 receives the temporary labels from the matching device 500 and passes them to the matrix generation unit 151. The temporary labels are received together with the signal or radio wave feature associated with the temporary labels and reliability information, and the matrix generation unit 151 generates a relationship matrix indicating the relationship between the temporary labels a to f and the four cluster numbers based on the signal or radio wave feature. The visualization unit 152 displays the generated relationship matrix. The relationship matrix can be generated, for example, as shown in the table in FIG. 9. In this table, temporary label names a to f and cluster numbers 00 to 03 are arranged in a matrix, and the frequency of occurrence is indicated at the intersections. The intersections are as described with reference to FIG. 4, and values other than the frequency of occurrence, for example, can also be used.
[0136] In addition, temporary label reliability information can be displayed when an intersection is selected during display, or can be displayed in the name of the temporary label as "high reliability" or "reliability level: 5 (maximum)", allowing the operator to recognize the reliability of the temporary label.
[0137] As illustrated here, the estimated information can include temporary label reliability information associated with the temporary label, indicating the reliability of the temporary label, and the relationship matrix can be generated with the temporary label reliability information associated with the temporary label. Then, by displaying such a relationship matrix automatically or by an operator's operation, the operator can know the reliability of the temporary label.
[0138] The relationship matrix can also be generated by adding information indicating at least one of the sample feature values for each cluster and the inter-cluster distance between each sample feature value. The added information can be calculated by the clustering unit 148 of the feature clustering unit 140. The inter-cluster distance can be, for example, at least one of Mahalanobis distance, spatial mapping, similarity, and the mean and / or variance of each cluster. The added information can be displayed on the display unit together with the relationship matrix, and the timing of display can be any, such as when the relationship matrix is displayed or when an instruction operation from the operator is received.
[0139] Although various examples of clustering methods have been given, the information indicating the sample features of each cluster (feature information) and the information indicating the distance between clusters may change depending on the clustering method. For example, when vbgmm is used, Mahalanobis distances can be output. In addition, in the case of hierarchical clustering, information indicating a dendrogram may also be output.
[0140] In the example of the relationship matrix shown in FIG. 9 and FIG. 4, for 85 signals or radio wave features with temporary label a, 40, 40, and 5 radio wave features are classified into clusters with cluster numbers 00, 01, and 03, respectively. Also in this example, for 90 signals or radio wave features with temporary label b, all radio wave features are classified into cluster with cluster number 03. Furthermore, in this example, for 88 signals or radio wave features with temporary label c, 3, 5, and 80 radio wave features are classified into clusters with cluster numbers 00, 01, and 03, respectively. Furthermore, in this example, for 85 signals or radio wave features with temporary label d, 3, 2, and 80 radio wave features are classified into clusters with cluster numbers 00, 01, and 03, respectively. Furthermore, in this example, for 90 signals or radio wave features with temporary label e, 50 and 40 radio wave features are classified into clusters with cluster numbers 00 and 01, respectively. Furthermore, in this example, for the 90 signals or radio wave features with the temporary label f, all of the radio wave features are classified into the cluster with cluster number 02.
[0141] An example of the operator setting the correct label in the example of Fig. 9 will be described with reference to Fig. 10. Fig. 10 shows an example of changing the clustering threshold Th2 and the resultant change in the relationship matrix.
[0142] The operator assigns a correct label to the data associated with a highly reliable intersection shown in the relationship matrix. Of course, if there is no highly reliable intersection, it is possible not to assign a correct label at all. The operator can determine that the greater the value at the intersection (for example, a frequency value such as the number of data corresponding to the intersection or other statistical value), the higher the reliability of the intersection. The operator can also determine whether an intersection is highly reliable based on the temporary label reliability information attached to the temporary label.
[0143] Furthermore, by comparing multiple relationship matrices with different clustering thresholds Th2, the operator can determine whether an intersection has a high degree of reliability. Here, the clustering threshold Th2 can be determined, for example, based on the matching threshold Th1 used when matching sample features with template features. For example, the clustering threshold Th2 can be set equal to the matching threshold Th1, and can be varied between 0.8×Th1 and 1.2×Th1, for example.
[0144] In this example, the signal with temporary label a is a signal with a somewhat high degree of reliability that has been estimated as "almost certainly one individual" as a result of estimating the frequency band and position. After confirming this information, the operator determines that the reliability of this temporary label a is higher than the reliability of the first clustering result, and can set the correct label A so that the temporary label a becomes the correct label for the group of intersections of the box indicated by A. Note that for the correct label A, since the frequency of the cluster with cluster number 02 is 0, the result will be the same even if the correct label A is set excluding the intersections of the cluster with cluster number 02.
[0145] The low reliability of the first clustering result can be determined by the operator by, for example, confirming that slightly lowering the clustering threshold Th2 results in two clusters being merged into one cluster.
[0146] An example of temporary label reliability information and the reliability of the first clustering result will be described with reference to FIG. 10. In FIG. 10, a value of 1 to 5 indicating the reliability of each temporary label is added as temporary label reliability information next to each temporary label. In this example, 1 indicates the lowest reliability and 5 indicates the highest reliability. The temporary label reliability information can be expressed as a level like this, or in easily understandable terms such as high, medium, or low, but the expression method and the degree of reliability are not important. Of course, the display method of the temporary label reliability information is not limited to the example shown.
[0147] Although temporary label reliability information may not be added to the temporary labels, adding it and displaying it together with the relationship matrix allows the operator to accurately grasp the reliability of the temporary labels, i.e., the estimation accuracy of the estimation device. This makes it possible to avoid situations where an operator mistakenly assigns a correct label because they believe it to be highly reliable when in fact it is incorrect, and ensures that correct labels are assigned to the correct model / individual, thereby improving the accuracy of matching unknown signals.
[0148] 10 also shows the first clustering result (feature clustering result) and the generated relationship matrix when the clustering threshold Th2 is changed between 0.81 and 0.9. Note that the clustering threshold Th2 in this example is determined as 0.9×Th1 or 1.0×Th1 when the matching threshold Th1 is 0.9.
[0149] When the clustering threshold Th2 is 0.9, the results are classified into four clusters, as shown in the feature clustering result CL-2 and the relationship matrix RE-2. Note that Figure 9 also shows the state shown by the feature clustering result CL-2 and the relationship matrix RE-2. In contrast, when the clustering threshold Th2 is 0.81, the results are classified into three clusters, as shown in the feature clustering result CL-1 and the relationship matrix RE-1.
[0150] The example in FIG. 10 shows that when the clustering threshold Th2 is lowered from 0.9 to 0.81, two clusters (cluster numbers 00 and 01) are merged into one cluster (cluster number 10). That is, when the clustering threshold Th2 is 0.9 for each of the temporary label names a, c, d, and e, the clusters are separated into two clusters with cluster numbers 00 and 01. On the other hand, when the clustering threshold Th2 is lowered to 0.81 for each of the temporary label names a, c, d, and e, the clusters are merged into one cluster with cluster number 10. In this case, the reliability of the clustering can be determined to be low. On the other hand, if the clusters do not change even when the clustering threshold Th2 is changed, the reliability of the clustering can be determined to be high. As an example, the signal of the temporary label f, which will be described later, will be used. The operator can determine that the intersection (combination) of the temporary label f and cluster number 02 (cluster number 11), which are one cluster even when the clustering threshold Th2 is lowered, is a combination of a highly reliable temporary label and cluster.
[0151] In this way, the operator can determine the reliability of the clustering by visually checking the relationship matrix with the clustering threshold Th2 changed. To enable such visual checking, for example, the feature amount clustering result CL-1, the relationship matrix RE-1, the feature amount clustering result CL-2, and the relationship matrix RE-2 in FIG. 10 can be displayed on a display unit. The display unit may be configured to switch between the upper and lower sections of FIG. 10 by an operator's operation, or the entirety of FIG. 10 may be displayed on the display unit at once, or only the feature amount clustering results CL-1 and CL-2 may be displayed on the display unit. Alternatively, only the relationship matrices RE-1 and RE-2 in FIG. 10 may be displayed on the display unit.
[0152] 9, the operator can know from the temporary label reliability information that the signals with temporary labels b to d have a reliability of "1" as a result of estimating the frequency band and position, and are therefore low-reliability signals. After checking this information, the operator can determine that the reliability of these temporary labels b to d is at least lower than the reliability of the feature clustering result, and can reject any of the temporary labels b to d as correct labels. Although not shown in FIG. 9, if the operator recognizes that the reliability of the feature clustering result for cluster number 03 is high, the operator can also set correct labels for the three intersections between the temporary labels b to d and cluster number 03.
[0153] Furthermore, the operator can learn from the temporary label reliability information that the reliability of the signals with temporary label a and temporary label e is somewhat high (reliability level "4") and high (reliability level "5"), respectively. The operator can then determine that, in the example of Figure 9, they should definitely be separated into separate individuals, that is, into two clusters. Meanwhile, the relationship matrix RE-2 shows that the feature clustering results for cluster numbers 00 and 01 are classified into the same cluster. From this result, the operator can learn that the reliability of the feature clustering is low. Here, the low reliability of the feature clustering for this unknown signal means that the learning model has low reliability for this unknown signal.
[0154] Furthermore, the operator can know from the temporary label reliability information that the reliability of the signal with temporary label e is high. Based on this high reliability, the operator can determine that the frequency band and position estimation results are definitely one individual, that is, that they should be separated into one cluster. Meanwhile, in the relationship matrix RE-2, the feature clustering results are separated into cluster numbers 00 and 01. Therefore, the operator can learn from this result that the reliability of the clustering is low. Having confirmed this information, the operator can determine that the reliability of this temporary label e is higher than the reliability of the feature clustering, and set correct label B so that the temporary label e becomes the correct label for the group of intersections of the frame indicated by B.
[0155] However, the relationship matrix RE-1 shows that the temporary label e is integrated into a single cluster as a result of feature clustering performed with the clustering threshold Th2 lowered to 0.81. Therefore, it can be said that the reliability of the feature clustering results for this unknown signal is high with this clustering threshold Th2. Here, when generating a single training data set, the first clustering process is not performed with a different clustering threshold Th2 for each unknown signal. Therefore, the reliability of the feature clustering for the temporary label e should be determined while taking into account the results of the other temporary labels, and whether or not to set a correct label should be determined based on the determination result. On the other hand, from the perspective of generating more training data, it is possible to set a correct label for the temporary label e with the clustering threshold Th2 set to 0.81. For example, if the data of the unknown signal or unknown radio wave features can be corrected to fill the difference between the clustering thresholds Th2 of 0.81 and 0.9, the correct label and the data can be generated as a set of training data.
[0156] Furthermore, the operator can know from the temporary label reliability information that the reliability of the signal with temporary label f is high. Based on the high reliability, the operator can determine that the frequency band and position estimation results are definitely a single individual, i.e., that the signal should be divided into one cluster. Furthermore, since the signal with temporary label f is classified into one cluster (cluster numbers 02 and 11) even when the clustering threshold Th2 is lowered, the reliability of the clustering can also be said to be high. In other words, the operator can determine that the intersection (combination) of temporary label f and cluster number 02 (cluster number 11) is a combination of a highly reliable temporary label and cluster with a one-to-one correspondence. Having confirmed this information, the operator can determine that both the reliability of this temporary label f and the reliability of the feature clustering results are high, and can set a correct label C so that the temporary label f is the correct label only for the intersection (cell) of the frame indicated by C.
[0157] In this way, the data generation unit 153 generates additional training data in which the correct label C is added only to highly reliable data. Specifically, the data generation unit 153 generates additional training data for updating the training model 11b based on the correct label C and the unknown signal or unknown radio wave feature associated with the corresponding intersection. If the operator also determines that the tentative labels a and e have high reliability, they are similarly set as correct labels A and B, respectively, and additional training data is generated based on the unknown signal or unknown radio wave feature associated with the corresponding intersection. In this case, a total of three sets of additional training data can be generated for the correct labels A, B, and C. Here, the operator can generate additional training data for only the correct label C to obtain reliable training data, or generate additional training data for the correct labels A, B, and C to obtain a large amount of training data.
[0158] Although an example has been given here in which a correct label is set for an unknown signal or unknown radio wave feature that is lower than the matching threshold Th1, a correct label can also be set for a signal (i.e., a known signal) or its radio wave feature that is higher than the matching threshold Th1. This makes it possible to increase the amount of training data to which correct correct labels are assigned. In this case, the possibility that a temporary label will be set as a correct label increases.
[0159] Then, by performing re-learning using the additional training data generated in this way, the following effects can be achieved, for example: Fig. 11 shows the feature clustering result CL-3 and relationship matrix RE-3 before re-learning, and the feature clustering result CL-4 and relationship matrix RE-4 after re-learning.
[0160] Before re-learning, if a feature clustering result CL-3 is obtained and a relationship matrix RE-3 is generated, then additional training data is generated by setting correct labels A, B, and C to the temporary labels a, e, and f as described in Fig. 10. Then, the results of re-learning based on this additional training data are, for example, as shown by the feature clustering result CL-4 and the relationship matrix RE-4.
[0161] At the time of correct label assignment before relearning, the tentative labels a, e, and f have slightly high reliability (reliability of "4"), high reliability (reliability of "5"), and high reliability (reliability of "5"), respectively, as explained in Figure 10. At this point, the reliability taking into account the feature clustering results for the intersections indicated by A, B, and C can be said to be medium reliability, slightly high reliability, and high reliability, respectively.
[0162] In contrast, re-learning can improve the matching accuracy for unknown signals, as shown in the feature clustering result CL-4 and relationship matrix RE-4, for example. In the relationship matrix RE-4, the temporary labels a, e, and f are re-learned with correct labels assigned. Therefore, the feature clustering results for the temporary labels a, e, and f are one-to-one, and it can be said that the accuracy of extracting sample features for unknown signals corresponding to these temporary labels can be improved. Increasing the accuracy of extracting sample features means that accurate matching is possible if template features are registered for those sample features. Therefore, it can be said that re-learning has improved the matching accuracy for signals corresponding to the temporary labels a, e, and f.
[0163] Furthermore, an important point is that in the relationship matrix RE-4 after re-learning, there is also a one-to-one relationship with the feature clustering results for temporary label b, indicating that the matching accuracy for unknown signals has improved. Specifically, before re-learning, the estimation results for whether signals with temporary labels b, c, and d refer to three individual transmitting terminals were unreliable, and there was a risk that the matching accuracy would be low if the learning model was used as is. In contrast, by adding signals with other temporary labels a, e, and f and re-learning, temporary label b is classified into a single cluster. This means that if the signals with temporary labels b, c, and d are actually all signals from a single transmitting terminal, it can be said that the matching accuracy for unknown signals has improved.
[0164] In this way, in this system 10, by visualizing the relationship between the feature clustering results of transmitting terminal matching and the temporary labels estimated by a separate estimation device, highly reliable labeling becomes possible, and learning using signals from unknown transmitting terminals becomes possible. As a result, this system 10 can improve matching accuracy.
[0165] Furthermore, the present system 10 can be configured to perform the above-described processing on currently receivable unknown signals, and if, for example, a 1:1 correspondence is detected, notify the operator that no further generation of additional learning data is necessary.
[0166] Next, a processing example and effect of the present system 10 will be described with reference to Fig. 12. Fig. 12 is a flow chart for explaining a processing example in the present system 10.
[0167] In the system 10, the receiving unit 111 receives a signal transmitted from a transmitting terminal, the radio wave feature generating unit 112 generates a radio wave feature from the signal, and the matching unit 130 performs a matching process (step S101). In the matching process, first, the sample feature extracting unit 132 extracts a sample feature from the radio wave feature using a learning model. In the matching process, the first feature matching unit 134 then calculates the similarity between the sample feature and the template feature stored in the template feature storage unit 135, and performs matching by comparing the similarity with a matching threshold Th1.
[0168] The first feature matching unit 134 determines whether the sample features have been registered as a result of the matching process (step S102). If the result is YES, that is, if the sample features are registered, the matching result is output from the output unit 136 (step S103), and the process ends. If there is a template feature whose similarity is greater than the matching threshold Th1, the first feature matching unit 134 determines that matching is successful and outputs the matching result, including the ID of the transmitting terminal that is the source of the radio waves, from the output unit 136. On the other hand, if the result is NO in step S102, that is, if the sample features are unregistered, the processes of steps S104 to S111 are executed. Note that the processes from step S102 onwards are executed for all unregistered sample features. Furthermore, if the result is NO in step S102, that is, if the sample features are unregistered, it is also possible to output this as the matching result.
[0169] If the result of step S102 is NO, first, the first feature matching unit 134 stores the target sample feature as a matching result in the sample feature temporary storage unit 146 (step S104). The loop process between steps S105s and S105e is executed for the sample feature stored here.
[0170] In this loop process, first, a clustering threshold Th2 is set (step S106). Note that the clustering threshold Th2 used initially as an initial setting can be set in advance, and from the next time onwards, the clustering threshold Th2 is changed by a predetermined number of changes, for example, by increasing and / or decreasing it at predetermined intervals. Note that in the example of Fig. 10, this change is made once, and as a result, a total of two relationship matrices are generated in step S109, which will be described later.
[0171] Furthermore, the clustering threshold Th2 is a threshold for clustering performed to generate additional training data to be used for matching, and is therefore set to a value that is not far from the matching threshold Th1. Therefore, the clustering threshold Th2 can be set within a range of, for example, ±10% or ±5% of the matching threshold Th1.
[0172] Next, the feature clustering unit 140 performs a first clustering process on the sample features using the clustering threshold Th2 set in step S106 to obtain a feature clustering result (step S107). The first clustering process can be performed by, for example, the second feature matching unit 147 and the clustering unit 148 as described above.
[0173] Thereafter, the visualization unit 152 visualizes the feature clustering results in a graph format, such as a relationship matrix RE-3 (step S108). Next, the matrix generation unit 151 generates a relationship matrix indicating the relationship between the temporary label information, which is the estimation result by the estimation device, and the feature clustering results, and passes the relationship matrix to the visualization unit 152 to display it (step S109). This ends the above-mentioned loop processing.
[0174] Next, the operator checks the information visualized by the visualization unit 152 and considers whether there are any intersections to which a correct label can be set, i.e., highly reliable intersections. Here, the operator can recognize the reliability of the temporary label at the time of receiving the temporary label from temporary label reliability information or the performance of the estimation device known in advance. On the other hand, the operator does not know the reliability of the feature clustering, and therefore there are cases where the reliability of the feature clustering cannot be grasped only from the relationship between the temporary label and the feature clustering result.
[0175] Therefore, in the present system 10, the loop processing described above is configured to allow the operator to check the feature clustering results while changing the clustering threshold Th2. This allows the operator to grasp the reliability of the feature clustering and make appropriate choices when setting the correct label, thereby increasing the reliability. Furthermore, as described above, the appropriate value of the clustering threshold Th2 varies depending on the target unknown signal. However, the operator can recognize that the clustering threshold Th2 has been set to an appropriate value, for example, from at least one of the relationship with the temporary label and the temporary label reliability information. From this perspective, configuring the clustering threshold Th2 to be variable can also increase the reliability of the feature clustering.
[0176] As a result of the investigation, the operator performs an operation to set a correct label for the intersection or intersection group for which the operator has determined that a correct label may be set, to the label setting unit 154, and the label setting unit 154 sets the correct label in accordance with the operation (step S110). Finally, additional learning data is generated based on the set correct label and the data of the unknown signal or unknown radio wave feature associated with the intersection or intersection group (step S111), and the process ends.
[0177] Furthermore, while step S106 is based on the premise that the clustering threshold Th2 is automatically changed a predetermined number of times, this is not limitative. For example, instead of changing the clustering threshold Th2 a predetermined number of times, the clustering threshold Th2 can be increased and / or decreased at predetermined intervals until a change occurs in the feature clustering results. Here, a change in the feature clustering results can be automatically determined according to a predetermined determination criterion. The predetermined determination criterion can refer to, for example, until a change occurs in the number of clusters, until a change occurs in the number of clusters corresponding to at least one temporary label, or until a change occurs in the number of clusters corresponding to a temporary label indicated with a predetermined reliability. Alternatively, the clustering threshold Th2 may be changed up to a predetermined number of times until a change occurs in the feature clustering results.
[0178] Furthermore, for example, the operator may be requested to change the clustering threshold Th2 a predetermined number of times or until there is a change in the clustering results, and the change may be implemented in response to the change. Alternatively, step S106 may be implemented by the operator inputting or selecting the clustering threshold Th2 each time. In this case, the loop processing can be terminated when the setting of the necessary correct answer labels is completed without the operator setting a new clustering threshold Th2. Alternatively, after the correct answer label is set for a certain intersection, the processing from step S106 onwards can be executed again by the operator changing the clustering threshold Th2.
[0179] Furthermore, as long as correct labels can be set in steps S110 and S111 and additional training data can be generated based on the correct labels, the processing flow is not limited to that described in Fig. 12. For example, if the reliability of the feature clustering results is not important when setting correct labels, or if the evaluation can be performed simply by checking the relationship matrix, the cluster visualization processing in step S108 can be omitted.
[0180] With the above configuration, the system 10 can generate additional training data that has been accurately labeled by an operator based on the results of checking the relationship matrix, and can then use that additional training data for re-training. Therefore, the system 10 can also improve the matching accuracy of unknown transmitting terminals that are unrelated to the transmitting terminals added as training data. For example, the system 10 is expected to improve the mismatch rate after re-training by about 15% compared to the mismatch rate before re-training.
[0181] 5 to 12, the description is based on the assumption that the matching device 500 generates estimated information using a predetermined estimation method, but this is not limited to this. For example, if such a method does not improve the matching accuracy or for the purpose of verification, the present system 10 can also adopt the following alternative configuration example 1. In alternative configuration example 1, the estimation method is changed so that the matching device 500 generates estimated information using a new estimation method or by adding a new estimation method. This can further narrow down the targets for which correct labels are set, improve reliability, and be used during verification.
[0182] In Alternative Configuration Example 1, the system 10 and the verification device 500 are connected so that an instruction to change the estimation method can be given to the verification device 500 in response to an instruction from, for example, an operator. In response to the instruction, the verification device 500 changes or adds an estimation method and returns estimated information. For example, as initially described with reference to FIG. 9 , the verification device 500 outputs estimated information as a result of estimating the bandwidth and location. Upon receiving a change instruction from the system 10, the verification device 500 changes the estimation method as follows. That is, upon receiving the change instruction, the verification device 500 returns estimated information as a result of estimation using a transmission power estimation function, or estimated information as a result of estimation using both the bandwidth and location estimation function and the transmission power estimation function. Here, the estimation methods before and after the change can include at least one of the functions listed as the functions of the estimation unit 12. Furthermore, for clusters that do not have sufficient reliability in the matrix based on the feature clustering results and the location estimation, further clustering using bandwidth estimation may be performed.
[0183] Furthermore, in a configuration in which correct labels are set using temporary labels based on the feature clustering results and position estimation, verification can be performed by verifying whether or not there are any problems with the correct labels using temporary labels based on the feature clustering results and band estimation. Similarly, clustering based on the feature clustering results and band estimation may be further clustered using position estimation. Furthermore, clustering based on the feature clustering results and band estimation may be verified using clustering based on the feature clustering results and position estimation.
[0184] Furthermore, in Alternative Configuration Example 1, an estimation method is selected from among the functions originally provided in the matching device 500, but as a further alternative, Alternative Configuration Example 2, it is also possible to configure the system 10 to also input estimated information resulting from estimation by a matching device other than the matching device 500. As a result, in Alternative Configuration Example 2 of the system 10, it is possible to generate a relationship matrix based on both or one of the pieces of estimated information and set a correct label. Of course, the above-mentioned separate matching devices may be two or more.
[0185] 5 to 12 and alternative configuration examples 1 and 2, the verification device 500 receives a signal radio wave and performs estimation and temporary label assignment based on the received signal, but this is not limited to this. As alternative configuration example 3, the verification device 500, or the verification device 500 and the other verification device, can also receive an unknown signal from the system 10 via a wired network, etc.
[0186] That is, the system 10 may transmit a signal or its radio wave feature determined to be unregistered by the system 10 to the verification device 500 via a wired network or the like, and the verification device 500 may estimate the received signal and assign a temporary label to it. In this case, the resulting temporary label assigned by the verification device 500 will be the temporary labels a to f for the unknown signal regarded as an unregistered signal by the system 10, and even in this case, the number of temporary labels may end up being the same as or different from the number of clusters.
[0187] In this way, the alternative configuration example 3 of the present system 10 can transmit the n unknown signals received by the receiving unit 111 or the n unknown radio wave features generated by the radio wave feature generating unit 112 based on the n unknown signals to an estimation device such as the matching device 500. Then, the estimation information acquiring unit exemplified as the temporary label acquiring unit 150 can input the m pieces of estimation target information by receiving the m pieces of estimation target information from the estimation device.
[0188] By adopting such alternative configuration example 3, it is possible to perform processes such as generating a relationship matrix using the clustering results and the estimation results for the same unknown signal or radio wave feature, and the information included in the intersections of the relationship matrix can also be made more accurate. Note that all of alternative configuration examples 1 to 3 can also be applied to the first embodiment, which is not limited to including temporary labels in the estimation information.
[0189] <Third embodiment> The third embodiment will be described with reference to Fig. 13, focusing on differences from the second embodiment. However, the various examples described in the first and second embodiments can also be applied to the third embodiment. Fig. 13 is a block diagram showing an example of the functional configuration of a transmission device verification system according to the third embodiment. Note that, among the components shown in Fig. 13, those with the same names as the components described in Fig. 5 basically have similar functions, and therefore, descriptions of the similar functions will be omitted except for some.
[0190] In a transmission device matching system 20 (hereinafter referred to as the present system 20) according to this embodiment shown in FIG. 13, information is input in the form of radio wave features before temporary labeling, rather than in the form of temporary labels. Then, in the present system 20, the radio wave features before temporary labeling are clustered in the same manner as the matching sample features, and a relationship matrix is then generated. In this embodiment, information equivalent to the temporary labels is estimated by clustering from the radio wave features before temporary labeling, and it can be said that the present system 20 includes an estimation device.
[0191] Therefore, the present system 20 includes a feature acquisition unit 250 and a second feature clustering unit 260 in place of the temporary label acquisition unit 150 in the transmission device verification system 10. Furthermore, the present system 20 includes a first feature clustering unit 240 in place of the feature clustering unit 140 in the transmission device verification system 10.
[0192] 1 to 3, and includes a sample feature temporary storage unit 146, a second feature matching unit 147, and a clustering unit 148, as well as a threshold variable control unit 241. The threshold determination unit 133 in FIG. 5 only sets the matching threshold Th1. The threshold variable control unit 241 corresponds to a part that sets (controls the change of) the clustering threshold Th2 in the threshold determination unit 133.
[0193] The feature amount acquiring unit 250 and the second feature amount clustering unit 260 are an example of the estimated information acquiring unit 1d in Figures 1 to 3. The second feature amount clustering unit 260 is a unit that executes a second clustering process, and can be called a second clustering unit.
[0194] The feature acquisition unit 250 inputs m radio wave features as m pieces of estimation target information. The second feature clustering unit 260 executes a second clustering process on these m radio wave features and obtains M pieces of estimation information as the clustering result (second clustering result). This second clustering process is performed by the clustering unit 262. The second clustering process may be exactly the same as the first clustering process, or may be a process in which only the clustering threshold is changed, or a process in which the algorithm is changed.
[0195] However, the clustering unit 262 outputs information including M temporary labels indicating each of the M transmitting terminals as the M pieces of estimated information. The clustering unit 262 can assign a temporary label to each of the clusters resulting from the second clustering process, and can assign five temporary labels when the data are classified into five clusters, for example. The temporary labels of this embodiment are generated using a different procedure from the temporary labels of the second embodiment, but can be said to have basically the same meaning.
[0196] Also in this embodiment, the clustering unit 148 can output information indicating at least one of the sample feature for each cluster and the inter-cluster distance of each sample feature, and add it to the relationship matrix. Similarly, the clustering unit 262 can output information indicating at least one of the radio wave feature for each cluster (the radio wave feature acquired by the feature acquisition unit 250) and the inter-cluster distance of each radio wave feature, and add it to the relationship matrix.
[0197] The estimated information obtained by the clustering unit 262 may also include temporary label reliability information for the temporary label. In this embodiment, the feature acquisition unit 250 can acquire reliability information for each of the m radio wave features in association with the m radio wave features. The clustering unit 262 can generate temporary label reliability information for one or more radio wave features included in each cluster resulting from the second clustering process performed on the radio wave features, i.e., for one or more radio wave features to which each temporary label is assigned. The temporary label reliability information can be generated based on the reliability information acquired by the feature acquisition unit 250. For example, if a certain temporary label includes four radio wave features, the clustering unit 262 can calculate temporary label reliability information for the temporary label by performing statistical processing such as averaging or median calculation on the reliability information associated with the four radio wave features. Then, the clustering unit 262 can associate the temporary label reliability information with the temporary label when outputting it.
[0198] The second feature clustering unit 260 can also include a threshold variable control unit 261 that controls changing the clustering threshold Th3 in the clustering unit 262. The threshold variable control unit 261 can set the clustering threshold Th3 in accordance with an operator's input or automatically, that is, can control the clustering unit 262 to change the clustering threshold Th3. By making the clustering threshold Th3 variable in this way, the second clustering process can be performed using a plurality of clustering thresholds Th3. Here, an appropriate clustering threshold Th3 may change each time depending on the input unknown signal or unknown radio wave feature, and therefore making the clustering threshold Th3 variable is also beneficial in this respect.
[0199] The matrix generation unit 151 and the units performing subsequent processing are similar to those in the transmission device verification system 10 of Fig. 5. Briefly explained, in this system 20, a relationship matrix is generated based on temporary labels or temporary labels and temporary label reliability information as the second clustering result and the first clustering result in the clustering unit 148. Furthermore, examples of display of the first clustering result, the second clustering result (temporary labels, etc.), and the relationship matrix can also be examples of display of the first clustering result, the second clustering result (temporary labels, etc.), and the relationship matrix, as exemplified in the second embodiment. Then, in this system 20, a correct label is set by the operator based on the relationship matrix, and additional learning data for the correct label is generated.
[0200] According to this embodiment, in addition to the effects of the second embodiment, even when the estimation device does not have a function to assign a temporary label or when the temporary label assigned by the estimation device does not meet the standards of the system 20 and is difficult to apply, it can be said that according to this embodiment, in addition to the effects of the second embodiment, it is possible to improve the versatility of a device that can be used as an estimation device.
[0201] Furthermore, in this system 20, by making the clustering threshold Th2 variable, it is possible to know whether the number of clusters will change, etc., so that the operator can check the reliability of the first clustering result and set a highly reliable correct label. Furthermore, by making the clustering threshold Th3 variable, it becomes possible to set an even more reliable correct label. Furthermore, it is also possible to automatically raise or lower at least one of the clustering thresholds Th2 and Th3, so that the computer can recognize whether the number of clusters of the temporary labels will change, and automatically set an appropriate threshold.
[0202] In this system 20, the threshold variable control unit 241 is not essential, and the clustering threshold Th2 can also be changed by the threshold determination unit 133, as in the second embodiment. However, as in the second embodiment, this embodiment does not exclude a configuration in which the clustering threshold Th2 cannot be changed. Similarly, in this system 20, the threshold variable control unit 261 is not essential.
[0203] Furthermore, in this embodiment, the system 20 has been described as an example of a transmission device verification system, but this embodiment can also be implemented as an example of the learning data generation system in FIG. 1 or an example of the learning system 2 in FIG.
[0204] <Fourth embodiment> The fourth embodiment will be described with reference to Fig. 14, focusing on the differences from the second embodiment. However, the various examples described in the first to third embodiments can also be applied to the fourth embodiment as long as they do not result in contradictory processing. Fig. 14 is a block diagram showing an example of the functional configuration of a transmission device verification system according to the fourth embodiment. Note that, among the components shown in Fig. 14, those with the same names as the components described in Fig. 5 and Fig. 13 basically have similar functions, and therefore, descriptions of the similar functions will be omitted except for some.
[0205] In a transmission device matching system 30 (hereinafter referred to as the present system 30) according to this embodiment shown in FIG. 14, information is input in the form of radio wave features before temporary labeling, rather than in the form of temporary labels. Then, the present system 30 standardizes (normalizes) the radio wave features before temporary labeling, or the extracted sample features and the radio wave features before temporary labeling. Then, the present system 30 clusters the normalized radio wave features before temporary labeling together with the matching sample features, and generates a relationship matrix. In this embodiment, information equivalent to the temporary labels is estimated by clustering from the radio wave features before temporary labeling, and the present system 20 can be said to include an estimation device.
[0206] Therefore, in the transmitting device matching system 10, the present system 30 has the feature acquisition unit 250 and some of the functions of the feature clustering unit 340 instead of the temporary label acquisition unit 150, and has the functions of the feature clustering unit 340 other than some of the functions described above instead of the feature clustering unit 140.
[0207] The feature amount clustering unit 340 includes a sample feature amount temporary storage unit 146, a second feature amount matching unit 147, a weight control unit 341, and a clustering unit 348. Of the feature amount clustering unit 340, the portion having functions other than the above-mentioned part of the clustering unit 348 is an example of the first clustering unit 1c in FIGS.
[0208] The above-described functional parts of the feature amount acquiring section 250 and the clustering section 348 are an example of the estimated information acquiring section 1d in FIGS.
[0209] The feature acquisition unit 250 inputs m radio wave features as m pieces of estimation target information. Then, as part of its functions, the clustering unit 348 executes a third clustering process on the m radio wave features, and obtains M pieces of estimation information as a result (hereinafter referred to as temporary label clustering result).
[0210] The clustering unit 348 outputs, as the M pieces of estimated information, information including M temporary labels indicating each of the M transmitting terminals. The temporary labels of this embodiment are generated using a different procedure than the temporary labels of the second embodiment, but the labels basically have the same meaning. The clustering unit 348 can assign a temporary label to each of the clusters resulting from performing the third clustering process on the m radio wave features. For example, if the m radio wave features are classified into five clusters, the clustering unit 348 can assign five temporary labels. That is, for input from the feature acquisition unit 250, the clustering unit 348 can output temporary labels instead of cluster numbers, or output cluster numbers with a predetermined symbol attached as temporary labels.
[0211] Furthermore, the clustering unit 348 can output the estimated information so as to include temporary label reliability information for the temporary labels. For this purpose, in this embodiment, the feature acquisition unit 250 can acquire reliability information for the radio wave features in association with each of the m radio wave features. The clustering unit 348 can generate temporary label reliability information for one or more radio wave features included in each cluster resulting from the third clustering process being performed on the radio wave features, i.e., one or more radio wave features to which each temporary label is assigned.
[0212] The temporary label reliability information can be generated based on the reliability information acquired by the feature acquisition unit 250. For example, when a certain temporary label includes four radio wave features, the clustering unit 348 can calculate the temporary label reliability information for the temporary label by performing statistical processing such as averaging or median calculation on the reliability information associated with the four radio wave features. Then, the clustering unit 348 can associate the temporary label reliability information with the temporary label and output it when outputting it.
[0213] Furthermore, the clustering unit 348 performs a first clustering process on the sample features output from the sample feature extraction unit 132 and the radio wave features acquired by the feature acquisition unit 250 to obtain a first clustering result. In other words, the clustering unit 348 can output both the temporary label clustering result and the first clustering result, and both clusterings to obtain these results can be performed at the same time or at different times.
[0214] However, in the first clustering process and the third clustering process in this embodiment, standardization (normalization) is performed on the radio wave features before temporary labeling and the extracted sample features by the weight control unit 341. Before describing the temporary label clustering results, the weight control unit 341 will be described.
[0215] The weight control unit 341 performs weight control processing (normalization processing and weighting processing) on the n sample features extracted by the sample feature extraction unit 132 and the m radio wave features input as m pieces of estimation target information by the feature acquisition unit 250. This normalization processing is processing known as a function such as a standard scaler or a min-max scaler, and is processing for eliminating the influence of the measurement units of both on the size of data used as input data for the first clustering processing and the third clustering processing. In other words, this weight control processing is processing for normalizing (standardizing) so that the clustering unit 348 can handle both feature values as values having the same maximum value or other magnitude for clustering, and then weighting one of the feature values. This normalization processing is also referred to as standardization processing. The weight control unit 341 can, for example, control the normalization processing by changing the weighting between the sample feature and the m radio wave features.
[0216] However, there may be cases where it is not necessary to originally perform normalization processing on the radio wave feature output from the estimation device and the sample feature output from the sample feature extraction unit 132. Therefore, normalization processing is not necessary in such cases, and therefore, it can be said that a configuration not including the weight control unit 341 can also be employed in this embodiment.
[0217] In this way, the clustering unit 348 performs the first clustering process on the n sample features and m radio wave features after the weight control process, and performs the third clustering process on the m radio wave features. As described in the first embodiment, the m radio wave features are information used to generate training data, and an estimation device with high estimation accuracy can be used. When the estimation accuracy is high, the number indicated by n and the number indicated by m will match. Therefore, in this case, the clustering unit 348 performs the first clustering process on the n sample features after the weight control process and the m radio wave features corresponding thereto, and performs the third clustering process on the m radio wave features.
[0218] Furthermore, the third clustering process can be performed using the same algorithm as the first clustering process, but with different input node dimensions and output nodes, and the clustering threshold can also be changed. The first clustering process and the third clustering process are described separately, but they can also be performed as a single clustering process. The first clustering process and the third clustering process in this embodiment can also use various algorithms exemplified as the algorithm for the first clustering process in the second embodiment. The first clustering process in this embodiment is basically the same as the first clustering process in the second embodiment, although the number of input dimensions increases by the number of dimensions of the radio wave features. The number of output dimensions (number of clusters) varies depending on the clustering results, making it difficult to make a general comparison with the second embodiment. However, the number of input dimensions may increase compared to the second embodiment.
[0219] Then, the clustering unit 348 can output the first clustering result and the temporary labels so as to distinguish between the combined features and the individual features as a result of performing the first and third clustering processes on the n sample features and m radio wave features after the weight control process. When outputting the clustering result of the clustering unit 348, for example, the clusters indicated by the first clustering result and the temporary labels on the same graph can be output so as to be distinguishable from each other.
[0220] The clustering unit 348 may perform the first and third clustering processes on the n sample features and m radio wave features after the weight control process and output the results so that each clustering result can be distinguished. For example, the clustering unit 348 may output two types of data: a first clustering result using both features (combined features) and a temporary label based on the temporary label clustering result (third clustering result). To achieve this, the clustering unit 348 performs the first and third clustering processes on the n sample features and m radio wave features after the weight control process, obtaining the result (clustering result with a large amount of data) and also obtaining the following result. That is, the clustering unit 348 obtains a clustering result with a large amount of data and also obtains a result of performing the third clustering process on only the m radio wave features after the weight control process (clustering result with a small amount of data).
[0221] However, even if normalization or weighting is performed on the feature amounts that are the basis for obtaining large clustering result data, normalization or weighting may not be performed on the m radio wave feature amounts for obtaining small clustering result data. This increases the likelihood that, for example, when the clustering results are displayed in a graph, the small clustering result data will be displayed in a manner that makes it easy to distinguish from the large clustering result data.
[0222] Furthermore, the clustering unit 348 in this embodiment can output information indicating at least one of the sample feature for each cluster and the inter-cluster distance for each sample feature, and add it to the relationship matrix. Furthermore, the clustering unit 348 can also output information indicating at least one of the radio wave feature for each cluster (the radio wave feature acquired by the feature acquisition unit 250) and the inter-cluster distance for each radio wave feature, and add it to the relationship matrix.
[0223] Furthermore, in this embodiment, the clustering threshold Th2 can include a clustering threshold Th2-1 for clustering the radio wave features acquired by the feature acquisition unit 250 and a clustering threshold Th2-2 for clustering the sample features. Alternatively, in the example described above in which a clustering result with a large amount of data and a clustering result with a small amount of data are obtained, the clustering threshold Th2 can include a clustering threshold Th2-1 for obtaining the former result and a clustering threshold Th2-2 for obtaining the latter result. Furthermore, as with the second embodiment, this embodiment does not exclude a configuration in which the clustering threshold Th2 cannot be changed.
[0224] The matrix generation unit 151 in the present system 30 generates a relationship matrix based on the clustering result output from the clustering unit 348. Specifically, the matrix generation unit 151 generates the relationship matrix based on the temporary labels or the temporary labels and temporary label reliability information, and the first clustering result for both feature amounts (combined feature amounts).
[0225] The units that perform processing downstream of the matrix generation unit 151 are basically the same as those in the transmission device verification system 10 of Fig. 5. Briefly explained, in this system 30, an operator sets a correct label based on the relationship matrix, and additional learning data is generated for the correct label. Furthermore, examples of displaying the clustering results and the relationship matrix can be the same as those exemplified in the second embodiment.
[0226] According to this embodiment, in addition to the effects of the second embodiment, the system 30 can assign temporary labels even when the estimation device does not have a function for assigning temporary labels or when the temporary labels assigned by the estimation device do not meet the standards of the system 30 and are difficult to apply. Furthermore, it becomes possible to visualize the relationship between the clustering result of the feature when the sample feature and the radio wave feature are combined, with the clustering result of the radio wave feature, which is equivalent to the temporary label to be assigned. Furthermore, it becomes possible to visualize the changes and trends in the relationship based on changes in weighting. Therefore, according to this embodiment, in addition to the effects of the second embodiment, it can be said that the versatility of the device that can be used as an estimation device can be improved.
[0227] Furthermore, in this system 30, by making the clustering threshold Th2-1 variable, it is possible to know whether the number of clusters will change, etc., so that the operator can check the reliability of the first clustering result and set a highly reliable correct label. Furthermore, by making the clustering threshold Th2-2 variable, it becomes possible to set an even more reliable correct label. Furthermore, it is also possible to automatically raise or lower at least one of the clustering thresholds Th2-1 and Th2-2, so that the computer can recognize whether the number of clusters of the temporary labels will change, and automatically set an appropriate threshold.
[0228] Furthermore, in this embodiment, the system 30 has been described as an example of a transmission device verification system, but this embodiment can also be implemented as an example of the learning data generation system in FIG. 1 or an example of the learning system 2 in FIG.
[0229] <Other embodiments> As described in the first to fourth embodiments, the present disclosure can also take the form of a training data generation method, a training method, or a transmission device matching method.
[0230] Although configuration examples of the components of the systems according to the first to fourth embodiments have been shown, the configuration is not limited to the illustrated examples as long as the functions of the components can be realized. For example, the configuration example in Fig. 5 can be modified so that the matching unit is equipped with a radio wave feature generation unit, and it is sufficient that the system as a whole has the necessary functions.
[0231] Furthermore, the systems according to the first to fourth embodiments or the devices constituting the systems may each have the following hardware configuration: Fig. 15 is a diagram showing an example of the hardware configuration included in a device.
[0232] The device 1000 illustrated in Fig. 15 may be a training data generation system, a training system, or a transmission device verification system according to the first to fourth embodiments, or each device constituting these systems. The device 1000 may be configured as an information processing device (a so-called computer), and may include, for example, a processor 1001, a memory 1002, an input / output interface 1003, and a wireless communication circuit 1004. Note that a wired communication circuit may also be included in addition to the wireless communication circuit 1004. The components such as the processor 1001 are connected by an internal bus or the like, and are configured to be able to communicate with each other.
[0233] The processor 1001 is a programmable device such as a central processing unit (CPU), a micro processing unit (MPU), a digital signal processor (DSP), or a graphics processing unit (GPU). Alternatively, the processor 1001 may be a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or other such device. The processor 1001 can execute various programs including an operating system (OS).
[0234] The memory 1002 is a storage device such as a random access memory (RAM), a read only memory (ROM), a hard disk drive (HDD), a solid state drive (SSD), a memory card, etc. The memory 1002 stores an OS program, application programs, and various data.
[0235] The input / output interface 1003 is an interface for a display device or an input device (not shown). The display device is, for example, a liquid crystal display, an organic electroluminescence display, a printer, etc. The input device is, for example, a device that accepts user operations such as a keyboard, a mouse, a touch panel, etc.
[0236] The wireless communication circuit 1004 is a circuit, module, or the like that performs wireless communication with other devices. For example, the wireless communication circuit 1004 includes an RF (Radio Frequency) circuit, etc. Note that part or all of the device 1000 can be realized by one or more integrated circuits. The device 1000 may also be realized by being divided into one or more parts, and for example, each of the components of the device 1000, such as the processor 1001 and the memory 1002, may also be realized by being divided into one or more parts.
[0237] The functions of device 1000 as a training data generation system, a learning system, or a transmission device verification system, or as each device constituting these systems, can be realized by various processing modules. The processing modules are realized, for example, by processor 1001 executing a program stored in memory 1002. In this case, the program may refer to a training data generation program, a learning program, or a transmission device verification program. Furthermore, the processing modules may be realized by a semiconductor chip.
[0238] The various programs described above include instructions (or software code) that, when loaded into a computer, cause the computer to perform one or more functions described in the embodiments. The programs may be stored in a non-transitory computer-readable medium or a tangible storage medium. By way of example and not limitation, computer-readable media or tangible storage media include random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD) or other memory technology, CD-ROM, digital versatile disc (DVD), Blu-ray disc or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device. The programs may also be transmitted on a transitory computer-readable medium or communication medium. By way of example and not limitation, transitory computer-readable media or communication media include electrical, optical, acoustic, or other forms of propagated signals.
[0239] The present disclosure is not limited to the above-described embodiments, and may be modified as appropriate without departing from the spirit and scope of the present disclosure. In addition, the present disclosure may be implemented by appropriately combining the respective embodiments. [Explanation of symbols]
[0240] 1. Training data generation system 1a Input section 1b Extraction part 1c First Clustering Department 1d Estimated information acquisition part 1e generator 1f, 154 Label setting section 1g, 153 Data Generation Section 2. Learning System 2a, 113 Learning Department 3, 10, 20, 30 Transmitter verification system 3a, 130 Collation section 111 Receiving unit 112 Radio wave feature generation unit 130 Collation Unit 132 Sample feature extraction unit, 133 Threshold determination unit 134 First feature matching unit 135 Template feature memory unit 136 Output section 140, 340 Feature Clustering 146 Sample feature temporary storage unit 147 Second feature matching unit 148, 262, 348 Clustering Department 151 Matrix Generation Unit 152 Visualization section 240 First feature clustering unit 241, 261 Threshold variable control section 250 Feature Acquisition Unit 260 Second feature clustering unit 341 Weight control section 500 Collation Device 900a, 900b transmitting terminal 1000 devices 1001 processor 1002 memory 1003 Input / Output Interface 1004 Wireless communication circuit A1 Target Area
Claims
1. an input unit that inputs n pieces of first information, each of which is either n unknown signals that are signals wirelessly transmitted from N unknown transmission devices or n unknown radio wave features that are radio wave features generated from the n unknown signals, where N and n are positive integers; an extraction unit that inputs the n pieces of first information into a supervised learning model generated from learning data including second information, the second information being either a known signal that is a signal wirelessly transmitted from a transmitting device whose transmission source is known or a known radio wave feature that is a radio wave feature generated from the known signal, and a correct label associated with the second information, and extracts n sample features corresponding to each of the n pieces of first information; a first clustering unit that executes a first clustering process on the n sample features that are the results extracted by the extraction unit; an estimation information acquisition unit that receives m pieces of estimation target information to be estimated by an estimation device that executes a process different from the first clustering process on the n unknown signals or the n unknown radio wave features, where M and m are positive integers, and acquires M pieces of estimation information associated with any one of the n unknown signals or any one of the n unknown radio wave features; a generation unit that generates a relationship matrix indicating a relationship between a first clustering result indicating a result of classification into K groups by the first clustering process, where K is a positive integer, and the M pieces of estimated information; A training data generation system comprising:
2. the estimation information acquisition unit receives m radio wave feature quantities or signals as the m pieces of estimation target information, and receives the M pieces of estimation information associated with the m radio wave feature quantities or signals; The M pieces of estimated information include M temporary labels indicating M respective transmitting devices. The training data generation system according to claim 1 .
3. the estimation information acquisition unit includes a second clustering unit that receives m radio wave feature quantities as the m pieces of estimation target information, and performs a second clustering process on the m radio wave feature quantities to obtain the M pieces of estimation information; The M pieces of estimated information include M temporary labels indicating M respective transmitting devices. The training data generation system according to claim 1 .
4. the first clustering unit performs a weight control process on the n sample features resulting from extraction by the extraction unit and m radio wave features input by the estimation information acquisition unit as the m pieces of estimation target information, and performs the first clustering process on the n sample features and the m radio wave features after the weight control process; the estimated information acquisition unit inputs the m radio wave feature quantities as the m pieces of estimation target information, and obtains the M pieces of estimated information as a result of the third clustering process; The M pieces of estimated information include M temporary labels indicating M respective transmitting devices. The training data generation system according to claim 1 .
5. the estimation information includes temporary label reliability information associated with the temporary label and indicating reliability of the temporary label; the relationship matrix is generated in a state in which the temporary labels are associated with the temporary label reliability information; The training data generation system according to any one of claims 2 to 4.
6. the first clustering unit executes the first clustering process a plurality of times with different clustering thresholds to obtain a plurality of first clustering results, the generation unit generates the relationship matrix for each of the plurality of first clustering results. The training data generation system according to any one of claims 1 to 5.
7. Further comprising a collation unit, the input unit inputs third information, which is either a signal wirelessly transmitted from any of a plurality of transmission devices or a radio wave feature generated from the signal, as information including the first information; the extraction unit inputs the third information into the learning model and extracts sample features for the third information; the matching unit matches the sample feature of the third information with a pre-registered template feature; the first clustering unit performs the first clustering process on sample features determined by the matching unit to be unknown or unregistered, among the n sample features; The training data generation system according to any one of claims 1 to 6.
8. a template feature registration unit that generates and registers template features from the sample features determined by the matching unit to be unknown or unregistered, The training data generation system according to claim 7 .
9. transmitting the n unknown signals or the n unknown radio wave features input by the input unit to the estimation device; the estimation information acquisition unit receives the m pieces of estimation target information from the estimation device, thereby inputting the m pieces of estimation target information; The training data generation system according to any one of claims 1 to 8.
10. further comprising a display unit that displays the relationship matrix, or the relationship matrix and the first clustering result. The training data generation system according to any one of claims 1 to 9.
11. The relationship matrix is added with information indicating at least one of the sample feature amount for each cluster and the inter-cluster distance of each sample feature amount. The training data generation system according to any one of claims 1 to 10.
12. The estimated information is information estimated by the estimation device, which includes at least one of the following: the position of each of the M transmitting devices; the band of the wireless transmission signal that is the signal wirelessly transmitted by each of the M transmitting devices; the frequency or frequency band of the wireless transmission signal; the modulation method used by each of the M transmitting devices; the power value of the wireless transmission signal; the frequency at which the wireless transmission signal is transmitted; the time occupancy rate at which the wireless transmission signal is transmitted; the transmission packet length of the wireless transmission signal; the amount of data transmitted by the wireless transmission signal; the frequency switching pattern when the wireless transmission signal is transmitted by a frequency hopping method; the spectrogram of the wireless transmission signal; and the spectrum of the wireless transmission signal. The training data generation system according to any one of claims 1 to 11.
13. a label setting unit that sets correct labels to at least some of the intersections of the relationship matrix; a data generation unit that generates learning data for updating the learning model based on unknown signals or unknown radio wave features associated with each of the at least some of the intersections and correct labels set for each of the at least some of the intersections; The training data generation system according to any one of claims 1 to 12, further comprising:
14. The training data generation system according to claim 13; a learning unit that performs machine learning based on the learning data generated by the learning data generation system and the learning data including the second information and a correct label linked to the second information, and updates the learning model; A learning system comprising:
15. A processor executes an input process to input n pieces of first information, where N and n are positive integers, which are either n unknown signals that are signals wirelessly transmitted from N unknown transmitting devices or n unknown radio wave features that are radio wave features generated from the n unknown signals; the processor inputs the n pieces of first information into a supervised learning model generated from learning data including second information, the second information being either a known signal that is a signal wirelessly transmitted from a transmitting device whose transmission source is known or a known radio wave feature that is a radio wave feature generated from the known signal, and a correct answer label associated with the second information; and executes an extraction process to extract n sample features corresponding to each of the n pieces of first information; the processor executes a first clustering process on the n sample features that are the extracted results; the processor, where M and m are positive integers, inputs m pieces of estimation target information to be estimated by an estimation device that executes a process different from the first clustering process on the n unknown signals or the n unknown radio wave features, and executes an estimation information acquisition process to obtain M pieces of estimation information associated with any one of the n unknown signals or any one of the n unknown radio wave features; the processor executes a generation process to generate a relationship matrix indicating a relationship between a first clustering result indicating a result classified into K groups by the first clustering process and the M pieces of estimated information, where K is a positive integer. Training data generation method.
16. The processor inputs m radio wave features or signals as the m pieces of estimation target information, and inputs the M pieces of estimation information associated with the m radio wave features or signals; The M pieces of estimated information include M temporary labels indicating M respective transmitting devices. The training data generation method according to claim 15.
17. the estimated information acquisition process includes inputting m radio wave feature quantities as the m pieces of estimation target information, performing a second clustering process on the m radio wave feature quantities, and obtaining the M pieces of estimated information; The M pieces of estimated information include M temporary labels indicating M respective transmitting devices. The training data generation method according to claim 15.
18. The processor performs a weight control process on the n sample features that are the extracted results and the m radio wave features that are input as the m pieces of estimation target information; the first clustering process executes a clustering process on the n sample features and the m radio wave features after the weight control process; the estimated information acquisition process includes inputting the m radio wave feature quantities as the m pieces of estimation target information, and obtaining the M pieces of estimated information as a result of a third clustering process; The M pieces of estimated information include M temporary labels indicating M respective transmitting devices. The training data generation method according to claim 15.
19. the estimation information includes temporary label reliability information associated with the temporary label and indicating reliability of the temporary label; the relationship matrix is generated in a state in which the temporary labels are associated with the temporary label reliability information; The training data generation method according to any one of claims 16 to 18.
20. The first clustering process is performed a plurality of times with different clustering thresholds to obtain a plurality of first clustering results, the generation process generates the relationship matrix for each of the plurality of first clustering results. The training data generation method according to any one of claims 15 to 19.
21. The processor performs a matching process, the input process includes inputting third information, which is either a signal wirelessly transmitted from any of a plurality of transmitting devices or a radio wave feature generated from the signal, as information including the first information; the extraction process includes inputting the third information into the learning model and extracting sample features for the third information; the matching process includes matching the sample feature of the third information with a template feature registered in advance; the first clustering process is performed on sample features determined to be unknown or unregistered in the matching process, among the n sample features; The training data generation method according to any one of claims 15 to 20.
22. The processor generates template features from sample features determined to be unknown or unregistered in the matching process, and registers the template features. The training data generation method according to claim 21.
23. The processor transmits the n unknown signals or the n unknown radio wave features input in the input processing to the estimation device; the estimation information acquisition process receives the m pieces of estimation target information from the estimation device, thereby inputting the m pieces of estimation target information; The training data generation method according to any one of claims 15 to 22.
24. The processor displays the relationship matrix, or the relationship matrix and the first clustering result, on a display device. The training data generation method according to any one of claims 15 to 23.
25. The relationship matrix is added with information indicating at least one of the sample feature amount for each cluster and the inter-cluster distance of each sample feature amount. The training data generation method according to any one of claims 15 to 24.
26. The estimated information is information estimated by the estimation device, which includes at least one of the following: the position of each of the M transmitting devices; the band of the wireless transmission signal that is the signal wirelessly transmitted by each of the M transmitting devices; the frequency or frequency band of the wireless transmission signal; the modulation method used by each of the M transmitting devices; the power value of the wireless transmission signal; the frequency at which the wireless transmission signal is transmitted; the time occupancy rate at which the wireless transmission signal is transmitted; the transmission packet length of the wireless transmission signal; the amount of data transmitted by the wireless transmission signal; the frequency switching pattern when the wireless transmission signal is transmitted by a frequency hopping method; the spectrogram of the wireless transmission signal; and the spectrum of the wireless transmission signal. The training data generation method according to any one of claims 15 to 25.
27. The processor sets correct labels for at least some of the intersections of the relationship matrix; the processor generates learning data for updating the learning model based on the unknown signals or unknown radio wave features associated with each of the at least some of the intersections and the correct labels set for each of the at least some of the intersections. The training data generation method according to any one of claims 15 to 26.
28. The processor performs machine learning based on the learning data generated by the learning data generation method described in claim 27 and the learning data including the second information and a correct label linked to the second information, and updates the learning model. How to learn.
29. A processor inputs first information which is an unknown wireless signal or an unknown radio wave feature generated from the unknown wireless signal; the processor inputs the first information into a supervised learning model generated from learning data including second information, which is a known wireless signal or a known radio wave feature generated from the known wireless signal, and a correct label associated with the second information, and extracts sample features corresponding to each of the first information; the processor performs a first clustering process on the extracted sample features; the processor outputs estimated information associated with the unknown wireless signal or the unknown radio wave feature by a process different from the first clustering process; the processor generates a first clustering result indicating the results of classification into groups by the first clustering process, and a relationship matrix indicating the relationship between the first clustering result and the estimated information. Training data generation method.
30. execute an input process of inputting n pieces of first information, which are either n unknown signals that are signals wirelessly transmitted from N unknown transmission devices or n unknown radio wave features that are radio wave features generated from the n unknown signals, where N and n are positive integers; inputting the n pieces of first information into a supervised learning model generated from learning data including second information, the second information being either a known signal that is a signal wirelessly transmitted from a transmitting device whose transmission source is known or a known radio wave feature that is a radio wave feature generated from the known signal, and a correct answer label associated with the second information; and executing an extraction process to extract n sample features corresponding to each of the n pieces of first information; performing a first clustering process on the n sample feature amounts obtained as a result of the extraction; an estimation device that performs a process different from the first clustering process on the n unknown signals or the n unknown radio wave features, where M and m are positive integers, inputting m pieces of estimation target information to be estimated, and performing an estimation information acquisition process to obtain M pieces of estimation information associated with any one of the n unknown signals or any one of the n unknown radio wave features; executing a generation process for generating a relationship matrix indicating a relationship between the first clustering result indicating the results classified into K groups by the first clustering process and the M pieces of estimated information, where K is a positive integer; A program that causes a computer to execute the learning data generation process.
Citation Information
Patent Citations
Radio wave specification learning device and target identifying device
JP2020173171A
Radio wave specification learning device and target identifying device
JP2020173172A
Learning model generation system and learning model generation method
JP2021179859A
Transmitting device verification device, transmitting device verification system, transmitting device verification method, and computer-readable medium
WO2021070248A1