Calculation method, calculation device

The arithmetic method addresses the challenge of large data sizes by integrating similar data through a dimension reduction process, resulting in a smaller data size that enhances resource efficiency and processing speed.

JP7693375B2Active Publication Date: 2025-06-17HITACHI LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021070139
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-04-19
Publication Date
2025-06-17
Estimated Expiration
2041-04-19

AI Technical Summary

Technical Problem

Existing methods for processing large amounts of data in industrial and natural science fields result in increased data size, leading to resource inefficiencies and slower processing speeds.

Method used

An arithmetic method and device that acquire independent measurement data from multiple sensors, create mask data to indicate the presence or absence of measurement values, and perform a dimension reduction process through embedding to integrate similar data, resulting in a reduced embedding table with smaller data size.

Benefits of technology

The method effectively reduces data size by integrating similar data, thereby reducing resource requirements and improving processing speed and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007693375000001
    Figure 0007693375000001
  • Figure 0007693375000002
    Figure 0007693375000002
  • Figure 0007693375000003
    Figure 0007693375000003
Patent Text Reader

Abstract

To reduce data size.SOLUTION: An operation method includes an acquisition step of acquiring independent measurement data including a combination of a non-uniform measurement time and a measurement value for a plurality of sensors, a preprocessing step of generating mask data which is a numeric value indicating presence / absence of the measure value and measurement data which is a numeric value based on the measurement value for the measurement time and generating a preprocessing table including the measurement data and the mask data of the plurality of sensors for all of the measurement times, and a dimension reduction step of generating a reduced embedded table by performing embedding processing on the measurement data in the preprocessing table, integrating the plurality of measurement data, and integrating the plurality of mask data.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an arithmetic method and an arithmetic device.

Background Art

[0002] Calculations using huge amounts of data are being utilized in the industrial and natural science fields. Although the processing capacity of arithmetic devices has improved with technological advancements, it is desirable for the data size of the processing target to be small in order to save arithmetic resources and improve processing speed. Patent Document 1 discloses a method that can be executed by a device that calculates and displays a risk evaluation value of an event series, which is a partially ordered set showing a part of an event group consisting of a finite number of M types of events in chronological order. The method includes generating an M-dimensional sparse order matrix based on the event series, calculating a dense order matrix by interpolating between elements of the generated sparse order matrix, calculating a mapping matrix that maps the similarity relationship between event series to a two-dimensional or three-dimensional space by an embedding method based on the calculated dense order matrix, calculating corresponding points of each event series on the two-dimensional or three-dimensional space using the calculated mapping matrix, and displaying and outputting the calculated corresponding points in the two-dimensional or three-dimensional space.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the invention described in Patent Document 1, the data size becomes large.

Means for Solving the Problems

[0005] The arithmetic method according to the first aspect of the present invention is An arithmetic method executed by an arithmetic device,An acquisition step of acquiring independent measurement data including combinations of inconsistent measurement times and measurement values for a plurality of sensors, and for each measurement time, creating mask data which is a numerical value indicating the presence or absence of the measurement value and measurement data which is a numerical value based on the measurement value, and creating a preprocessing table including the measurement data and the mask data of the plurality of sensors for all the measurement times; and a dimension reduction step of performing an embedding process on the measurement data in the preprocessing table, integrating the plurality of measurement data and integrating the plurality of mask data based at least on the similarity of the measurement data to create a reduced embedding table. The arithmetic unit according to the second aspect of the present invention includes an acquisition unit that acquires independent measurement data including combinations of inconsistent measurement times and measurement values for a plurality of sensors, and for each measurement time, creates mask data which is a numerical value indicating the presence or absence of the measurement value and measurement data which is a numerical value based on the measurement value, and creates a preprocessing table including the measurement data and the mask data of the plurality of sensors for all the measurement times; and a dimension reduction unit that performs an embedding process on the measurement data in the preprocessing table, integrates the plurality of measurement data and integrates the plurality of mask data based at least on the similarity of the measurement data to create a reduced embedding table.

Advantages of the Invention

[0006] According to the present invention, since similar data is integrated, the data size can be reduced.

Brief Description of the Drawings

[0007]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Embodiments for Carrying Out the Invention

[0008] - First Embodiment - Hereinafter, with reference to FIGS. 1 to 12, a first embodiment of an arithmetic unit will be described.

[0009] (Hardware Configuration) FIG. 1 is a hardware configuration diagram of an arithmetic unit according to the present invention. The arithmetic unit 1 includes a CPU 2 which is a central processing unit, a ROM 3 which is a read-only storage device, a RAM 4 which is a readable and writable storage device, a storage device 5, a user interface 6, and a communication module 7. The CPU 2 expands and executes a program stored in the ROM 3 in the RAM 4 to perform arithmetic operations described later.

[0010] The arithmetic unit 1 may be realized by an FPGA (Field Programmable Gate Array) which is a rewritable logic circuit or an ASIC (Application Specific Integrated Circuit) which is an application-specific integrated circuit, instead of the combination of the CPU 2, the ROM 3, and the RAM 4. Further, the arithmetic unit 1 may be realized by a combination of different configurations, for example, a combination of the CPU 2, the ROM 3, the RAM 4, and the FPGA, instead of the combination of the CPU 2, the ROM 3, and the RAM 4.

[0011] The storage device 5 is a readable and writable non-volatile storage device, for example, a hard disk drive. The user interface 6 is an interface device for exchanging information with a human who operates the arithmetic unit 1, and is, for example, a touch panel in which a liquid crystal display and a pointing device are integrated. The user interface 6 may be realized by a combination of a liquid crystal display and a pointing device. The communication module 7 communicates with other arithmetic units or sensors by wireless communication or wired communication.

[0012] (Functional Configuration) Figure 2 is a functional block diagram of the arithmetic unit 1. The arithmetic unit 1 includes, as its functions, an acquisition unit 11, a preprocessing unit 12, a dialogue unit 13, a parameter calculation unit 14, and a dimensionality reduction unit 15. Also, in either the RAM 4 or the storage device 5 of the arithmetic unit 1, independent measurement data 16, a preprocessing table 17A, a time normalization table 17B, an embedding table 17C, and a reduced embedding table 17D are stored. However, in this embodiment, for the sake of simplicity of description, it is described that all of the above-described information is stored in the RAM 4. Although details will be described later, the preprocessing table 17A, the time normalization table 17B, the embedding table 17C, and the reduced embedding table 17D include the information itself included in the independent measurement data 16 or information obtained by processing the information included in the independent measurement data 16.

[0013] The acquisition unit 11 acquires measurement values from each sensor and stores them in the RAM 4 as independent measurement data 16. Note that the acquisition unit 11 may receive a measurement value and a time from the sensor and store them in the RAM 4, or the acquisition unit 11 may receive a measurement value from the sensor and store the current time at the time of reception and the received measurement value in the RAM 4. Note that the acquisition unit 11 may collect not only the output of the sensor existing outside the arithmetic unit 1 shown in FIG. 1 but also the output of an arithmetic program (not shown) included in the arithmetic unit 1.

[0014] The preprocessing unit 12 reads the independent measurement data 16 and creates the preprocessing table 17A. The dialogue unit 13 prompts the user for input using the user interface 6 and presents the calculation result to the user. The parameter calculation unit 14 reads the preprocessing table 17A and two parameters input by the user, and calculates another parameter necessary for creating the reduced embedding table 17D. Also, the parameter calculation unit 14 creates the time normalization table 17B in the process of processing. The dimensionality reduction unit 15 creates the reduced embedding table 17D based on the user's input and creates the embedding table 17C in the process of processing.

[0015] The independent measurement data 16 stores combinations of measurement values measured by each sensor and the measurement times. The timing at which each sensor measures is not synchronized and is random. For example, even if the measurement value of sensor S1 is measured at time t1, the measurement value of sensor S2 may not be recorded in the independent measurement data 16 at time t1.

[0016] The preprocessing table 17A stores measurement data based on the measurement values stored in the independent measurement data 16 and mask data indicating the presence or absence of the measurement values for all measurement times. The mask data is created to correspond one-to-one with each measurement data. The preprocessing table 17A contains all the information included in the independent measurement data 16. For example, if the arithmetic unit 1 receives 10 measurement values from 10 sensors respectively and the measurement times associated with all the measurement values are different, the preprocessing table 17A will have a size of "20x100". The breakdown is that since the measurement data and the mask data are both "10" respectively, the total is "20", and in time series it is "10" x "10" which is "100".

[0017] The time normalization table 17B is obtained by the parameter calculation unit 14 integrating a plurality of data in the preprocessing table 17A in time series according to time steps. This time step is a value specified by the user or a value calculated by the parameter calculation unit 14 based on other parameters specified by the user. However, if the time step is very fine, the time normalization table 17B will be the same as the preprocessing table 17A.

[0018] The embedded table 17C is created by the dimensionality reduction unit 15 based on the embedded table 17C. The reduced embedded table 17D is created by the dimensionality reduction unit 15 based on the embedded table 17C. The reduced embedded table 17D is an integration of some measurement data and mask data of the embedded table 17C. For example, when the measurement data of sensor S1 and sensor S2 are integrated, the mask data of sensor S1 and the mask data of sensor S2 are also integrated. The reduced embedded table 17D is a so-called deliverable in this embodiment, and the data volume is reduced while retaining the characteristics of the independent measurement data 16. Although not described in this embodiment, the reduced embedded table 17D can be used in various applications, for example, it can be used in Neural ODE (Ricky T. Q. Chen et al., Neural Ordinary Differential Equations, 2018).

[0019] (Preprocessing table) Figure 3 is a diagram showing the relationship between the independent measurement data 16 and the preprocessing table 17A. An example of the independent measurement data 16 is shown at the upper part of Figure 3, and an example of the preprocessing table 17A is shown at the lower part of Figure 3. In Figure 3, the outputs of sensor S1 and sensor S2 are used for explanation, but it is also assumed that the independent measurement data 16 includes the outputs of hundreds, thousands, or tens of thousands of sensors.

[0020] Sensors S1 and S2 operate independently and measure irregularly to output measurement values. However, calling each of them a "sensor" is only for convenience, and each "sensor" does not necessarily have to be equipped with hardware for sensing. The information output may not be the measurement value itself, but may be a calculation result or an anomaly detection signal. When outputting an anomaly detection signal, it may be specified that no output is made when there is no anomaly, and a value of, for example, "1" is output when an anomaly is detected. Here, the outputs of sensor S1 and sensor S2 are any integer between +10 and -10.

[0021] In the example shown in FIG. 3, independent measurement data 16 indicates that sensor S1 measured "5" at time t1 and "2" at time t8, and sensor S2 measured "0" at time t3 and "3" at time t8. In the independent measurement data 16, the outputs of each sensor are expressed as independent tables. Next, an attempt is made to integrate these two tables. Since time t8 is common to the two tables, the times of the integrated table are the three times t1, t2, and t8. Then, when the measured values at each time are described, it becomes like the table in the middle of FIG. 3 (hereinafter referred to as the "intermediate table"). However, in the intermediate table, the columns where no measured values exist, specifically the value of sensor S2 at time t1 and the value of sensor S1 at time t2, are represented as "NaN" (Not a Number; missing value), indicating that no numerical values exist.

[0022] The measured values are directly described in the intermediate table, and the columns where no measured values exist are clearly distinguished from others by storing "NaN", so no misunderstanding occurs. However, from the perspective of data processing, if some numerical values are stored, batch processing becomes possible regardless of the presence or absence of measured values, making it easier to handle. Therefore, in this embodiment, a concept of "measurement data" corresponding to the measured values is introduced, and a concept of "mask data" indicating whether the measurement data is the measured value itself is also introduced. That is, it can be said that the measurement data and the mask data exist in a one-to-one correspondence, and each measurement data has a corresponding mask data.

[0023] For example, "1" in the mask data indicates that the corresponding measurement data is a measured value, and "0" in the mask data indicates that the corresponding measurement data is not a measured value. In the example shown in FIG. 3, the name of the mask data corresponding to sensor S1 is "M1", and the name of the mask data corresponding to sensor S2 is "M2". In the intermediate table, the column of mask M1 at time t1 is set to "1" because the corresponding measured value "5" exists. Also, the column of mask M1 at time t2 is set to "0" because the corresponding measured value is "NaN", that is, it does not exist.

[0024] After clarifying the presence or absence of measurement with this mask vector, as shown in the lower part of FIG. 3, in preprocessing table 17A, "NaN" in the intermediate table is replaced with an arbitrary numerical value, here "0". By this replacement, all data is represented numerically, facilitating data processing. In this case, just looking at the measurement data in preprocessing table 17A, it is impossible to determine whether the measured value was "0" or whether the measurement was not performed at all. However, this determination becomes possible by referring to the corresponding mask data.

[0025] In FIG. 3, for reasons of drawing, the time is described very briefly as "t1" etc., but actually, specific dates and times such as "February 3, 2021 04:05:06" are described. Also, the granularity of time is arbitrary and can be, for example, in units of 1 hour, 1 minute, 1 second, or 1 / 100 second. Also, the granularity of time may differ for each sensor or each measurement. Also, hereinafter, all time-series mask data corresponding to a specific sensor is called "mask information". For example, FIG. 3 includes the mask information of mask M1 and the mask information of mask M2.

[0026] (Processing) FIG. 4 is a flowchart showing the overall optimization process by the arithmetic unit 1. First, in step S30, the preprocessing unit 12 performs preprocessing in the arithmetic unit 1 and combines the outputs of a plurality of sensors into one table. The preprocessing will be described in detail later. In the subsequent step S31, the dialogue unit 13 performs screen display using the user interface 6 and waits for input from the user. The user inputs two of the three: the number of dimensions, the data utilization rate, and the time step.

[0027] FIG. 5 is a diagram showing an example of the screen displayed by the user interface 6 in step S31. FIG. 5 has input interfaces for specifying the aforementioned three, namely, the number of dimensions, the data utilization rate, and the time step. When the user sets two of these three, the parameter calculation unit 14 calculates the remaining one and displays it on the screen as described later. Returning to FIG. 4, the explanation continues.

[0028] In the subsequent step S32, the parameter calculation unit 14 identifies the items to be calculated by the arithmetic unit 1 based on the user input. Specifically, if the parameter calculation unit 14 determines that the number of dimensions is to be calculated since the data utilization rate and the time step are set, it proceeds to step S33. Also, if the parameter calculation unit 14 determines that the data utilization rate is to be calculated since the time step and the number of dimensions are set, it proceeds to step S34. Further, if the parameter calculation unit 14 determines that the time step is to be calculated since the number of dimensions and the data utilization rate are set, it proceeds to step S35.

[0029] In step S33, the parameter calculation unit 14 executes the dimension number calculation process described later and proceeds to step S36. In step S34, the parameter calculation unit 14 executes the data utilization rate calculation process described later and proceeds to step S36. In step S35, the parameter calculation unit 14 executes the time step calculation process described later and proceeds to step S36. In step S36, the dialogue unit 13 further displays on the screen the item calculated in any one of steps S33 to S35 and proceeds to step S37. In step S37, the dimension reduction unit 15 executes the dimension reduction process as described later and ends the process shown in FIG. 4.

[0030] FIG. 6 is a flowchart showing the preprocessing by the preprocessing unit 12 and is a diagram showing the details of step S30 in FIG. 4. First, in step S301, the arithmetic unit 1 reads all the data to be processed and sorts all the recorded times, for example, sorts them in ascending order. In this sorting process, if there are exactly the same times, one of them is left and the others are deleted. For example, in the example shown in FIG. C, the data of sensor S1 records time t1 and time t8, and the data of sensor S2 records time t2 and time t8. Since time t8 is duplicated between them, one report is deleted, and the sorting result is t1, t2, t8.

[0031] In the subsequent step S302, the arithmetic unit 1 creates a table with a width that is twice the number of sensors and a height that is the number of unique time instances. In the example shown in FIG. C, since the number of sensors is "2" and the number of unique time instances is "3" for t1, t2, and t8 described above, a table with a width of "4" (twice "2") and a height of "3" is created.

[0032] In the subsequent step S303, the arithmetic unit 1 transfers the measured values of each sensor to the left half of the table created in step S302, and enters "NaN" in the columns where no measured values exist. However, the vertical direction of the table indicates ascending time, and the horizontal direction indicates the type of sensor in left-justified form. In the subsequent step S304, the arithmetic unit 1 writes mask data to the right half of the table created in step S302. At this time, the arithmetic unit 1 sets the mask data corresponding to the measurement data being "NaN" to "0", and sets the mask data corresponding to the measurement data being other than "NaN", that is, where any measured value exists, to an arbitrary value. This arbitrary value may be either "0" or "1".

[0033] In the subsequent step S305, the arithmetic unit 1 replaces the "NaN" written in step S303 with "0" and ends the process shown in FIG. 6. That is, when the process of this step is completed, the preprocessing table 17A shown at the bottom in the example shown in FIG. C is completed.

[0034] FIG. 7 is a flowchart showing the dimension calculation process by the parameter calculation unit 14 and is a diagram showing the details of step S33 in FIG. 4. First, in step S331, the arithmetic unit 1 reads the time step and data utilization rate values specified by the user. Hereinafter, the time step specified by the user is referred to as "time step DS", and the data utilization rate specified by the user is referred to as "data utilization DU".

[0035] In the subsequent step S332, the arithmetic unit 1 corrects the time in the preprocessing table according to the time step specified by the user. Specifically, each time described in the preprocessing table is divided by the time step DS, and the rounded or truncated result is used as the new time information. For example, when the time step DS is 60 seconds, it may be handled by deleting the information in seconds from the time information. In this case, "March 4, 2020, 5:45:23" is changed to "March 4, 2020, 5:45". If there is a duplication in the time of the preprocessing table due to this process, it will be processed in the next step.

[0036] In the subsequent step S333, the arithmetic unit 1 integrates the information in the preprocessing table that has the same time as follows. That is, the average value of the measured values of the sensors is calculated, and the value of the mask is subjected to an OR operation, which is a bit operation. However, when the corresponding mask value is "0", since no measurement is performed, the measured value of that sensor may be ignored and not included in the population for calculating the average value. By the process of this step, the number of time-series data in the preprocessing table, that is, the number in the vertical direction of the table, is maintained or decreased. A specific example of the process in this step will be described with reference to the drawings.

[0037] FIG. 8 is a diagram showing a specific example of step S333. The upper part of FIG. 8 extracts and shows a part of the preprocessing table, and for convenience of explanation, the record number is described at the left end. In the example shown in FIG. 8, the time information of both sensor S1 and sensor S2 is recorded in units of 1 second. Record "52" is the measured value at 5:45:23 on March 4, 2020, and as can be seen from the mask value, only the measured value of sensor S2 is recorded. Record "53" is the measured value at 5:45:48 on March 4, 2020, and as can be seen from the mask value, the measured values of sensor S1 and sensor S2 are recorded.

[0038] Here, when the time step DS is set to 60 seconds, the information in seconds is omitted, and record "52" and record "53" are treated as the same time. In this case, since the values of M1 for record "52" and record "53" are "0" and "1", the value of S1 corresponding to "1", which is "2", is set as the value of S1 in the integrated record. Since the values of M2 for both record "52" and record "53" are "1", the value of S2 in the integrated record is set to "7", which is the average of "8" and "6". Returning to FIG. 7, the description will continue.

[0039] In the subsequent step S334, the arithmetic unit 1 tabulates the patterns of the mask information, which is the information of each time series included in the mask data, for the preprocessing table 17A after being processed in step S333. In the subsequent step S335, the arithmetic unit 1 sorts the patterns tabulated in step S334 in descending order of the number of occurrences. In the subsequent step S336, the arithmetic unit 1 identifies the minimum number of patterns P that exceeds the specified data utilization rate. Here, the processing of steps S334 to S336 will be described in detail with a specific example.

[0040] FIG. 9 is a diagram for explaining the operations of steps S334 to S336. In this specific example, it is assumed that there are 800 pieces of time information in the preprocessing table after being processed in step S333 and 1000 sensors to be targeted. In this case, when the mask information corresponding to the output of each sensor is treated as vector data, it is 800-dimensional vector data, and there are 1000 pieces of this 800-dimensional vector data, which is the same number as the sensors. In the table shown in FIG. 9, the specific vector values are arranged from top to bottom in descending order of the corresponding number, and a column of "rank" is also provided. Further, in FIG. 9, for the sake of convenience, the cumulative count up to that rank is described on the right end.

[0041] As described above, since the pattern of the mask information is aggregated by the process of step S334, when step S334 is completed, the state where "vector" and "count number" shown in FIG. 9 are associated is obtained. However, in this state, the sorting as shown in FIG. 9 has not been performed yet. When the subsequent step S335 is completed, the sorted state as shown in FIG. 9 is obtained. Here, when the data utilization rate specified by the user is "85%", since the cumulative total up to the rank "123" is "847" cases, it does not reach 850 cases which is 85% of 1000 cases. Therefore, including the next rank "124" and beyond exceeds the specified 85%, and thus the minimum pattern number P to be specified in step S336 is specified as "124". Returning to FIG. 7, the description continues.

[0042] In step S337 which is executed next to step S336, the arithmetic unit 1 calculates the dimension number ceil(log2P) and ends the process shown in FIG. 7. Hereinafter, the dimension number calculated in this step is referred to as "dimension number D". The function ceil is a function that performs the ceiling operation on the real number which is the argument. When the argument is a decimal number, it outputs the result of the ceiling operation, and when the argument is an integer, it outputs that integer. The value calculated in step S337 is the dimension number that this flowchart aims for.

[0043] FIG. 10 is a flowchart showing the data utilization rate calculation process and is a diagram showing the details of step S34 in FIG. 4. In FIG. 10, the same step numbers are assigned to the processes common to FIG. 7, and the description of those steps will be omitted hereinafter. First, in step S341, the arithmetic unit 1 reads the values of the time step and the dimension number specified by the user. Hereinafter, the time step specified by the user is referred to as "time step DS", and the dimension number specified by the user is referred to as "dimension number D".

[0044] The steps S332 to S335 executed after step S341 are processes using the time step DS read in step S341, and since they are as described with reference to FIG. 7, the description is omitted here. In the subsequent step S346, the arithmetic unit 1 calculates the ratio of the total number of the top 2 D-power patterns sorted to the total number of all patterns as the data utilization rate, and ends the process shown in FIG. 10. The value calculated in step S346 is the data utilization rate that this flowchart aims for.

[0045] Here, a case where the specified number of dimensions D is "3" and the number of sorted patterns is as shown in FIG. 9 above will be specifically described as an example. 2 to the power of 3 is "8", and the total of the top "8" patterns in FIG. 9 is "150". And since the total number of all cases in FIG. 9 is "1000", the data utilization rate is calculated as "15%".

[0046] FIG. 11 is a flowchart showing the time step calculation process and is a diagram showing the details of step S35 in FIG. 4. First, in step S351, the arithmetic unit 1 reads the values of the data utilization rate and the number of dimensions specified by the user. Hereinafter, the data utilization rate specified by the user will be referred to as "data utilization rate DU", and the number of dimensions specified by the user will be referred to as "number of dimensions D". In the subsequent step S352, the arithmetic unit 1 reads the upper limit value Tmax and the lower limit value Tmin of the time step that can be specified. These pieces of information may be read from another device, or the values stored in the storage device 5 in advance may be read.

[0047] In the subsequent step S353, the arithmetic unit 1 sets Ttmp to the average value of Tmin and Tmax. In the subsequent step S354, the arithmetic unit 1 executes the process of step S34, that is, the process shown in FIG. 10, for each combination of Tmin, Tmax, Ttmp and the specified number of dimensions D, and calculates Umin, Umax, and Utmp respectively. In other words, in step S354, the process shown in FIG. 10 is executed three times by changing only the parameters of the time step. Next, the arithmetic unit 1 proceeds to step S355.

[0048] In step S355, the arithmetic unit 1 determines whether the calculated value of Utmp is greater than the specified data utilization rate Du. If the arithmetic unit 1 determines that the calculated value of Utmp is greater than the specified data utilization rate Du, it proceeds to step S356. If it determines that the calculated value of Utmp is less than or equal to the specified data utilization rate Du, it proceeds to step S357. In step S356, the arithmetic unit 1 substitutes Utmp into Umin, substitutes the average value of Ttmp and Tmax into Utmp, executes the process of FIG. 10 using the dimension number D specified in step S351 and the Dtmp updated in this step to update Utmp, and returns to step S355.

[0049] In step S357, the arithmetic unit 1 determines whether the calculated value of Utmp is less than the specified data utilization rate Du. If the arithmetic unit 1 determines that the calculated value of Utmp is less than the specified data utilization rate Du, it proceeds to step S358. If it determines that the calculated value of Utmp is not less than the specified data utilization rate Du, that is, equal to the specified data utilization rate Du, it proceeds to step S359. In step S358, the arithmetic unit 1 substitutes Utmp into Umax, substitutes the average value of Ttmp and Tmin into Utmp, executes the process of FIG. 10 using the dimension number D specified in step S351 and the Dtmp updated in this step to update Utmp, and returns to step S355.

[0050] In step S359, the arithmetic unit 1 executes the processes of steps S332 and S333 shown in FIG. 7 using the value of Ttmp calculated immediately before as the time step, and ends the process shown in FIG. 11. The preprocessing table is updated by the process of this step. The value used in this step is the time step that this flowchart aims for. To outline the process shown in FIG. 11, it is to fix the dimension number to the specified value and sequentially change the time step to find the value of the time step that matches the specified data utilization rate.

[0051] FIG. 12 is a flowchart showing the dimensionality reduction calculation process and is a diagram showing the details of step S37 in FIG. 4. As shown in FIG. 4, either step S34 or S35 is executed before step S37 is executed. In any case where a step is executed, based on the pretreatment table 17A, a time normalization table 17B is created using the specified time step or the calculated time step.

[0052] In step S370, the arithmetic unit 1 performs a known embedding process for the purpose of dimensionality reduction on the measurement data in the time normalization table 17B, and creates an embedding table 17C by mapping it to a vector in a low-dimensional space. Examples of this embedding method include the method of Wang et al. (Jizhe Wang, et.al., Billion-scale Commodity Embedding for E-commerce Recommendation in Alibaba, 2018), the method of Nalmpantis et al. (Christoforos Nalmpantis, et.al., Signal2Vec: Time Series Embedding Representation, 2019), and the method of Kazemi et al. (Seyed Mehran Kazemi, et.al., Time2Vec: Learning a Vector Representation of Time, 2019).

[0053] In the subsequent step S371, the arithmetic unit 1 calculates the cosine similarity between the measurement data in the embedding table 17C and proceeds to step S372. After step S372, the embedding table 17C is processed to create a reduced embedding table 17D. That is, until the process shown in FIG. 12 is completed, the reduced embedding table 17D is not completed, but hereinafter, even if it is being created for convenience, it is called the reduced embedding table 17D. The initial state of the reduced embedding table 17D is the embedding table 17C itself created in step S370.

[0054] In step S372, the arithmetic unit 1 determines whether the number of types of measurement data in the current reduced embedding table 17D is greater than the specified or calculated number of dimensions D. If the arithmetic unit 1 determines that the number of types of measurement data in the current reduced embedding table 17D is greater than the specified or calculated number of dimensions D, it proceeds to step S373. If the arithmetic unit 1 determines that the number of types of measurement data in the current reduced embedding table 17D is not greater than the specified or calculated number of dimensions D, it determines that the current reduced embedding table 17D is the final reduced embedding table 17D and ends the process shown in FIG. 12.

[0055] In step S373, the arithmetic unit 1 identifies a plurality of pieces of measurement data with the highest cosine similarity among the unintegrated measurement data among the cosine similarities calculated in step S371 and proceeds to step S374. In step S374, the arithmetic unit 1 selects the oldest unprocessed time among the measurement data identified in step S373. In the subsequent step S375, the arithmetic unit 1 counts the number of pieces of mask data whose values corresponding to the measurement data identified in step S373 and at the time selected in step S374 are "1". If the arithmetic unit 1 determines that the value of the mask data of "1" is zero, it proceeds to step S376. If it determines that the value is "1", it proceeds to step S377. If it determines that the value is "2" or more, it proceeds to step S378.

[0056] In step S376, the arithmetic unit 1 sets the value of the measurement data to be integrated to zero and proceeds to step S379. In step S377, the arithmetic unit 1 sets the value of the measurement data for which the corresponding mask data is "1" to the value of the measurement data to be integrated and proceeds to step S379. In step S378, the arithmetic unit 1 sets the average value or logical sum of the measurement data for which the corresponding mask data is "1" to the value of the measurement data to be integrated and proceeds to step S379. Note that in each of steps S376 to S378, the mask data uses the value of the logical sum of the integration targets.

[0057] In step S379, the arithmetic unit 1 determines whether all the data at all times has been processed for the plurality of measurement data identified in step S373. If it is determined that all the data at all times has been processed, the process returns to step S372. If it is determined that the processing of all the data at all times is not complete, the process returns to step S374. The above is the description of FIG. 12.

[0058] According to the first embodiment described above, the following operational effects can be obtained. (1) The arithmetic method executed by the arithmetic unit 1 includes an acquisition step executed by the acquisition unit 11 that reads independent measurement data including combinations of non-uniform measurement times and measurement values for a plurality of sensors, and for each measurement time, creates mask data that is a numerical value indicating the presence or absence of a measurement value and measurement data that is a numerical value based on the measurement value, and creates a preprocessing table 17A including the measurement data and mask data of a plurality of sensors for all measurement times. A preprocessing step executed by the preprocessing unit 12, and an embedding process is performed on the measurement data in the preprocessing table 17A, and based on at least the similarity of the measurement data, a plurality of measurement data are integrated, and a reduced embedding table 17D in which a plurality of mask data are integrated is created. It includes a dimensionality reduction step executed by the dimensionality reduction unit 15. Therefore, since the reduced embedding table 17D is smaller than the preprocessing table 17A by the dimensionality reduction step executed by the dimensionality reduction unit 15, the data size becomes smaller, and it is expected that the operation using the reduced embedding table 17D will reduce the required resources and improve the processing load.

[0059] (2) The dimensionality reduction step executed by the dimensionality reduction unit 15 determines a plurality of measurement values to be integrated based on the cosine similarity of the measurement data. Therefore, the measurement data can be efficiently integrated.

[0060] (3) The arithmetic method executed by the arithmetic unit 1 includes a time integration step (S332 and S333 in FIGS. 7, 10, and 11) that integrates a plurality of measurement values and a plurality of mask data combined at different measurement times in the preprocessing table 17A according to a time step that is a time width. Therefore, since the measurement data is arranged in time units of a predetermined time step, unified processing becomes possible.

[0061] (4) The calculation method executed by the arithmetic unit 1 includes an input step (step S31 in FIG. 4) that accepts specifications of two out of three: a time step that is the integration time width, a data utilization rate that is the utilization rate of measurement data, and a dimensionality that is the number of data; a parameter calculation step by the parameter calculation unit 14 that calculates one of the three that is not specified; and a time integration step that integrates a plurality of measurement values and a plurality of mask data combined at different measurement times in the preprocessing table 17A according to the time step. Therefore, three mutually related parameters can be determined without contradiction.

[0062] (5) In the preprocessing step, the mask data at the measurement time when there is no measurement value is set to zero for each sensor, and the measurement data corresponding to the non-existent measurement value is set to zero.

[0063] (Modification 1) In the first embodiment, the mask data is set for the purpose of discriminating the presence or absence of a measurement value and whether the measurement value is zero. That is, in the first embodiment, when the value of the sensor vector is "0", mask data is created to discriminate whether the measurement value is "0" or the measurement has not been performed. Therefore, when the measurement data is other than "0", the value of the mask data can be set arbitrarily. However, the design concept of the mask data may be different from that of the first embodiment.

[0064] That is, the mask data may be used for discriminating the presence or absence of a measurement value. For example, "1" in the mask data indicates that the corresponding measurement value exists, that is, the measurement has been performed, and "0" in the mask data indicates that the corresponding measurement value does not exist, that is, the measurement has not been performed, and the mask data may be created in this way. In this case, the value of the sensor data corresponding to "0" in the mask data can be set arbitrarily.

[0065] FIG. 13 is a flowchart showing a second method for creating the preprocessing table 17A. This flowchart is different in that steps S304 and S305 are replaced by steps S304A and S305 compared to FIG. 2. The difference will be described below.

[0066] In step S304A, the arithmetic unit 1 writes the mask data to the right half of the table created in step S302. At this time, the arithmetic unit 1 sets the mask data corresponding to the measurement data being "NaN" to "0", and sets the mask data corresponding to the measurement data being other than "NaN", that is, where there is some measurement value, to "1". In the subsequent step S305A, the arithmetic unit 1 replaces the "NaN" written in step S303 with an arbitrary value and ends the process shown in FIG. 6. This arbitrary value may be a real number within a predetermined numerical range, and can take at least a value between the minimum value and the maximum value of the measurement values.

[0067] In this Modification 1, the following operational effects can be obtained. (6) The preprocessing step sets the mask data at the measurement times when there are no measurement values for each sensor to zero, and sets the mask data at the measurement times when there are measurement values for each sensor to "1". Therefore, in accordance with the handling of the mask data in the application that processes the reduced embedded table 17D, the mask data can be created by the second method instead of the first method.

[0068] (Modification 2) Although the design concept of the preprocessing table is different between the first method described in the first embodiment and the second method described in Modification 1, the preprocessing table may be created so as to correspond to both.

[0069] FIG. 14 is a flowchart showing a method for creating a preprocessing table corresponding to both the first method and the second method. This flowchart can also be said to be obtained by changing step S304 to step S304A with respect to FIG. 6, or by changing step S305A to step S305 with respect to FIG. 13. By creating the preprocessing table using the method shown in FIG. 14, a preprocessing table corresponding to both the first method and the second method can be created.

[0070] (Modification Example 3) The processing of the dimensionality reduction unit 15 may be changed as follows. FIG. 15 is a flowchart showing the processing of the dimensionality reduction unit 15 in Modification Example 3. FIG. 15 is different from FIG. 12 in that steps S370 and S371 are changed to steps S370A and S371A. In step S370A, the dimensionality reduction unit 15 performs embedding processing not only on the measurement data but also on the mask data. In the subsequent step S371A, the dimensionality reduction unit 15 concatenates each measurement data and mask data and calculates the cosine similarity. For example, the measurement data of sensor S1 and the mask data of the corresponding mask M1 are concatenated into one vector, and the cosine similarity with the vector obtained by concatenating the measurement data of sensor S2 and the mask data of mask M2 is calculated. The processing after step S372 is the same as that in the first embodiment, and thus the description thereof is omitted.

[0071] According to this Modification Example 3, the following operational effects can be obtained. (7) The dimensionality reduction step executed by the dimensionality reduction unit 15 determines a plurality of measurement values to be integrated based on the cosine similarity of the measurement data and the mask data. Therefore, a more accurate comparison is possible compared to the case where only the cosine similarity between the measurement data is calculated.

[0072] (Modification Example 4) The arithmetic unit 1 does not necessarily communicate directly with each sensor. For example, the arithmetic unit 1 may communicate with the sensors via one or more other devices to receive measurement values, or may read the past measurement values of each sensor received by other devices. The arithmetic unit 1 may receive the past measurement values via the communication module 7, or may read the past measurement values stored in the storage medium via a medium reading device (not shown). That is, the function of the acquisition unit 11 may be data reading from the storage medium.

[0073] (Modification 5) In the first embodiment, the processing executed by the arithmetic unit 1 may be realized by a plurality of hardware devices. That is, it is only necessary that the system as a whole has the functions provided by the arithmetic unit 1, and the number of devices for realizing the functions is arbitrary.

[0074] - Second Embodiment - With reference to FIGS. 16 to 21, a second embodiment of the arithmetic unit will be described. In the following description, the same components as those in the first embodiment are denoted by the same reference numerals, and the differences will be mainly described. Points not particularly described are the same as those in the first embodiment. In this embodiment, it is mainly different from the first embodiment in that analysis is performed in addition to the processing of the first embodiment.

[0075] FIG. 16 is a functional configuration diagram of the arithmetic unit 1A in the second embodiment. The arithmetic unit 1A further includes an analysis unit 19 and a correspondence table 18 in addition to the configuration of the arithmetic unit 1 in the first embodiment. The analysis unit 19 reads the reduced embedded table 17D and identifies the explanatory variables corresponding to the target variable. The target variable may be specified by the user each time the analysis unit 19 operates, or may be specified in advance. The correspondence table 18 stores the correspondence of the sensor types in the preprocessing table 17A and the reduced embedded table 17D. Note that the hardware configuration of the arithmetic unit 1A is the same as that in the first embodiment, so the description thereof is omitted.

[0076] FIG. 17 is a diagram showing an example of the correspondence table 18. The correspondence table 18 shows the correspondence between the measurement data before and after reduction by the dimensionality reduction process, in other words, the correspondence between the preprocessing table 17A and the reduced embedded table 17D. In the example shown in FIG. 17, it is shown that the sensors 3, 6, and 9 in the preprocessing table 17A and the embedded table 17C are integrated into the third position from the left in the reduced embedded table 17D.

[0077] FIG. 18 is a flowchart showing the overall optimization process by the arithmetic unit 1A in the second embodiment. The differences from the eleventh embodiment are that the step S31 including the screen display is changed to S31A, the step S41 is added, and the step S37 which is the dimensionality reduction process is changed to the step S37A. These will be described in detail below.

[0078] FIG. 19 is a diagram showing an example of the screen display in the second embodiment. In FIG. 19, it is different from the first embodiment in that the context period range can be set. That is, in the present embodiment, it has four parameters: the context period range, the time step, the number of dimensions, and the data utilization rate. And in the present embodiment, it is essential for the user to specify the context period range, and the point that the user specifies two of the three parameters of the time step, the number of dimensions, and the data utilization rate is the same as in the first embodiment.

[0079] FIG. 20 is a flowchart showing the limiting process of limiting the preprocessing table to the context period range. In step S411, the arithmetic unit 1 identifies reference data which is any of the measurement data serving as the reference for the context period. The reference data may be specified by the user each time, or may be specified in advance. In the subsequent step S412, the arithmetic unit 1 calculates the occurrence interval of the reference data identified in step S411 with reference to the preprocessing table 17A or the independent measurement data 16.

[0080] In the subsequent step S413, the arithmetic unit 1 calculates the time that satisfies the ratio specified by the user. For example, if the average generation interval of the reference data is "2 hours" and the ratio corresponding to the context period specified by the user is "1σ", then "1 hour 22 minutes", which is 68% of "2 hours", is calculated. In the subsequent step S414, the arithmetic unit 1 deletes unnecessary rows from the preprocessing table, that is, data other than the period calculated in step S413 from the generation time of the reference data, and ends the process shown in FIG. 20.

[0081] For example, if the reference data is S1 and the time calculated in step S413 is "1 hour 22 minutes", the rows to be deleted are determined as follows. That is, the arithmetic unit 1 targets the data at the time when the measurement data of the sensor S1 in the preprocessing table 17A is not zero or the mask data of the sensor S1 is not zero, and the data from that time to the time "1 hour 22 minutes" before that time, and deletes the other data.

[0082] FIG. 21 is a diagram showing the dimensionality reduction process in the second embodiment. Compared with FIG. 12 in the first embodiment, FIG. 21 is different in that step S421 is added between step S373 and step S374. In step S421, the arithmetic unit 1 writes the plurality of sensor data specified in step S373 in the correspondence table 18 and proceeds to step S374. Specifically, the arithmetic unit 1 adds to the correspondence table 18 the name of the sensor that measured the sensor data specified in step S373 and information indicating the position of the integrated data, for example, which value from the left. The description of the processing after step S374 is the same as that in the first embodiment, so it is omitted.

[0083] FIG. 22 is a flowchart showing the analysis process by the analysis unit 19. First, in step S431, the analysis unit 19 reads the reduced embedded table 17D. In the subsequent step S432, the analysis unit 19 identifies unmeasured measurement data. The unmeasured measurement value is the value in the column described as "NaN" in FIG. 3, or when the preprocessing table 17A is created by the method shown in FIG. 6, it is the value in the column of sensor data where the target variable is "0" and the corresponding mask data is "0". When the target variable is integrated with other sensor data, the data column containing the target variable is identified with reference to the correspondence table 18.

[0084] In the subsequent step S433, the analysis unit 19 estimates the measurement value identified in step S432. Various known techniques can be used for this estimation, for example, interpolation in Word2Vec can be used. In the subsequent step S434, the analysis unit 19 estimates the explanatory variable for the target variable. However, if sensor data is integrated during the creation process of the reduced embedded table 17D, in this step, the identification of the explanatory variable is not achieved, and it only stays at the identification such as the fourth data column from the left.

[0085] In the subsequent step S435, the analysis unit 19 refers to the correspondence table 18 and identifies the names of the data included in the data column identified in step S434. However, if the explanatory variable identified in step S434 is not integrated with other sensor data, the processing of this step can be omitted. In the subsequent step S436, the variables identified in step S436 are displayed on the screen and the process shown in FIG. 22 is terminated.

[0086] Also, in this embodiment, the dimensionality reduction process shown in FIG. 12 is changed as follows. That is, in step S373, the names of the identified plurality of sensor data and the storage positions of the integrated data are stored in the correspondence table 18.

[0087] According to the second embodiment described above, the following operational effects can be obtained. (8) It includes an analysis step of reading the reduced embedded table 17D and identifying sensors having other measurement values that are highly correlated with the measurement values of a predetermined sensor. The analysis step identifies a predetermined sensor and sensors having other measurement values with reference to a correspondence table 18 showing the correspondence between the preprocessing table 17A and the reduced embedded table 17D. Therefore, since the arithmetic unit 1A reads the reduced embedded table 17D and performs analysis, analysis with fewer resources becomes possible.

[0088] (Modification Example 1 of the Second Embodiment) In the second embodiment, the processing of the dimensionality reduction unit 15 may be changed as follows. FIG. 23 is a diagram showing the processing of the dimensionality reduction unit 15 in Modification Example 1 of the second embodiment. In step S441, the dimensionality reduction unit 15 calculates the co-occurrence probability between the measurement data in the embedded table 17C. In the subsequent step S442, the dimensionality reduction unit 15 determines, as integration targets, the measurement data pairs whose co-occurrence probability calculated in step S441 is equal to or greater than a predetermined value, and the corresponding mask data pairs.

[0089] In the subsequent step S443, the dimensionality reduction unit 15 calculates the co-occurrence probability between the mask data in the embedded table 17C. In the subsequent step S444, the dimensionality reduction unit 15 determines, as integration targets, the mask data pairs whose co-occurrence probability calculated in step S443 is equal to or greater than a predetermined value, and the corresponding measurement data pairs. In the subsequent step S445, the dimensionality reduction unit 15 sets, as the integrated value, the logical sum of the average value of the measurement data and the mask data for the measurement data and mask data to be integrated.

[0090] (Modification Example 1 of the Second Embodiment) The interpolation process by the analysis unit 19, that is, the estimation of the measurement value in step S433 of FIG. 22, may utilize the "interpolate" function or the "fillna" function of the library "pandas" in python.

[0091] In each of the above-described embodiments and modifications, the configuration of the functional blocks is merely an example. Some of the functional configurations shown as separate functional blocks may be integrated, or the configuration represented by one functional block diagram may be divided into two or more functions. Further, a part of the functions of each functional block may be provided by other functional blocks.

[0092] In each of the above-described embodiments and modifications, although the program of the arithmetic unit 1 is assumed to be stored in a ROM (not shown), the program may be stored in the storage device 5. Further, the arithmetic unit 1 may include an input / output interface (not shown), and when necessary, the program may be read from another device via a medium that can be used by the input / output interface and the arithmetic unit 1. Here, the medium refers to, for example, a storage medium detachable from the input / output interface, or a communication medium, that is, a wired, wireless, optical, etc. network, or a carrier wave or digital signal propagating through the network. Further, part or all of the functions realized by the program may be realized by a hardware circuit or an FPGA.

[0093] The above-described embodiments and modifications may be combined with each other. Although various embodiments and modifications have been described above, the present invention is not limited to these contents. Other aspects conceivable within the scope of the technical idea of the present invention are also included in the scope of the present invention.

Explanation of Reference Numerals

[0094] 1, 1A... Arithmetic unit 11... Acquisition unit 12... Preprocessing unit 13... Interaction unit 14... Parameter calculation unit 15... Dimension reduction unit 16... Independent measurement data 17A... Preprocessing table 17B... Time normalization table 17C... Embedding table 17D... Reduced embedding table 18... Correspondence table 19... Analysis unit

Claims

1. A calculation method executed by a calculation device, an acquisition step of acquiring independent measurement data including combinations of inconsistent measurement times and measurement values for a plurality of sensors; a preprocessing step of creating mask data, which is a numerical value indicating the presence or absence of the measurement value, and measurement data, which is a numerical value based on the measurement value, for each of the measurement times, and creating a preprocessing table including the measurement data and the mask data of the plurality of sensors for all the measurement times; a dimension reduction step of performing embedding processing on the measurement data in the preprocessing table, integrating the plurality of measurement data based on at least the similarity of the measurement data, and creating a reduced embedding table in which the plurality of mask data are integrated.

2. In the calculation method according to Claim 1, the dimension reduction step determines the plurality of measurement data to be integrated based on the cosine similarity of the measurement data.

3. In the calculation method according to Claim 1, the dimension reduction step determines the plurality of measurement data to be integrated based on the cosine similarity of the measurement data and the mask data.

4. In the calculation method according to Claim 1, further including a time integration step of integrating the plurality of measurement values and the plurality of mask data combined at different measurement times in the preprocessing table according to a time step which is a time width.

5. In the calculation method according to Claim 1, an input step of receiving specifications of two out of three: a time step which is an integration time width, a data utilization rate which is a utilization rate of measurement data, and a number of dimensions which is a number of data; a parameter calculation step of calculating one of the three that is not specified; A time integration step of integrating a plurality of the measurement values and a plurality of the mask data combined at different measurement times in the preprocessing table according to the time step, and an arithmetic method further including the same.

6. In the arithmetic method according to claim 1, In the preprocessing step, the mask data at the measurement time when the measurement value does not exist for each sensor is set to zero, and the measurement data corresponding to the non-existing measurement value is set to zero. An arithmetic method.

7. In the arithmetic method according to claim 1, In the preprocessing step, the mask data at the measurement time when the measurement value does not exist for each sensor is set to zero, and the mask data at the measurement time when the measurement value exists for each sensor is set to "1". An arithmetic method.

8. In the arithmetic method according to claim 1, The method further includes an analysis step of reading the reduced embedded table and identifying a sensor having other measurement values strongly correlated with the measurement value of a predetermined sensor, The analysis step identifies the predetermined sensor and the sensor having the other measurement values with reference to a correspondence table showing the correspondence between the preprocessing table and the reduced embedded table. An arithmetic method.

9. An acquisition unit that acquires independent measurement data including combinations of inconsistent measurement times and measurement values for a plurality of sensors, For each measurement time, mask data that is a numerical value indicating the presence or absence of the measurement value and measurement data that is a numerical value based on the measurement value are created, and the measurement data and the mask data of the plurality of sensors for each measurement time are created. A preprocessing unit that creates a preprocessing table including the same, An arithmetic device including a dimensionality reduction unit that performs an embedding process on the measurement data in the preprocessing table, integrates a plurality of the measurement data, and creates a reduced embedded table in which a plurality of the mask data are integrated based on at least the similarity of the measurement data.

Citation Information

Patent Citations

  • Time-series data analysis system, method and program

    JP2009187293A

  • Data transmitter and remote monitoring system

    JP2016039610A

  • Data acquisition system, input device, data acquisition device, and data synthesis device

    JP2019179005A

  • Method to combine partially aggregated sensor data in a distributed sensor system

    US20200049677A1

  • Method, device, and computer program for visualizing risk assessment valuation of sequence of events

    WO2013084779A1