Time-series data processing equipment and time-series data processing methods

The time-series data processing device addresses spurious correlations by preprocessing, grouping, and stratifying data to identify causal relationships, improving prediction accuracy and user trust through visualization.

DE112023006807T5Pending Publication Date: 2026-06-03MITSUBISHI ELECTRIC CORP

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
MITSUBISHI ELECTRIC CORP
Filing Date
2023-10-20
Publication Date
2026-06-03

AI Technical Summary

Technical Problem

Existing time series data processing technologies automatically select explanatory variables based on spurious correlations, leading to unreliable predictions, as they fail to establish causal relationships between variables.

Method used

A time-series data processing device that preprocesses, groups, and stratifies time-series data to identify causal relationships, using waveform and semantic analysis, and generates visualization diagrams to facilitate user selection of valid explanatory variables.

Benefits of technology

Prevents the selection of purely spurious correlated variables, enhancing the persuasiveness of predictions by incorporating domain knowledge and reducing user distrust through clear visualization of causal relationships.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A time series data processing unit comprises: a data input unit (110) for obtaining a first time series data set containing a plurality of time series data parts to be treated as candidates for explanatory variables; a preprocessing unit (120) for converting the obtained first time series data set into a format capable of calculating relationships between the plurality of time series data parts contained in the first time series data set and for generating a second time series data set containing the plurality of time series data parts after the conversion; a relationship calculation unit (140) for calculating waveform and semantic relationships between the plurality of time series data parts contained in the second time series data set;a stratification unit (150) for determining temporal relationships between the multitude of parts of time series data contained in the second time series data set, for stratifying the multitude of parts of time series data contained in the second time series data set according to the determined temporal relationships, and for outputting a result of the stratification; and a visualization unit (160) for generating a visualization diagram that visualizes the multitude of parts of time series data contained in the second time series data set based on the calculated waveform and semantic relationships and the output result of the stratification.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL AREA

[0001] The present disclosure relates to a time series data processing technology. STATE OF THE ART

[0002] In time series data processing using a predictive model, the selection of explanatory variables is extremely important and significantly influences the accuracy of the prediction. However, manually selecting the optimal explanatory variables from a large number of candidates is difficult. Accordingly, technologies are proposed to automatically perform the selection / elimination of explanatory variables during prediction according to a pre-specified algorithm (e.g., patent literature 1).In the technology according to patent literature 1, the selection / elimination of the explanatory variables is carried out automatically by comparing (the absolute values) of the regression coefficients between the response variable and the explanatory variables with a threshold value, assuming that explanatory variables with larger regression coefficients are more suitable (claim 4 and paragraph 0093 in patent literature 1). REFERENCE LIST PATENT LITERATURE

[0003] Patent literature 1: WO 2013 / 187295 SUMMARY OF THE INVENTIONAL PROBLEM

[0004] The technology disclosed in patent literature 1, in which the selection / elimination of explanatory variables is performed automatically, has a problem: in some cases, only explanatory variables are selected that show a spurious correlation with the response variable. That is, there is a problem that explanatory variables that cannot be said to be causally related to the response variable are classified as having a causal relationship based on a specific factor, and in some cases, only such explanatory variables are selected.

[0005] The present disclosure serves to solve such a problem and aims to provide a time series data processing technology that is able to prevent the selection of purely explanatory variables that show a spurious correlation. TECHNICAL SOLUTION

[0006] One aspect of a time-series data processing device according to an embodiment of the present disclosure comprises: a data input unit for obtaining a first time-series data set containing a plurality of time-series data parts to be treated as candidates for explanatory variables; a preprocessing unit for converting the obtained first time-series data set into a format capable of calculating relationships between the plurality of time-series data parts contained in the first time-series data set and for generating a second time-series data set containing the plurality of time-series data parts after the conversion; a relationship calculation unit for calculating waveform and semantic relationships between the plurality of time-series data parts contained in the second time-series data set;a stratification unit for determining temporal relationships between the multitude of parts of time series data contained in the second time series dataset, for stratifying the multitude of parts of time series data contained in the second time series dataset according to the determined temporal relationships, and for outputting a stratification result; and a visualization unit for generating a visualization diagram that visualizes the multitude of parts of time series data contained in the second time series dataset based on the calculated waveform and semantic relationships and the output result of the stratification. ADVANTAGEOUS EFFECTS OF THE INVENTION

[0007] The time series data processing device according to the embodiment of the present disclosure presents a plurality of candidates for explanatory variables, thereby preventing the selection of only explanatory variables that show a spurious correlation. BRIEF DESCRIPTION OF THE DRAWINGS Fig. Figure 1 is a drawing showing a configuration example for a time series data processing facility and a time series data processing system. Fig. Figure 2 is a schematic drawing showing a process carried out by a grouping unit. Fig. Figure 3 is a drawing that shows an overview of a relationship calculation performed by a relationship calculation unit. Fig. Figure 4 is a drawing that shows an overview of a stratification carried out by a stratification unit. Fig. 5A and Fig. 5B are drawings that show one variation of the layering. Specifically, they are Fig. 5A and Fig. 5B drawings showing an example of a month-shifted correlation. Fig. 5A is a drawing showing the original waveforms. Fig. 5B is a drawing showing a case where a waveform is shifted by one month. Fig. Figure 6 is a drawing that shows one variation of the layering. Specifically, it is Fig. 6 a drawing showing a specific example of a case in which the layering is carried out according to preset temporal relationships. Fig. Figure 7A is a drawing showing a configuration example for the hardware of the time series data processing facility. Fig. Figure 7B is a drawing showing a configuration example for the hardware of the time series data processing facility. Fig. Figure 8 is a flowchart of a time series data processing procedure. Fig. Figure 9 is a drawing that shows an example of a visualization diagram. Fig. Figure 10 is a drawing that shows an overview of a process carried out by a recalculation unit. DESCRIPTION OF THE EXECUTION FORMS

[0008] Various embodiments according to the present disclosure are described in detail below with reference to the accompanying drawings. It should be noted that components designated in the drawings with identical or similar reference numerals are components having identical or similar configurations or functions, and that overlapping descriptions of such components are omitted. Unless otherwise stated, the term "or" in the present disclosure is used to mean an inclusive logical disjunction.

[0009] Furthermore, “causal relationship” as used in this disclosure means a relationship selected by a user from temporal relationships determined by statistical analysis between two pieces of time-series data. The user selects a temporally ordered relationship that convinces the user. First embodiment. <konfiguration>

[0010] A time-series data processing device and a time-series data processing system according to a first embodiment of the present disclosure are described with reference to Fig. 1 explained. The in Fig. The depicted time-series data processing system comprises a time-series data processing unit 100, a storage unit 200, and a storage unit 300. Storage unit 200 is a unit for storing time-series data to be treated as candidates for explanatory variables. Storage unit 300 is a unit for storing additional data. (Time series data processing facility)

[0011] The time series data processing unit 100 comprises: a data input unit 110 for obtaining a first time series data set, containing a multitude of time series data parts to be treated as candidates for explanatory variables, from the storage unit 200; a preprocessing unit 120 for converting the obtained first time series data set into a format capable of calculating relationships between the multitude of time series data parts contained in the first time series data set and for generating a second time series data set containing the multitude of time series data parts after the conversion; a grouping unit 130 for classifying the multitude of time series data parts contained in the second time series data set into groups according to the nature of the multitude of time series data parts and generating a representative value for each group;a relationship calculation unit 140 for calculating waveform and semantic relationships between the multitude of parts of time series data contained in the second time series dataset; a stratification unit 150 for determining temporal relationships between the multitude of parts of time series data contained in the second time series dataset, for stratifying the multitude of parts of time series data contained in the second time series dataset according to the determined temporal relationships, and for outputting a result of the stratification; a visualization unit 160 for generating a visualization diagram that visualizes the multitude of parts of time series data contained in the second time series dataset based on the calculated waveform and semantic relationships and the output result of the stratification;and a recalculation unit 170 to accept additional time series data and perform a recalculation to add the accepted additional time series data to the visualization chart. (Data input unit)

[0012] The data input unit 110 acquires a time series dataset (first time series dataset) D1, which contains a variety of time series data segments to be treated as candidates for leading indicators. From these candidates, a leading indicator, or preceding indicator, is selected to be used for forecasting a future indicator or time series data of the leading indicator. An example of a future indicator used for forecasting is the delivery volume of air conditioners.

[0013] The purpose of the time-series data processing device 100 and the time-series data processing system of the present disclosure will be explained in more detail below. To this end, a case is considered in which the quantity of apples consumed is used as an explanatory variable when predicting the delivery volume of air conditioners. Even if this prediction succeeds, it is difficult to establish a causal relationship between the "quantity of apple consumption" and the "delivery volume of air conditioners" based on domain knowledge. Accordingly, there is a possibility that the result of the prediction in this case will not be convincing to a user of the time-series data processing device 100 and may arouse suspicion in the user.

[0014] In contrast, consider a case where the number of construction starts for new buildings is used to explain the forecast of air conditioning unit delivery volume. In this case, domain knowledge establishes a causal relationship: "The number of buildings increases. → The demand for newly installed air conditioning units in the buildings increases. → The delivery volume for air conditioning units increases." Accordingly, the forecast result in this case should be convincing for the user.

[0015] Selecting explanatory variables that convince the user is difficult to achieve through statistical analysis alone, making manual evaluation essential. On the other hand, there are hundreds of thousands of types of economic indicator data used to forecast product demand, and manually selecting explanatory variables from such a vast amount of data that are both domain-specific and statistically valid is extremely challenging.

[0016] One goal of the time series data processing device 100 and the time series data processing system of the present disclosure is to make a procedure for selecting / eliminating explanatory variables convincing to the user and to reduce feelings of distrust in the procedure regarding prediction by improving the procedure. The user's satisfaction is the significance of the causal relationship between a multitude of indicators.

[0017] The explanation returns to data input unit 110. It is assumed that the time series dataset D1 comprises a plurality of time series datasets, and each time series dataset is given a name indicating the type of data it contains. Furthermore, it is assumed that the time series of all the plurality of time series datasets contained in time series dataset D1 share the same unit. For example, it is assumed that all of the plurality of time series datasets in time series dataset D1 are monthly data, yearly data, or similar, sharing the same unit. Data input unit 110 delivers the acquired time series dataset D1 to preprocessing unit 120. (Preprocessing unit)

[0018] The preprocessing unit 120 is a functional unit for performing data preprocessing of the time series data set D1, acquired at the data input unit 110, to generate a time series data set D2 and to output the generated time series data set D2 to the grouping unit 130. That is, the preprocessing unit 120 converts the acquired first time series data set D1 into a format capable of calculating relationships between the multitude of time series data parts contained in the first time series data set D1 and the second time series data set, which contains the multitude of time series data parts after the conversion. The data preprocessing includes, in particular, a process related to missing values ​​and a process related to data standardization and similar tasks. (Group unit)

[0019] The grouping unit 130 is a functional unit for classifying the time series data set D2 into a multitude of groups according to the type of each part of data, for generating a representative value for each group, for transmitting information about it (representative value) to the group, and for outputting a group D3 to the relationship calculation unit 140 after the representative value has been transmitted.

[0020] More specifically, the grouping unit 130 first performs clustering of the time series data set D2 obtained from the preprocessing unit 120, based on the degrees of similarity between waveforms, the degrees of similarity between names, and the like. The clusters obtained as a result of the clustering are treated as groups, and then data processing is carried out that treats the groups as units. Subsequently, a representative value is generated that reflects the characteristics of each group in order to treat each group as time series data.

[0021] In the first embodiment, the classification of the time series dataset D2 is performed as follows. First, a cross-correlation is calculated according to the following formula (1) in a state where the time series of each part of the data are aligned. Then, the parts of the data whose correlation coefficient is equal to or greater than a certain value (predefined; D3-N1) are classified into the same group. It is assumed that each group has information about a list of names (D3-1) of the parts of the data contained in the group. r=Σi(xi−x¯)(yi−y¯)∑i(xi−x¯)2∑i(yi−y¯)2

[0022] It should be noted that the index i represents a date / time t, and e.g. t = 1 January, 1 February, 1 March, ..., 1 December.

[0023] Furthermore, the generation of representative values ​​is carried out as follows. The period for the representative value of each group is the period from the oldest to the most recent point in time of the data contained in the group and is treated as period T1. Additionally, the average value of the data at each point in time within period T1, which is contained in the group, is treated as the representative value of the group at that point in time. In this way, a representative value can be generated that reflects the characteristics of the data contained in each group. It is assumed that the groups D3 have information (D3-2) about the representative values.

[0024] Fig. Figure 2 is a schematic drawing showing a process carried out by grouping unit 130. As in Fig. As shown in Figure 2, the time series dataset D2 comprises a multitude of time series datasets, such as a time series dataset "production volume of iron resources", a time series dataset "production volume of electricity", a time series dataset "number of construction starts" and a time series dataset "food export volume". The grouping unit 130 performs a grouping of the multitude of time series data contained in the time series dataset D2 according to the type of data. Fig. Figure 2 shows a state in which the time series datasets "Production volume of iron resources" and "Production volume of electricity" are classified into group 1, and the time series datasets "Number of construction starts" and "Food export volume" are classified into group 2. In this way, the time series datasets are classified into groups, and each group is designated as group D3 after classification. (Relationship calculation unit)

[0025] The relationship calculation unit 140 is a functional unit for calculating the relationships between the groups (waveform and semantic relationships) D4 between the groups D3 generated by the grouping unit 130, and for outputting information about the relationships D4 and the groups D3 to the stratification unit 150. Specifically, the relationships between a multitude of the groups D3 are calculated based on representative values ​​of the respective groups. The relationships between the groups provide information about the degree of similarity of the waveforms between the groups (D4-1). Waveform and semantic relationships refer to waveform relationships and semantic relationships, respectively. Waveform relationships are indicators that represent the degree to which the waveforms of data are similar to one another. Semantic relationships are, for example, the depth of association between data, which can be defined by domain knowledge or similar factors.Even though the waveforms of the air conditioning demand volume data and the average temperature data are not similar, a concrete example of semantic relationships from domain knowledge is that the air conditioning demand volume data and the average temperature data are obviously related in some cases.

[0026] In the first embodiment, the calculation of the relationships between the groups is performed as follows. Since a representative value of each group can be handled as time series data, the relationships are defined by the magnitude of the coefficients of a cross-correlation between the data, similar to the process performed by the grouping unit 130. It should be noted that at the time of the cross-correlation calculation in the relationship calculation unit, the calculation is performed using data obtained by shifting the representative values ​​of the groups by up to ±N (a predefined integer; e.g., 12) in one-month increments, and that the largest correlation coefficient among the calculated correlation coefficients is treated as the degree of waveform similarity (D4-1) between the groups.Furthermore, the shift width at this time is treated as a temporal precedence relationship (D4-2) between the groups.

[0027] Fig. Figure 3 is a diagram that shows an overview of the relationship calculation performed by the relationship calculation unit 140. As in Fig. As shown in Figure 3, the relationships D4 between the multitude of groups D3 are calculated, and groups with strong relationships are connected by lines according to the result of the calculation. Fig. Group 1 has strong relationships with Group 2 and Group 3, and Group 1 is connected to Group 2 and Group 3 by lines. In addition to Group 1, Group 2 also has strong relationships with Group 4 and Group 5, and Group 2 is also connected to Group 4 and Group 5 by lines. In addition to Group 2, Group 4 also has a strong relationship with Group 6, and Group 4 is also connected to Group 6 by a line. (Layer unit)

[0028] The stratification unit 150 is a functional unit for determining temporal relationships between the multitude of parts of time series data contained in the time series data set D2, for stratifying the multitude of parts of time series data contained in the time series data set D2 according to the determined temporal relationships, and for outputting the result of the stratification.

[0029] The stratification unit 150 can process the multitude of groups D3 generated by the grouping unit 130 as processing targets. In this case, the stratification unit 150 determines temporal relationships between the multitude of groups D3, stratifies the multitude of groups D3 according to the determined temporal relationships, and outputs the result of the stratification.

[0030] A more detailed explanation is given for the case where the stratification unit 150 treats the multitude of groups D3 as processing targets. To clarify causal relationships between the multitude of groups D3, the stratification unit 150 specifies temporal relationships between the groups D3 based on the sequence of various activities or statistical analyses, such as a month-shifted correlation, and transmits information (D3-3) about the temporal relationships to the groups D3.

[0031] Fig. Figure 4 is a drawing that shows an overview of an example of stratification carried out by a stratification unit 150. As in Fig. As shown in Figure 4, the stratification unit 150 rearranges the groups D3 linked by the relationship calculation unit 140 layer by layer according to the temporal relationships. The example of stratification in Fig. Figure 4 shows that group 2 is ahead of group 4 and group 5 in time, and group 1 is ahead of group 2 and group 3 in time.

[0032] In the first embodiment, the generation of layers (temporal relationships) of the groups is carried out, for example, as follows. That is, the time series data are shifted in time to maximize the correlation coefficient, and a temporally ordered relationship between the data is specified based on the temporal precedence relationship obtained from the shift width. This point is discussed with reference to Fig. 5A and Fig. 5B explained. Fig. 5A and Fig. Figure 5B shows an example of a month-shifted correlation. Fig. 5A is a drawing showing the original waveforms and Fig. 5B is a drawing showing a case where a waveform is shifted by one month. As in Fig. 5A and Fig. As shown in Figure 5B, a relationship is assumed to exist between certain time series data X (represented by a bold line) and other specific time series data Y (represented by a thin line), such that the cross-correlation coefficient between time series data X and Y reaches its maximum when time series data Y is shifted forward by one month. This means that time series data Y precedes time series data X by one month. Accordingly, a temporal precedence relationship arises, in which time series data Y precedes time series data X by one month. The stratification unit 150 also captures such temporal precedence relationships between other data, determines temporal relationships from the captured temporal precedence relationships, stratifies the multitude of groups D3 according to the determined temporal relationships, and outputs the result of the stratification.Similarly, in a case where the time series data set D2 is treated as the processing target, the stratification unit 150 determines temporal relationships between the multitude of parts of time series data contained in the time series data set D2, stratifies the multitude of parts of time series data contained in the time series data set D2 according to the determined temporal relationships, and outputs the result of the stratification.

[0033] Another implementation of stratification, for example, a temporally ordered relationship defined using words such as demand → production → sales, is preset according to the domain knowledge. The domain knowledge is defined and implemented through user input. The preset domain knowledge comprises a multitude of words and the definition of temporal relationships between these words. The stratification unit 150 retrieves the configured domain knowledge. The user input regarding the preset can be retrieved directly by the stratification unit 150 or indirectly via the data input unit 110. The stratification unit 150 stratifies data whose names contain the words, based on the retrieved temporal relationships. Fig. Figure 6 shows a specific example of a case where stratification is performed according to predefined temporal relationships. The stratification unit 150 obtains the predefined temporal relationships, demand → production → sales, and stratifies the groups "car demand volume", "car production volume", "car sales volume", "metal demand volume", "semiconductor demand volume", "PC production volume", and "household appliance sales volume" according to the temporal relationships. Since the groups "car demand volume", "metal demand volume", and "semiconductor demand volume" contain the term "demand", the stratification unit 150 classifies the groups "car demand volume", "metal demand volume", and "semiconductor demand volume" into the first layer.Since the groups "Passenger Car Production Volume" and "PC Production Volume" contain the term "Production," the stratification unit 150 classifies the groups "Passenger Car Production Volume" and "PC Production Volume" as belonging to the second stratification. Since the groups "Passenger Car Sales Volume" and "Household Appliance Sales Volume" contain the term "Sales," the stratification unit 150 classifies the groups "Passenger Car Sales Volume" and "Household Appliance Sales Volume" as belonging to the third stratification.

[0034] Furthermore, the stratification unit 150 can perform stratification using a Granger causality test approach. (Visualization unit)

[0035] The visualization unit 160 is a functional unit for generating a visualization diagram that visualizes the relationships between the groups using information about the groups D3 (including the name list D3-1 of the data contained in each group and the information D3-3 about the step sequence of each group) and the (semantic) relationships between the groups D4 (including the degrees of similarity of the waveforms between the groups D4-1 and the temporal precedence relationships D4-2 between the respective groups) in such a way that a person can easily grasp the relationships between the groups, and for outputting the generated visualization diagram.

[0036] In the first embodiment, the visualization of the relationships between the groups is carried out as follows. The visualization of the relationships between groups is performed in a graphical format in which the respective groups (D3) are represented as vertices and the relationships between the groups (D4) are represented as edges.

[0037] First, a layered structure similar to the layers (D3-3) defined by the layering unit 150 is prepared.

[0038] Subsequently, each group is assigned to a layer based on the data in the name list (D3-1) contained within the group. Each group is visualized as a vertex.

[0039] The relationships between the groups are then visualized based on the information (D4-1) about the degree of similarity of the waveforms between the groups. In the visualization, the groups, represented by the vertices, are connected by edges that vary in appearance, such as line thickness, line color, or line type (e.g., solid or dashed line), depending on the strength / weakness of the relationships.

[0040] For example, the list, the name list (D3-1) of the data contained in each group, is made easily visible by displaying the list in a list format when an operation such as clicking on a group at a corner point is performed, and so on. Such operation and display are performed via input and output devices not shown. The visualization unit 160 obtains an operating instruction entered via the input device and performs a display control for the display on the output device. (Recalculation unit)

[0041] Once the series of processes carried out by the data input unit 110 to the visualization unit 160 has been executed, the recalculation unit 170 accepts additional time series data D1-2 and performs a process to output a visualization diagram in which the accepted additional time series data D1-2 has been added to the time series data set D1.

[0042] In the first embodiment, the recalculation at the time the additional time series data D1-2 were added is performed as follows. It should be noted that the additional time series data D1-2 were stored on the storage device 300 and the recalculation unit 170 retrieves the additional time series data D1-2 from the storage device 300.

[0043] It is assumed here that the additional time series data D1-2 is a single piece of time series data. If multiple pieces of time series data are added, the processes outlined below are performed on each piece of additional data, appending content appropriate to the additional data to a generated visualization chart. - The recalculation unit 170 delivers the additional time series data D1-2 to the preprocessing unit 120. - The preprocessing unit 120 performs preprocessing of the additional time series data D1-2. - The grouping unit 130 calculates the cross-correlation coefficients between the additional time series data D1-2, after preprocessing, and the multitude of already generated groups D3. The calculation of the cross-correlation coefficients can be performed using the same procedure as described above. - In cases where data has a correlation coefficient equal to or greater than the predefined constant (D3-N1) set in grouping unit 130, grouping unit 130 selects the groups with the highest correlation coefficient and adds the additional time series data D1-2 to the selected group. That is, grouping unit 130 adds the name of the additional data to the name list D3-1 of the data contained in the relevant group. Once this process is complete, the recalculation is finished. - In a case where no data with a correlation coefficient equal to or greater than the predefined constant (D3-N1) set in grouping unit 130 is available, grouping unit 130 treats the additional time series data D1-2 as an independent group. The relationship calculation unit 140 then calculates the relationships between this group of additional time series data D1-2 (which is a single group) and other groups D3. Based on the result, visualization unit 160 updates the visualization chart and completes the recalculation.

[0044] It should be noted that the recalculation process can also be performed using a different method than the one described above with the bullet points. The alternative method is as follows: Recalculation unit 170 combines the additional time series data D1-2 and the original time series data set D1 to generate a time series data set D1-3. Recalculation unit 170 then delivers the generated time series data set D1-3 to preprocessing unit 120. By processing the time series data set D1-3, preprocessing unit 120 then processes it, and visualization unit 160 performs the same process. The choice between the two methods is made based on user input. (Hardware)

[0045] Next, a configuration example of the hardware of the time series data processing unit 100 will be presented with reference to the Fig. 7A and Fig. 7B explains. The corresponding functions of the time series data processing unit 100 are implemented by a processing circuit. The processing circuit can be a Fig. 7A depicts a dedicated processing circuit 400 or a processor 500 for executing programs that are located in a Fig. 7B shows 600 GB of RAM stored.

[0046] If the processing circuitry is the dedicated processing circuit 400, then the dedicated processing circuit 400 is, for example, a single circuit, a composite circuit, a programmed processor, a parallel programmed processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a combination thereof. The functions of the time-series data processing device 100 can be implemented by a variety of separate processing circuits, or the functions of the time-series data processing device 100 can be implemented collectively by a single processing circuit.

[0047] In a case where the processing circuitry is the processor 500, the functions of the time-series data processing device 100 are implemented by software, firmware, or a combination of both. The software and firmware are written as programs and stored in the main memory 600. The processor 500 reads and executes a program stored in the main memory 600, thereby implementing a function of the time-series data processing device 100.Examples of memory type 600 include non-volatile or volatile semiconductor memory such as random-access memory (RAM), read-only memory (ROM), flash memory, erasable programmable read-only memory (EPROM), or electrically erasable programmable read-only memory (EEPROM), magnetic disk, flexible disk, optical disk, compact disk, mini-disk, and DVD. Memory type 600 can be implemented as the same device as memory type 200 or memory type 300.

[0048] It should be noted that some of the functions of the Time Series Data Processing Unit 100 may be implemented by dedicated hardware, while other functions may be implemented by software or firmware. Thus, the processing circuit can implement the functions of the Time Series Data Processing Unit 100 through hardware, software, firmware, or a combination thereof. <funktionsweise>

[0049] Next, an operation carried out by the time series data processing facility 100 will be described with reference to Fig. 8 explained. (Step ST1)

[0050] First, in step ST1, the data input unit 110 obtains the time series data set D1, which contains the candidates for explanatory variables to be used for the prediction. (Step ST2)

[0051] In the next step ST2, the preprocessing unit 120 performs preprocessing of the time series data set D1, which is obtained from the data input unit 110. The preprocessing unit 120 then delivers the time series data set D2 to the grouping unit 130, which is obtained after the preprocessing of the time series data set D1. (Step ST3)

[0052] In step ST3, the grouping unit 130 next classifies and groups the acquired time series data set D2 according to the degrees of similarity between the type of the respective parts of the data. (Step ST4)

[0053] Next, in step ST4, the relationship calculation unit 140 calculates the relationships between the groups. Furthermore, the relationship calculation unit 140 establishes temporal relationships between the groups by comparing the waveforms. (Step ST5)

[0054] In step ST5, the stratification unit 150 next determines the temporal relationships based on the sequence of the various activities or the statistical analysis, e.g., the month-shifted correlation, and stratifies the groups based on the determined temporal relationships. (Step ST6)

[0055] In step ST6, the visualization unit 160 next visualizes the relationships between groups obtained in step ST4 layer by layer, based on the temporal relationships obtained in step ST5. (Step ST7)

[0056] Following the processing sequence of steps ST1 to ST6, the recalculation unit 170 performs a recalculation of the relationships to add additional time series data to the existing groups and visualize it. When the time series data is added, the recalculation of the relationships can be performed quickly by comparing the additional data with the existing groups. <Vorteilhafte Effekte>

[0057] Manual evaluation is essential to selecting explanatory variables that will convince a user when predicting indicators of economic activity. However, manually selecting explanatory variables from such a large number of data points that are both domain-specific and statistically valid is extremely difficult.

[0058] Based on a visualization diagram depicting relationships output by the time-series data processing device 100 according to the present disclosure, it is possible to extract multiple candidates as statistically valid explanatory variables from a large number of data sets. Extracting a large number of candidates in this way reduces the number of targets where a human performs selection / elimination based on domain knowledge. This makes it easier for the person performing the prediction to select / eliminate explanatory variables. Accordingly, it becomes possible to incorporate human domain knowledge while ensuring the statistical utility of the explanatory variables used for prediction.In this way, distrust of the prediction obtained from the explanatory variables can be reduced and the persuasiveness of the prediction can be increased. <implementierungsbeispiel>

[0059] The following section explains a more concrete implementation example of the time series data processing unit 100. A visualization diagram illustrating the relationships output by the time series data processing unit 100 is shown, for example, in Fig. 9 shown. Fig. In section 9, each vertex represents a group of similar time-series data, and each edge connecting the vertices represents the strength of a relationship between the groups. Furthermore, a variety of layers are included in Fig. 9 layers, which are distinguished according to temporal relationships, and Fig. Figure 9 illustrates that groups positioned in higher strata are ahead in time of groups positioned in lower strata.

[0060] As an example, a case is considered in which the future trend of the data “delivery volume of air conditioners”, which is in group 2 in Fig. 9 are included, as predicted. (First step)

[0061] From the information about the relationships between the data in Fig. From point 9, it can be deduced that group 2, which contains the data "Air conditioner delivery volume", has strong relationships with groups 1, 4, and 5. In this way, it is possible to narrow down the candidates for explanatory variables to the data in any of groups 1, 4, and 5 as those suitable for use when predicting the "Air conditioner delivery volume" data. (Second step)

[0062] From the information about the layers in Fig. It can be deduced from point 9 that of the groups (1, 4, and 5) to which the candidates were narrowed down in the first step, only group 1 precedes group 2 in time, which contains the data "Air conditioner delivery volumes." In this way, it is possible to narrow down the candidates for explanatory variables to data in group 1 as one suitable for use when predicting the data "Air conditioner delivery volumes." If certain data Y are to be predicted, and if there are other data X that precede data Y in time, then data X constitutes an explanatory variable (leading indicator) useful for predicting data Y. (Third step)

[0063] It is assumed that the information about the groups in Fig. It can be deduced from section 9 that the data "Number of construction starts for new buildings", "Amount of apple consumption", and "Amount of mackerel harvest" belong to the groups of candidates for explanatory variables to which the candidates were narrowed down in the second step. The data "Number of construction starts for new buildings", "Amount of apple consumption", and "Amount of mackerel harvest" are presented to the user as the final candidates for explanatory variables. (Fourth step)

[0064] The user selects from the candidate data for explanatory variables "Number of new building construction starts," "Amount of apple consumption," and "Amount of mackerel harvest" the data that are most convincing when used for the forecast. For example, there is an intuitively easy-to-grasp causal relationship between the data "Air conditioning delivery volume" and the data "Number of new building construction starts," because "The number of buildings increases. → The demand for newly installed air conditioning units in the buildings increases. → The air conditioning delivery volume increases." Accordingly, the user can simply select the "Number of new building construction starts" as the explanatory variable for forecasting "Air conditioning delivery volume."By performing the prediction using data that exhibits a causal relationship that is intuitively easy to grasp in this way, it is possible to obtain prediction results that convince the user.

[0065] By grouping explanatory variables into grouping unit 130, the number of factors appearing in the visualization diagram can be reduced. As a result, it is possible to progressively narrow down the candidates for explanatory variables, and the effort required when selecting / eliminating explanatory variables can be reduced.

[0066] Visualizing the relationships between candidates for explanatory variables allows the selection / elimination of explanatory variables, which previously relied on domain knowledge or implicit competence of experts, to be performed without human skills.

[0067] When candidate data for explanatory variables or predictor target data are added, assigning the additional data to existing groups according to the process performed by the recalculation unit 170 by grouping explanatory variables in the grouping unit 130 eliminates the need to re-perform grouping, relationship calculation and visualization, and can reduce the time required to re-output the visualization graph.

[0068] It should be noted that embodiments can be combined and that each embodiment can be modified or omitted as required. INDUSTRIAL APPLICABILITY

[0069] The time series data processing device according to the present disclosure can be used as a device for predicting data relating to indicators such as the delivery volume of air conditioners. REFERENCE MARK LIST

[0070] 100: Time series data processing unit, 110: Data input unit, 120: Preprocessing unit, 130: Grouping unit, 140: Relationship calculation unit, 150: Stratification unit, 160: Visualization unit, 170: Recalculation unit, 200: Storage unit, 300: Storage unit, 400: Processing circuit, 500: Processor, 600: Main memory QUOTES INCLUDED IN THE DESCRIPTION

[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature

[0000] WO 2013 / 187295

[0003] < / implementierungsbeispiel> < / funktionsweise> < / konfiguration>

Claims

Time series data processing device, comprising: a data input unit for obtaining a first time series data set containing a plurality of time series data pieces to be treated as candidates for explanatory variables; a preprocessing unit for converting the obtained first time series data set into a format capable of calculating relationships between the plurality of time series data pieces contained in the first time series data set and for generating a second time series data set containing the plurality of time series data pieces after the conversion; a relationship calculation unit for calculating waveform and semantic relationships between the plurality of time series data pieces contained in the second time series data set;a stratification unit for determining temporal relationships between the multitude of time series data parts contained in the second time series dataset, for stratifying the multitude of time series data parts contained in the second time series dataset according to the determined temporal relationships, and for outputting a stratification result; and a visualization unit for generating a visualization diagram that visualizes the multitude of time series data parts contained in the second time series dataset based on the calculated waveform and semantic relationships and the output result of the stratification. Time series data processing device according to claim 1, wherein the layering unit determines the temporal relationships between the plurality of parts of time series data contained in the second time series data set according to the domain knowledge defined by user input. Time series data processing device according to claim 2, wherein the domain knowledge comprises a plurality of words and a definition of temporal relationships between the plurality of words. Time series data processing device according to one of claims 1 to 3, wherein the stratification unit performs a time shift such that correlation coefficients between the plurality of parts of time series data contained in the second time series data set are maximized, shift widths are determined, and temporal relationships between the plurality of parts of time series data contained in the second time series data set are determined from the determined shift widths. Time series data processing device according to one of claims 1 to 4, wherein the waveform and semantic relationships are cross-correlations between the plurality of parts of time series data contained in the second time series data set. Time series data processing device according to one of claims 1 to 5, wherein the visualization diagram is a graph in which all time series data of the plurality of parts of time series data contained in the second time series data set are represented as a vertex and the waveform and semantic relationships between the plurality of parts of time series data are represented as edges. Time series data processing device according to claim 6, wherein the edges vary in line thickness, line color or line type, such as solid line or dashed line, according to the waveform and semantic relationships between the plurality of parts of the time series data contained in the second time series data set. Time series data processing device according to one of claims 1 to 7, further comprising a grouping unit for classifying the plurality of parts of time series data contained in the second time series data set according to the nature of the plurality of parts of time series data into groups and generating a representative value for each group. Time series data processing device according to claim 8, wherein the grouping unit performs a grouping of the plurality of parts of time series data contained in the second time series data set according to degrees of similarity between waveforms of the plurality of parts of time series data contained in the second time series data set. Time series data processing device according to one of claims 8 or 9, wherein the grouping unit generates a representative value for each group that reflects a feature of the data contained in the group. Time series data processing device according to one of claims 1 to 10, further comprising a recalculation unit for accepting additional time series data and performing a recalculation to add the accepted additional time series data to the visualization diagram. Time series data processing device according to one of claims 8 to 11, wherein the recalculation unit determines a representative value of the assumed additional time series data, calculates a correlation coefficient between the determined representative value of the assumed additional time series data and the generated representative value of each group, and assigns the assumed additional time series data to a group whose calculated correlation coefficient is the largest. Time series data processing device according to one of claims 8 to 11, wherein the recalculation unit determines a representative value of the assumed additional time series data, calculates a correlation coefficient between the determined representative value of the assumed additional time series data and the generated representative value of each group, generates a new group containing the assumed additional time series data if none of the calculated correlation coefficients is lower than a predetermined threshold, and calculates the waveform and semantic relationships and temporal relationships between the generated new group and other groups. A time series data processing procedure performed by a time series data processing unit comprising a data input unit, a preprocessing unit, a relationship calculation unit, a stratification unit, and a visualization unit, wherein the time series data processing procedure comprises the following steps: by the data input unit, obtaining a first time series dataset containing a plurality of time series data pieces to be treated as candidates for explanatory variables; by the preprocessing unit, converting the obtained first time series dataset into a format capable of calculating relationships between the plurality of time series data pieces contained in the first time series dataset, and generating a second time series dataset containing the plurality of time series data pieces after the conversion;through the relationship calculation unit, calculating waveform and semantic relationships between the multitude of time series data parts contained in the second time series dataset; through the stratification unit, determining temporal relationships between the multitude of time series data parts contained in the second time series dataset, stratifying the multitude of time series data parts contained in the second time series dataset according to the determined temporal relationships, and outputting a stratification result; and through the visualization unit, generating a visualization graph that visualizes the multitude of time series data parts contained in the second time series dataset based on the calculated waveform and semantic relationships and the output stratification result.