Real-time, multi-frequency seasonal adjustment process for traditional and alternative data.
The STAHL method addresses the challenges of seasonal adjustment in high-frequency data by incorporating a Holiday Trend Decomposition Loop into the STL approach, resulting in improved reliability and speed of seasonal adjustment.
Patent Information
- Application Number
- FR2023013067
- Authority / Receiving Office
- FR · FR
- Patent Type
- Utility models
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2033-11-27
AI Technical Summary
Current methods for seasonal adjustment of high-frequency data face challenges such as complexity in handling multiple and non-integer seasonal patterns, sensitivity to high-frequency outliers, periodic or structural missing values, and the need for timely and unrevised data.
The STAHL method, which stands for Seasonal Trend And Holiday decomposition based on LOESS, adapts the STL approach by incorporating a Holiday Trend Decomposition Loop (HTL loop) to automatically evaluate holiday impacts and improve outlier detection and estimation.
STAHL enhances the reliability and speed of seasonal adjustment by effectively handling complex seasonal patterns, outliers, and missing values, producing accurate and timely seasonally adjusted data.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Method for real-time and multi-frequency seasonal adjustment of traditional data and alternative data.
[0001] Over the past decade, macroeconomic forecasting has increasingly relied on various types of big data. Alternative data refers to non-traditional sources of information that can provide unique insights. Examples of such sources include satellite imagery, IoT devices, sentiment analysis in social media, etc.
[0002] Alternative data can be grouped into four broad categories: text data, geospatial data, geolocation data, and structured data. Text data can be covered by multiple sources, mainly on the internet, such as social media data, professional blogs, news articles, job offers, web searches, or hotel and restaurant reviews, for example. The availability of high-resolution satellite images and the development of deep learning models have led to numerous applications for retrieving various geospatial data from earth observation satellite images, atmospheric data, or radar data. Geolocation data can take the form of maritime traffic, flights, mobility, or vehicle transit numbers.Structured data can include prices of goods and services, real estate prices, internet queries and web traffic.
[0003] The present invention aims to provide a competitive advantage by identifying hidden patterns, detecting emerging trends and improving predictive models by taking advantage of all this alternative data. Most of it is produced at a daily frequency, seven days a week, without exception. These time series have a fairly short history, ranging from five to a little over ten years. Their statistical analysis must therefore answer many specific questions.
[0004] A first problem to be solved is the complexity of the multiplicity and potential interactions of seasonal patterns (notably annual, monthly, weekly and daily cycles) and calendar effects (public holidays, time-varying non-Gregorian calendars such as the Chinese or Hijri calendars).
[0005] In addition, some alternative data may suffer from high sensitivity to high-frequency outliers (mobility ban during COVID periods for example), may be subject to periodic or structural missing values (clouds disrupting the quality of satellite images at re- regular).
[0006] Moreover, most users of indicators constructed from such data expect such alternative data to be timely and unrevised, which, from the perspective of seasonality extraction, constitutes a specific practical challenge.
[0007] The present invention aims to respond precisely to all these problems by introducing a new seasonal adjustment framework subsequently called “STAHL” for Seasonal Trend And Holiday decomposition based on LOESS in English.
[0008] In the following, a brief analysis of the current state of high-frequency seasonal adjustment procedures is presented, with their advantages and disadvantages from the perspective of practitioners and the industry. STATE OF THE ART
[0009] The digital transformation process of our modern economies and the emergence of the COVID-19 pandemic have increased interest in the statistical processing of alternative and high-frequency data. Sub-monthly economic data are in high demand to provide early warning and economic forecasting indicators more quickly, which require specific statistical processing. Webel (2022) provided the most recent general and very comprehensive review of these methods.
[0010] Note that Yt is a discrete time series with a seasonal periodicity of r. For example, r = 12 for monthly data, r = 52.18 for weekly series, and r = 365.2425 for annual data measured at a daily frequency. The consensus in the statistical literature calls high-frequency (HF) time series data observed at sub-monthly intervals (r > 12) (resp. low-frequency (LF) time series if the data are observed at monthly or lower periodicities (r < 12)).
[0011] While different parts of the literature have focused on data minute by minute, hour by hour, 4 hours by hour, the present invention focuses on daily and weekly data in this exercise, in order to be able to produce a wide variety of this type of HF series industrially (i.e. automatically), seven days a week, all year round. At such frequencies, there are already many difficulties due to the interactions of calendar and seasonal effects that are practically absent at low frequencies. For example, as Webel (2022) describes, the effects of fixed holidays / vacations and end-of-period events can depend on the days of the week on which the corresponding events fall. Christmas effects can be significantly different if December 24 or 26 falls from Tuesday to Thursday, from Friday to Sunday. The same applies to the fleeting rise in level at the end of the third quarter if the last day of that quarter had not been a Monday.
[0012] Calendar-related dynamics can become even more complex when secular and religious activities primarily follow different calendars with HF data. Campante and Yanagizawa-Drott (2015) demonstrate that the Muslim calendar has an impact on economic time series such as growth, mobility, and private consumption. In this research, we focus on the impact of the Chinese lunar calendar in the empirical sections to highlight a key improvement of STL by the method according to the invention.
[0013] As illustrated by Webel (2022) or Proietti and Pedregal (2022), the marine profile of HF time series is very complex due to the coexistence of multiple and overlapping seasonal patterns with integer or non-integer periodicities.
[0014] The issue of data accessibility poses another problem. Currently, high-frequency (HF) observations typically provide only a few years of history, which is often inadequate to reliably predict all aspects of HF dynamics, including possible interaction effects. Performing seasonal and calendar extraction on alternative data with an "industrial", i.e., highly standardized, approach is a challenge that must be addressed with as parsimonious a model as possible. Alternative data may also suffer from a significant number of missing values, due to measurement problems or the unavailability of holidays. This adds another constraint to the seasonal adjustment method which must be able to deal with regular or irregular patterns of missing values.
[0015] Extension of conventional methods of the XI1 family
[0016] Many statistical organizations around the world use one of the recognized methods to deseasonalize low-frequency (LF) time series. These methods include the Xl 1 approach, the ARIMA (AutoRegressive Integrated Moving Average) model-based strategy, and structural time series (STS) models.
[0017] Ladiray, Palate, et al. (2018) show that the main seasonal adjustment methods used for monthly and quarterly series, such as TRAMO-SEATS and X12-ARIMA, as well as STL, can be adapted to high-frequency data that have multiple and non-integer periodicities. For example, TRAMO-SEATS can be modified using fractional Arima models and more efficient numerical algorithms. The nonparametric and iterative processes of XI1 and STL can also be adapted after imputation of missing values induced by different month and year lengths. But the authors claim that tuning the multiple parameters of the methods could be tedious.
[0018] Indeed, according to the authors, the adjustment of the parameters of XI1 and STL, namely the length of the filters used in the decomposition, needs to be improved. In the real XI1 algorithm, and for monthly and quarterly series, the order of the moving averages is chosen based on a "signal to noise" ratio. Thresholds have been defined by simulations and by practice, which may constitute an obstacle to an industrial generalization of this procedure for series dealing with multiple high frequencies. For Ladiray, Palate, et al. (2018), large-scale simulations still need to be carried out to understand the behavior of these ratios and thresholds to develop a decision rule for high-frequency data.
[0019] Structural models of time series and their variants
[0020] Proietti and Pedregal (2022) advocate the use of structural time series (STS) models on HF time series. Unobserved component models (UCMs) can, by design, handle any type of high-frequency data, but the models tend to be very "series-specific," according to the authors. In particular, the selection of harmonics can be quite arbitrary and is not easy to achieve at scale. Therefore, UCM models lack the desired flexibility we seek to implement an industrial seasonal adjustment procedure that could be applied to a wide range of alternative data and different frequencies. In testing UCM approaches, the same difficulties were encountered as those indicated by Ollech (2021).We also faced serious identification and convergence problems in the maximum likelihood parameter estimation. The optimization method can indeed have convergence problems with high-frequency data, leading to different local minima depending on the initialization, as reported in Ladiray, Palate et al. (2018). This leads to problems in identifying the signal-to-noise ratios of the irregular component and in separating the trend component from the lower-frequency seasonal components when considering STS and UCM models as an industrial solution for high-frequency seasonal and calendar adjustment.
[0021] In the alternative data and machine learning community, Prophet, a Bayesian approach developed by Facebook's Core Data Science team (Taylor and Letham (2018)) is a frequently popular algorithm for performing seasonal adjustment extraction. It was specifically designed to provide a flexible and reliable forecasting tool that can be configured, interpreted, and evaluated by subject matter experts and analysts without extensive time series modeling expertise. The general idea is to specify models of relatively sparse unobserved components and impose priorities on the unknown parameters. Prophet has many advantages: it was designed to detect annual, weekly, and daily seasonality in the data and to be robust to outliers by using L1 regularization in the modeling process. It provides an intuitive way to include public holidays, provided their effects are known and well identified. However, Prophet was designed more as a forecasting package than as a general statistical decomposition framework. As such, it often operates as a black box, making it more difficult to understand the intricacies of the model. Furthermore, the Prophet framework does not currently allow the seasonal trend to evolve over time, making it too restrictive for our daily macroeconomic series. Furthermore, the trend extraction modeling is based on a saturating nonlinear growth and trend change detection approach, which may also seem too restrictive and not suitable for a one-time procedure.
[0022] The STL method
[0023] The semi-parametric STL approach introduced in R.B. Cleveland et al. (1990) offers an alternative technique for extracting seasonality from HF series. Rooted in a locally weighted regression smoother (LOESS), it differs from Xl 1 in its high flexibility regarding the frequency of the time series: STL can handle any type of frequency, not just monthly or quarterly. One of the main strengths of the STL algorithm over STS is its computational speed, which allows fitting varying seasonal frequencies in a unified iterative process, as well as its robustness to outliers, making it suitable for data with anomalies. The primary objective of STL is the decomposition of a time series into its trend, seasonal, and residual components, thus providing an easily interpretable decomposition.
[0024] Ollech (2021) contributes to the existing literature by designing a seasonal adjustment routine for daily data, but also describes the limitations of STL for single seasonal frequencies, as it cannot handle calendar effects such as the influences of moving holidays. Ollech (2021) also describes the convergence problems of the RegARIMA model for adjusting daily time series for holidays.
[0025] Webel (2022) also provides an update on the automatic procedures needed to produce HF seasonal adjustments. STL and Xl 1 currently lack automatic selection rules for trend filters and seasonal filters from HF data. Furthermore, while there have been many studies on point, i.e., unrevised, versus in-sample seasonal estimates and on the implicit biases associated with trend extraction, there is no procedure entirely dedicated to point seasonal adjustment.
[0026] BRIEF STATEMENT OF THE INVENTION
[0027] The invention therefore aims to overcome these drawbacks of the methods of treatment of data and previous.
[0028] The idea of the invention is to adapt the STL approach by proposing an extended framework to automatically evaluate the impacts of holidays in an integrated manner.
[0029] Thus, the invention aims in particular to improve the calculation speed and the reliability of the detection and estimation of outlier values.
[0030] The idea behind the invention therefore consists of starting from the STL approach and improving it, by inserting at the end of the STL method a loop for decomposing the outlier values of the holidays (hereinafter called "HTL loop") in order to generate a representative parameter, by modifying the first step of the STL to take into account this representative parameter, by performing a feedback loop between the end of the new HTL loop and the modified first step of the STL, and by implementing at least two iterations of the steps, the iterations being able to go up to 15.
[0031] The invention utilizes the discovery by the inventors that although it is useful to know the seasonal period, the STL can be used without this knowledge, at least initially, which is a major asset when designing an industrial statistical decomposition framework.
[0032] To this end, the invention relates to a method for processing at least one series of temporal data, such as daily data, representative of one or more physical quantity(ies) influenced by seasonal effects and implemented in a given industrial process, the method comprising the following steps: decomposition of the time series (Yt) into a trend-cycle component (Tt), a seasonal component (St), a calendar-holiday component (Ht) and an irregular component (It) as in STL using LOESS regressions and moving averages: [Equation 1] Y,= Tt+ St+ H,+ It, for t = 1 to N each value being regressed in LOESS regressions on a local neighborhood of a fitted linear or quadratic polynomial, the LOESS regressions combining the weights of the observations with a bi-weight function so that, for any point in time, the weight of the observation xi5 is given by: [Equation 2] 3 "paved Xq (x) = lx; - xq I the distance between the qth x; the most distant and x; the treatment method being characterized in that it comprises the following steps: a) Adjustment of trends and holidays (S + I = Y- T- 0.5H) b) Preliminary smoothing of the sub-series by LOESS of bandwidth ns to obtain a preliminary factor (^including both irregular and seasonal components; c) Low-pass filter composed of moving average and LOESS asymmetric filters of bandwidth nt to capture low frequency from; d) First seasonal component: — C1} - ; e) First series corrected for seasonal variations: SA^ = Y* - S*; f) First series of trends y* extracted from the SA* series with a LOESS filter; g) First irregular component: = y _ SA^ - ; h) Extraction of potential holiday dates from SA1); i) Estimation of vacancies, H1), by LOESS regression of the bandwidth nh; j) Selection of holidays with adjusted confidence interval; k) Adjustment of working days SA^ = SA, - Hf; 1) SA^)-based trend reestimation with nt bandwidth LOESS regression; m) Reconstruction of the irregular: 7* — Yt - SA} - ; n) Repeating steps a) to m) n0 times, n0 being an integer between 2 and 30, preferably between 4 and 15; o) After X iterations of steps a) to m), generate a trend indicator adapted to the determined industrial process, the indicator being accompanied by a predetermined threshold beyond which an action is programmed to be undertaken in said industrial process.
[0033] According to particular embodiments, one and / or the other of the following provisions is also used: - the threshold may be representative of a duration and / or a critical value; - the method may further comprise a step p) of controlling a means of regulating the industrial process when the predetermined threshold of the trend indicator is exceeded; - step n) may further comprise a calculation of a convergence indicator on the trend-cycle (Tt), and the seasonal component (St); - the method may further comprise a calculation of a confidence interval (CI) - the calculation of the confidence interval can be programmed as follows: 1 For i <= N do 2 For j <= n do 3 V.[ J] WLS(y, X, W[z], j) 4 ^[.d - U [h - y,L / l)2 5 SSTX; - E'JUwMj] - 6 ) *“ -¼ E? i^[ / ] v 7 ¢--2 ' l. / j 7 8 y SSTx, 9 Cf *- y [0] ± (insignificance, • SE, Or : - (Xi,y; ) refers to the set of values (x,y) observed in a kernel K; at point i; - a WLS (Weighted Least Square) regression is applied to obtain a vector of estimates; i - N is the total number of data points (including missing values); - q is the kernel width in LOESS regression - a is the significance level; and - SEi is the standard error for each estimation point in the LOESS regression.
[0034] The invention also relates to a method for producing electricity comprising an electrical production means, a distribution network and electricity consumption means, characterized in that the production means can be controlled by a computer programmed to implement the preceding daily data processing method. BRIEF DESCRIPTION OF THE FIGURES
[0035] [Fig. 1], a diagram illustrating the daily data processing method according to the invention and in which a decomposition is indicated intended for people who are not familiar with Fourier analysis;
[0036] [Fig.2], two comparative curves of the kernels with symmetric filter and asymmetric filter metric;
[0037] [Fig.3], a curve illustrating the revisions of the initial data relating to the compensation claims, from the first to the last vintage
[0038] [Fig.4], a curve illustrating the comparison of the method according to the invention (STAHL) with the first and last vintage of SA
[0039] [Fig.5], a curve illustrating the comparison of STAHL with the first and the last SA vintage - Covid and post-Covid period
[0040] [Fig.6], a curve illustrating electricity consumption over several years;
[0041] [Fig.7], a curve illustrating sub-series on Monday and Sunday electricity
[0042] [Fig.8], a curve illustrating electricity consumption and seasonal adjustment weekly ;
[0043] [Fig.9], a curve illustrating the seasonal, idiosyncratic and ten components electricity consumption trend
[0044] [Fig. 10], three curves illustrating electricity consumption according to three seasonal trends;
[0045] [Fig. 11], a curve illustrating electricity consumption corrected for seasonal variations;
[0046] [Fig. 12], a curve illustrating Internet searches for the Easter bunny
[0047] [Fig. 13], a curve illustrating the seasonal STL adjustment on the curve of the [Fig.12] ;
[0048] [Fig. 14], a curve illustrating Internet searches on the Easter bunny and the estimation of the holiday effect;
[0049] [Fig. 15], a curve illustrating Internet searches for the Easter Bunny and the estimated number of holidays and the Barnacle effect;
[0050] [Fig. 16], a curve illustrating Internet searches for the Easter bunny using HTL and STL methods;
[0051] [Fig. 17], a curve illustrating Internet searches on the Easter bunny using the HTL and STAHL methods;
[0052] [Fig. 18], two curves illustrating the raw data series on trucking in Germany and on pollution in China;
[0053] [Fig. 19], a curve illustrating the effects of holidays on German trucking and Chinese pollution from [Fig. 18];
[0054] [Fig.20], a curve illustrating trucking in Germany corrected for variations seasonal with / without cancellation of holidays;
[0055] [Fig.21], a curve illustrating pollution in Dongguan corrected for sai variations ringing with / without cancellation of holidays
[0056] [Fig.22], a curve illustrating the variance of the sub-series of German trucks;
[0057] [Fig.23], a curve illustrating the variance of the sub-series on pollution in Dongguan
[0058] [Fig.24], a curve illustrating the variance of Internet searches on the rabbit of Easter;
[0059] [Fig.25], a curve illustrating the significance for Easter and Thanksgiving with IC normal ;
[0060] [Fig.26], a curve illustrating the data for Easter and Thanksgiving with IC normal ;
[0061]
[0062]
[0063]
[0064]
[0065]
[0066]
[0067]
[0068]
[0069]
[0070]
[0071]
[0072]
[0073]
[0074]
[0075]
[0076] [Fig.27], a curve illustrating the significance for Easter and Thanksgiving with adjusted CI; [Fig.28], a curve illustrating the data for Easter and Thanksgiving with adjusted CI; [Fig.29], a curve illustrating the spectral density for STAHL SA (seasonally adjusted) and NSA (raw, i.e. not seasonally adjusted) data initial demand; [Fig.30], a curve illustrating the spectral density for STAHL SA and NSA Pollution over Dongguan, China; [Fig.31], a plot showing the spectral density for HTL SA and NSA and the Internet search volume for "Easter Bunny"; [Fig.32] a curve illustrating the spectral density for STAHL SA and NSA of German trucking in terms of mileage. Statement of the invention According to the invention, STAHL decomposes a time series (Yt) into a trend-cycle component (Tt), a seasonal component (St), a calendar-holiday component (Ht) and an irregular component (It) as in STL using LOESS regressions and moving averages: [Equation 1] Y;= Tt+ St+ Ht+1„ for t= LA Each value is regressed in LOESS regressions on a local neighborhood of a fitted linear or quadratic polynomial. The following STAHL implementation is restricted to the linear function of (weighted) observations x. The possibility of using quadratic fitted polynomials to estimate LOESS regressions is ignored. To minimize outlier contamination, LOESS regressions combine the weights of the observations with a bi-weight function. Thus, for any point in time, the weight of the observation x;, is given by: [Equation 2] with Xq (x) = lx; - xq I the distance between the qth furthest x; and x. The parameter q is essential to distinguish the trend from the seasonal component: by determining the number of neighboring observations in the local regression, the higher q, the smoother the identified seasonal factor. The x; close to x have the largest weights which finally decrease to zero at the qth furthest point. [Fig.l] is a diagram illustrating the day data processing process nalières according to the invention. In addition to the decomposition indicated for people who are not familiar with Fourier analysis, a method for automatic detection of seasonality is also included in the procedure.
[0077] Thus, as illustrated in [Fig.l], STAHL comprises: - a first modified STL internal loop composed of seven steps focused on the extraction of a first version of the trend and seasonal components, and - a second internal loop HTL (Holiday Trend Decomposition Loop) composed of six steps and which disentangles the outliers of the holiday estimates and leads to a systemic construction of series corrected for seasonal and holiday variations.
[0078] More specifically, the first modified STL inner loop comprises the following steps:
[0079] a) Adjustment of trends and holidays (S + I = Y- T- 0.5H)
[0080] Only half of the effect of the holidays (term H introduced by the invention) is removed for two reasons. First, to give the seasonal component priority over the holiday component. Second, to ensure that holidays always appear as outliers in outlier detection, which is why outlier detection takes +0.5H into account. In the first iteration of the process, the H term is equal to zero at this stage since it is calculated in subsequent stages. In the second iteration, it is replaced by the value calculated in the first iteration of the entire process, and so on. In the (n)th iteration, the H term is replaced by the value calculated in the (nl)th iteration.
[0081] b) Preliminary smoothing of the sub-series by LOESS of a bandwidth ns to obtain a preliminary factor Cf including both irregular and seasonal components.
[0082] c) Low-pass filter composed of moving average (section 6) and LOESS asymmetric filters of bandwidth nt to capture a low frequency from Cf;
[0083] d) First seasonal component: = - Lj;
[0084] e) First series corrected for seasonal variations: SA^ = - St '
[0085] f) First series of trends extracted from the SA^ series with a LOESS filter;
[0086] g) First irregular component: — Yt - SA^ - ;
[0087] At this stage, the irregular component may contain outliers and non-seasonal, but regular, holiday effects. One of the main innovations of the STAHL framework is the addition of a second inner loop HTL (Holiday Trend Decomposition Loop) which disentangles outliers from holiday estimates and leads to a systemic construction of seasonally adjusted series and a vacation.
[0088] As shown in [Fig.l], the HTL loop comprises the following steps:
[0089] h) Extracting potential vacation dates from SA1};
[0090] i) Estimation of vacancies, fl* by LOESS regression of bandwidth nh;
[0091] j) Selection of holidays with adjusted confidence interval (Annex A.3.2);
[0092] k) Adjustment of working days SA1} = SA1} - H1}
[0093] 1) Reestimation of the trend T1} based on SA1} with LOESS regression of width of band nt;
[0094] m) Reconstruction of the irregular: — Yt - SA} - T1}
[0095] At the end of this last step, an external loop is executed in order to avoid the contamination of outliers in seasonal variation correction. The loop consists of extracting the robustness weights œt from the idiosyncratic component I t using the bi-weight function:
[0096] [Equation 3] [oo^i _ JH ifÉ — 0 otherwise »
[0098] where h = 6 x median(It)
[0099] The robustness weights are then used as multipliers for the LOESS regression weights in the STL and HTL inner loops, which significantly reduces the weights of outliers.
[0100] In this context, the parameters of the procedure are as follows: - npla size of the seasonal period - nsla seasonal LOESS bandwidth - n, the bandwidth of the low-pass LOESS for trend extraction - nja holiday LOESS bandwidth - number of outer loops
[0101] Pretreatment and industrial approach
[0102] Identification of seasonality
[0103] Before defining any other parameters for the STAHL procedure, the first parameter to identify is np, which is the number of points in the seasonal pattern. For this, in the context of high-frequency data, spectral density analysis is very useful to visually identify peaks in the series spectrum, and therefore the seasonal frequency to be corrected.
[0104] The methodology consists of creating n random permutations fy,) z'e[0,n] of a time series y. The permutations eliminate any seasonal trend and the spectral power of y is therefore distributed over all frequencies as noise. In Taking the maximum of the spectral power of each permuted series ~yi, we obtain a threshold for the maximum noise power, which is smaller than the power associated with a seasonal pattern in y. Using n = 100 permutations, we take the 99th highest value in the series as the threshold above which the spectral power of the original series indicates seasonality.
[0105] In practice, there may be spectral leakage problems when trying to correctly identify peaks in the spectral density. However, since there is only a small set of possible seasonal frequencies (annual, quarterly, monthly, weekly), it is sufficient to check whether the spectral power around these frequencies is above a certain threshold.
[0106] Data resampling
[0107] By nature, the STAHL procedure can only be used for integer periodicities. While this is not a problem for monthly and quarterly data, in practice, the periodicities of weekly and daily data are most often non-integer. To apply the procedure to high-frequency data, we need to resample our data.
[0108] For daily data, we use the following methodology. To account for seasonality, the length of each month is extended to 31 days, either by extending the months over 31 points or by adding additional points at the end of each month using cubic spline interpolation, depending on the nature of the data. For annual seasonality, February 29 is removed to ensure consistency of 365 points per year before adjustment. Once the STAHL procedure has been applied to the resampled data, the resampling procedure is reversed: the added dates are removed and the removed dates are added back by interpolation.
[0109] In the case of weekly data, the problem is slightly different due to the staggering effect of weeks within the year: the year does not always start on a specific day of the week, meaning that the underlying annual seasonality is observed with a gradual shift, with 52 or 53 weeks per year. Furthermore, for economic data, the seasonal pattern usually runs exactly between January 1 and December 31 and is just very slightly distorted for leap years, justifying the suppression of February 29. This is not the case for weekly data, where the extent of days covered over 52 weeks varies from year to year.
[0110] Therefore, for the weekly data, we use a time warping procedure to ensure that we have 53 equally spaced points per year, by adding a zero to the Fourier transform of the series (see Appendix A1 for details).
[0111] This method of oversampling data comes with caveats:
[0112] With the time axis, holiday effects become less identifiable due to the higher sampling which creates a spillover effect. However, since the pre- and post-holiday effects are also corrected by the STAHL procedure via the Barnacle procedure (see below), the spillover effect is, in practice, limited.
[0113] The oversampling procedure is not punctual, the value of a point is influenced by the future. Therefore, to reproduce pseudo-real-time conditions, oversampling, seasonal adjustment, and then downsampling vintage by vintage must be carried out iteratively.
[0114] Estimation of STAHL parameters
[0115] In order to estimate the parameters of the STAHL procedure, our approach provides a method to select them: - For ns and nh, the selection is based on a leave-one-out cross-validation. - As for nt, the default value is: „ _ ~ 1 -n^ However, for some series we need a more rigid low-pass filter to separate the trend from the seasonal and holiday component, with a filter bandwidth equal to the size of the algorithm's burn-in period, i.e. nt= np x ns. - The number of outer loops no is determined by convergence criteria on the trend and the seasonal component.
[0116] Seasonal adjustment with missing values
[0117] Handling periodic or structural missing values (air pollution data in northern China regions affected by clouds every winter) is an innovative advantage of our procedure: each instance of a season can have a missing value, normal STL cannot handle such structural or periodic missing values.
[0118] The STL methodology excels at handling a few random missing values. When the smoothed subseries is constructed, LOESS regression can easily fill in the singular missing values. When the subseries are then combined in the low-pass filter, there are no more missing values. However, this process breaks down when the missing values are structural. By this, we mean that a specific subseries consists entirely of missing values. In this case, the moving average in the low-pass filter will encounter problems.
[0119] We address this problem by creating a moving average procedure Robust for missing values. Normally, if the moving average has a length of 15, you would add up the last 15 values and divide by 15. However, if there is a missing value, it is replaced by zero. Then, the last 15 values (including zero) are added up and divided by 14. If the point at which the value is estimated is missing, the output value will also be missing. Finally, to ensure that there are no extreme edge effects, each point is assigned a weight. The weight is 1 if there are no missing values in the moving average; for the previous example, the weight for that point would be 14 / 15. These weights are multiplied by the outlier weights before the LOESS portion of the low-pass filter.
[0120] Adjustment of public holidays within STAHL
[0121] The holiday procedure is one of the main contributions of the STAHL methodology. It allows for the removal of the effects of holidays that are not seasonal in nature. For example, Christmas always occurs at the same time in each seasonal cycle. December 25 is always the 360th day of the year. Thus, when annual seasonal adjustment is applied to a daily series, Christmas will have its own sub-series. However, other holidays, such as Easter, Ramadan, or Chinese New Year, do not occur on the same day each year in the Gregorian calendar cycle. This means that they will not be automatically taken into account during a normal STL procedure. STAHL proposes an innovative procedure to apply the same logic used in STL to remove these specific time-varying holiday effects.
[0122] Several steps are required to successfully extract holidays. The first step is to remove all seasonal holidays from the holiday list, as the STL procedure can handle these itself, as mentioned above. The next step is to test the statistical significance of the holiday dates. Some holidays may have no effect, and adding them may unnecessarily complicate the procedure. Including holidays that are not statistically significant may have a negative effect on seasonal extraction.
[0123] While some holidays, such as Ramadan, last for several days, others, such as Easter, do not. However, this does not mean that their effect is limited to the "official" dates of the holidays. To solve this overflow problem, we developed the Barnacle procedure.
[0124] The barnacle is a crustacean that attaches itself to other animals or the ocean substrate. We gave this name to the procedure because barnacle days are important because of their "attachment" to the events of the festival.
[0125] The Barnacle procedure allows us to process days that are significantly affected by their proximity to a major holiday. We apply the procedure Bamacle by checking whether the days adjacent to the holiday are significant. If these days are significant, the following day is checked, and so on, up to a strict limit of 46 days before and after each holiday. This approach differs from the impact models (constant impact, linear increase and decrease) used, for example, in Ladiray, Palate et al. (2018), because the shape of the impact is more flexible here.
[0126] Point decomposition of seasonal trends with STAHL
[0127] Preferably, when all data series are produced and delivered daily, they cannot be revised. To keep the structure of each series consistent, the historical parts of each series must be calculated in pseudo-real time. This means that each series can only use historical data for a single estimate. Furthermore, we do not want past values to change when additional data is added to the series.
[0128] In the classical STL procedure, there are three different steps in which forward bias can occur. The first, and most obvious, occurs at any step of the LOESS regressions. The second occurs during the low-pass filtering step, when a moving average is applied in both directions. Finally, during outlier detection, the median value that is normally calculated from the entire data set can introduce bias. To industrialize and create reproducible series, these issues must be addressed.
[0129] Asymmetric LOESS
[0130] As a kernel regression, LOESS applies a weight to points on either side of the estimation point. It also applies a non-zero weight to points in the future. We circumvent this problem by applying an asymmetric kernel regression.
[0131] LOESS regression differs from more conventional kernel regression in the way the kernel width is defined. In conventional kernel regression, the bandwidth is fixed and the number of points used to make an estimate varies depending on the number of data points present around the estimation point. In LOESS regression, the number of points is fixed and the kernel width is defined so that the same number of points are always present.
[0132] Starting from a kernel width of zero, in a normal symmetric LOESS, the kernel width is slowly expanded in either direction until the required number of data points are within the bandwidth. In the case of an asymmetric kernel, only the bandwidth on the left side is expanded, while the right side of the estimation point has zero bandwidth.
[0133] In the case of an edge, at the beginning or end of the data set, the bandwidth can only extend in one direction. This is why the symmetric and asymmetric LOESS filters have identical kernels at the beginning and end. At the beginning, this means that there is a slight phase shift. The comparison of the kernels at the end of [Fig.2] clearly shows how the asymmetric kernel reproduces the point estimate of a symmetric LOESS.
[0134] The LOESS asymmetric filter may not be the most optimal filter in terms of gain and phase shift. It is also possible to use more optimal asymmetric filters that balance the properties of the moving average in terms of accuracy, revisions and timeliness and that allow to construct asymmetric moving averages that exhibit almost no phase shift. It is also possible to use the non-parametric approach based on spectral theory with the family of dynamic asymmetric filters. This procedure, designed in the frequency domain, would allow to split the revision criterion into two distinct effects, one related to the gain and the other to the phase of the transfer functions.
[0135] Moving Average Extrapolation Procedure
[0136] In the STL procedure, the smoothed subseries for days is extended to two more values than the original subseries. The most distant values are determined by linear extrapolation. This ensures that when the bidirectional moving series
[0137] If the inverse moving average procedure is applied, the final series has a length equal to that of the original series. However, a problem arises in that the inverse moving average (MA) procedure uses future values. To circumvent this problem, instead of directly using the smoothed subseries values, the future values are extrapolated linearly to each point in time. Then, the moving average procedure is applied to these extrapolated values.
[0138] Weighting of outliers
[0139] When weighting outliers, all values are divided by six times the median value of the series. If the entire series is considered, this can be considered forward-looking behavior. To avoid this, we set a validation date, so that only the median of the series up to the validation date is considered. This approach ensures that the values in the series do not change when more values are added in the future.
[0140] Critical running-in period
[0141] The critical burn-in period corresponds to the time at the beginning of the estimation, during which we are not yet able to produce pseudo-real-time data. Each of the three sources of forward-looking bias has a different amortization period and affects the structure of the data generation process in a unique way.
[0142] During the early stages of asymmetric LOESS, the nucleus starts with a straight tail, then it becomes symmetrical before finally reaching the tail shape desired left. LOESS is used in subseries, low pass and trend.
[0143] The last LOESS to reach the left tail shape is the LOESS subseries. This occurs at point ns xnp. Before this burn-in is complete, the data will be heavily biased forward, and because the weighted center of the kernel is further forward, it will also have a slight phase shift compared to the data after burn-in.
[0144] The extrapolating moving average needs at least two values before it starts extrapolating. Since it extrapolated from the same point in each cycle, it needs 2xnp points before it completes its burn-in. The change in the seasonal extraction process is due to the fact that the extrapolated data tends to be slightly more volatile than the original values. Since they are averaged again, this does not have a significant impact on the quality of the estimate.
[0145] Finally, the validation date can be set freely and does not have a significant impact on the seasonal extraction process. It is therefore logical to set it to at least ns xnp, but it does not pose significant prospective problems if it is set further in the future. Once a series is put into production, this value should not be changed in order to guarantee the stability of the history.
[0146] Multiple seasonality
[0147] The STAHL procedure treats seasonality iteratively, starting by removing seasonality from the highest frequency and then moving to lower frequencies.
[0148] Empirical illustrations
[0149] Example 1: Traditional Weekly Data: Initial Claims for Unemployment Insurance in the United States
[0150] Description of data
[0151] As a first example, we applied the STAHL procedure to the series "US Initial Unemployment Claims." This weekly series, published by the U.S. Employment and Training Administration, covers the number of initial unemployment claims filed by unemployed U.S. workers in order to qualify for unemployment benefits. It is published every Wednesday as a seasonally adjusted series and as a non-seasonally adjusted series. This series is relevant to illustrate our seasonal adjustment procedure for the following reasons:
[0152] Its weekly frequency allows us to illustrate the resampling procedure. -Available since 1968, its long history allows us to illustrate the influence of the economic cycle on the trend. We also have access to all data vintages since 2009, which allows us to assess the impact of revisions in the raw data (NSA) and seasonally adjusted data (SA).
[0153] It presents outliers which may be due to external catastrophic events (hurricanes, Covid pandemic) or labor-related events (strikes).
[0154] This is a relevant series for monitoring the macroeconomic situation and the business cycle, used, for example, in high-frequency composite indicators such as the Aruoba-Diebold-Scotti Business Conditions Index. The Aruoba-Diebold-Scotti Business Conditions Index, published by the Federal Reserve Bank of Philadelphia, is designed to track real economic conditions at a high observation frequency. Its underlying economic indicators, which are seasonally adjusted, include weekly initial jobless claims, monthly payroll employment, monthly industrial production, monthly real personal income less transfer payments, monthly real manufacturing and trade sales, and quarterly real GDP. It is a mixture of high- and low-frequency data.
[0155] Official seasonal adjustment method
[0156] The official seasonal adjustment of initial requests (https: / / www.bls.gov / lau / seasonal-adjustment-for-weekly-unemployment-insurance-cl aims) is conducted by the Bureau of Labor Statistics (BLS) and is based on the methodology described in W.P. Cleveland, Evans, and Scott (2014) since 2002. The series has annual multiplicity, and seasonal factors are calculated using locally weighted regression on sine and cosine terms, with the weights extracted from state-space modeling of the series. There is no trend extraction per se, as the series are detrended by differencing before seasonal adjustment. The holiday procedure is based on the X-13-ARIMA-SEATS program, with time-constant holiday weights. The adjustment for outliers and interventions also uses the X-13-ARIMA-SEATS program.
[0157] Both raw and seasonally adjusted official data are subject to revision with each release. The main source of change is the revision of seasonal factors, as shown in [Fig. 3]. In this illustration, we focus on the demand values for the first release and the latest available vintage (end of October 2023).
[0158] The Covid-19 pandemic has also brought about other changes in the me methodology. Due to its extreme impact on unemployment claims, the BLS changed the seasonal factors from multiplicative to additive for the period March 2020 through June 2021, and then back to multiplicative after that date. The extreme nature of the event and its duration also necessitated an ex post update (in April 2023) of the outlier specification during the pandemic period, which resulted in significant revisions to seasonal factors from the end of 2021 onwards.
[0159] Implementation of STAHL
[0160] We will now apply the STAHL procedure to produce our own version of the seasonally adjusted unemployment claims data. In doing so, we work under real-time conditions, meaning there is no revision of the data, and we use real-time data for the vintages from May 2009 to November 2023.
[0161] First, following the official procedure, we assume a multiplicative seasonality pattern, meaning that a logarithmic transformation is applied. Given the weekly frequency of the data, we apply the resampling procedures described in Section 4.2 to obtain 53 equally spaced data points per year. The resampled holidays given to the HTL procedure are: New Year's Day, Washington's Birthday, Memorial Day, Independence Day, Labor Day, Columbus Day, Veteran's Day, Thanksgiving, Christmas, and Easter. The Bamacle procedure detects the effects of Bamacle days before and after Thanksgiving and Independence Day.
[0162] During the Covid period, we do not switch to the additive method and retain the multiplicative scheme. However, even with the built-in outlier robustness of the STAHL procedure, the unprecedented extreme behavior of the series during this period contaminates the seasonal adjustment for the years 2022 and 2023. Therefore, we follow the BLS choice and remove the period from March 2020 to June 2021 from the data when running the STAHL procedure starting in January 2022.
[0163] One-off seasonal adjustment with STAHL
[0164] We observe in [Fig.4] that the seasonal correction obtained by STAHL closely follows the official seasonally adjusted signal, both in terms of trend and idiosyncratic behavior. The procedure is robust to outliers. For example, the amplitude of the August 2017 spike due to Hurricane Harvey matches that measured in the official data. However, the volatility around the Thanksgiving period is higher for our model than for the official data, a consequence of oversampling.
[0165] Regarding the Covid period, as shown in [Fig.5], STAHL is sufficiently robust to outliers, so that maintaining a multiplicative scheme gives similar results to the BLS methodology for the period from April 2020 to December 2021. It should be noted, however, that STAHL's output briefly diverges from the official data in January 2021, before recover. For the post-Covid period, STAHL's real-time production is less volatile than initial estimates and closer to the latest vintage. This illustrates both the stability of our model over time and the benefit of hindsight provided by excluding the Covid period from 2022.
[0166] A visual analysis of the spectral density for the original unadjusted series and the STAHL-adjusted series (see [Fig.29] in Appendix A.4) shows that the procedure eliminates all seasonal peaks from the spectrum while preserving the low-frequency components.
[0167] The parameters presented in Table 1 confirm that, in the post-Covid period, the STAHL model is closer to the latest estimates than to the first vintage in terms of mean absolute log error. In the pre-Covid period, the mean log error of STAHL with respect to the first and latest vintages is close to 0.02, which is significant compared to the revised MLE. This suggests that the STAHL results tend to underestimate the underlying trend compared to the official data. This discrepancy could be attributed to differences in trend separation methods. This underestimation also results in a median absolute log error of about 0.03. This problem could be addressed by forcing the seasonal factor for each year to sum to zero.
[0168] [Tables 1] compared to the first estimate compared to the last estimate Official revisions Metrics MLEMALE MLEMALE MLEMALE 2009-2019 0.0160 .026 0.0220 .028 0.0030 .015 2020-2021 0.0060 .05 -0.0250 .061 -0.0090 .04 2022-2023 0.030 .065 0.0080 .045 0. 00.054
[0169] Table 1 represents the median logarithmic error and the median absolute logarithmic error of STAHL compared to official data.
[0170] With this illustration, we demonstrate STAHL's ability to reproduce the seasonal adjustment of the official methodology for high-frequency data using a robust, more parsimonious, and purely ad hoc method.
[0171] Example 2: Daily electricity data with multiple seasonality
[0172] Description of data
[0173] In this section, we analyze an empirical application of the STL procedure with multiple seasonality. We seasonally adjust the electricity consumption data in MWh of the US electric grid, published by the Energy Information Administration (https: / / www.eia.gov / electricity / gridmonitor / dashboard / electric_overview / US48 / US48), with a history from mid-2015 to early 2023. The data are daily and have a weekly and annual seasonal pattern. Although the electricity data contain useful economic information on consumption and production, their low signal-to-seasonality ratio can make their use difficult. This series seems to be a perfect candidate to demonstrate the effectiveness of the STL procedure. [Fig.6] presents the raw, non-seasonally adjusted electricity data.
[0174] When dealing with multiple seasonalities, the appropriate procedure is to deseasonalize the shortest cycle first, then the longest cycle. Since we are dealing with a weekly seasonality (7) and an annual seasonality (365), we start by removing the weekly seasonality.
[0175] Weekly seasonal adjustment
[0176] The first step in the STL procedure is to create subseries. Since our weekly seasonal cycle has 7 days, we create 7 subseries, one for each day of the week. Each subseries only includes Mondays, Tuesdays, etc. An example of subseries for Mondays and Sundays is shown in [Fig.7]. Here, we can clearly see that electricity consumption is, on average, lower on Sundays than on Mondays. To capture this effect, a LOESS regression is fitted to the data. This creates the smoothed subseries, which is conceptually similar to a seasonal variable for days like Monday or Sunday. The main advantage over a simple seasonal variable is that it is much more flexible and can account for changing seasonal dynamics.
[0177] [Fig.8] shows the first year of data with the weekly seasonality removed. The series appears much smoother and no longer exhibits weekly oscillations. However, much of the annual seasonality is still visible, as there is a peak during the summer related to air conditioning and a second, smaller peak in the winter for heating demand.
[0178] Annual seasonal adjustment
[0179] Following the same procedure, to obtain the seasonally adjusted week-year series, a second STL procedure is then applied. [Fig.9] illustrates how the different components of the series unfold. It shows how much of the variation in the series comes from the annual seasonal component and why it is so important to seasonally adjust such series before trying to use them for econometric procedures.
[0180] Seasonally adjusted series
[0181] [Fig.10] shows the weekly and annual seasonal components, as well as the extracted total seasonality. [Fig. 11] shows the difference between the original series and the fully deseasonalized series.
[0182] Example 3: Holiday Adjustment: Google Searches for “Easter Bunny”
[0183] Data Description
[0184] Easter is a difficult holiday to account for in a seasonal adjustment procedure because it does not have a fixed place in the Gregorian calendar. To illustrate its effect, we chose the Google search volume series for the keywords "Easter bunny." While most series have both seasonal and holiday effects, this series was specifically selected because all movements can clearly and easily be attributed solely to the holiday effect. It therefore provides an excellent test bed to demonstrate the problems that can arise when attempting to correct for holiday effects using the standard STL procedure, which is designed for regular seasonal effects. The Google search volume series starts in 2010 and extends to August 2023, comprising 4981 daily values.They range from 0 to 100, with the values representing the relative search volume for the period.
[0185] Application of classic STL
[0186] If we examine the first 500 days of Easter Bunny searches on Google in [Fig. 12], two of the Easter events can be easily isolated. The increase in searches starts quite a long time before Easter and quickly declines once Easter is over. If we examine the same period after deseasonalizing the series using the classic STL approach ([Fig. 13]), the results appear rather disappointing. For each of the two Easter events, STL assigns negative values: for the first day before Easter and for the second day after Easter. Obviously, the fact that Easter is moving disrupts the STL procedure, which tries to find some kind of "average" effect of Easter over a longer period.
[0187] Application of HTL
[0188] Since it is obvious that the classical STL procedure is not suitable for handling moving holiday effects, we developed the HTL procedure. Conceptually, the HTL procedure works very similarly to the STL procedure. The main difference is that the sub-series only include holiday events, instead of having a sub-series for each period of a seasonal cycle.
[0189] [Fig. 14] shows the estimated effect of the Easter holidays with the original series. Although Even if Easter is considered significant (and an "Easter effect" is therefore estimated), the results are not very impressive. The estimated Easter effect appears to be much smaller than the true effect and is limited to a single day per Easter event. To correct for this, we developed the Barnacle procedure (see Appendix). In the Barnacle procedure, we begin by testing the statistical significance of the holiday effect. If the holiday is found to be significant, we repeat the process with the days on either side of the holiday, i.e., the Barnacle days. We continue this process until we find a non-significant day or reach the 46-day limit.
[0190] In [Fig. 15], we can see that once we have included the Barnacle day effects, our estimates are much closer to the original series than before. [Fig. 16] illustrates the dramatic improvement brought about by switching from STL to HTL with the Barnacle procedure. Thus, we are able to reduce the positive and negative extremes of the series.
[0191] STAHL Easter Bunny Search Application
[0192] This Easter Bunny search series is simpler than many other series that can be used and processed daily by the method according to the invention. The main difference is that this series only has holiday effects, whereas most series have a mixture of holiday and seasonal effects. We will examine this in more detail in the next section.
[0193] [Fig. 17] shows the seasonally adjusted series with HTL and STAHL. We can observe that the STAHL seasonal adjustment has a little difficulty in handling the negative effects, as we observed in the STL procedure. However, it is also evident that STAHL can handle the dates close to the Easter event quite effectively. The effectiveness of the HTL adjustment is visually confirmed by the spectral density plot (see Appendix A.4, [Fig.31]), where we observe that the HTL procedure removes the rapid succession of peaks present in the density of the non-seasonally adjusted (NSA) data. These peaks correspond to the Fourier transform of a Dirac comb-shaped holiday pattern (A Dirac comb is a pattern of regularly spaced peaks in the time domain).
[0194] Example 4: STAHL application on alternative data
[0195] Description of data
[0196] This section covers the application of STAHL to alternative data sets. To illustrate the empirical uses of STAHL, we present two different sets. The first is a set on daily truck mileage on German autobahns, and the second is a set on air pollution (NO2) levels in the city of Dongguan in China. Both sets have very high levels of NO2. seasonality and significant moving holiday effects. [Fig. 18] shows the raw series, both of which exhibit clear annual seasonality.
[0197] Dongguan is an important industrial center, particularly for manufacturing and exports, in southern China. The pollution data series for this city is recorded by the Ozone Monitoring Instrument (OMI) aboard NASA's Aura satellite, which was launched in 2004. The series begins on January 1, 2005, and extends to December 31, 2022, providing a total of 6503 data points.
[0198] Daily truck mileage data in Germany are available online on the Destatis website (see https: / / www.destatis.de / EN / Service / EXSTAT / Datensaetze / truck-toll-mileage.html). In Germany, there is a general driving ban for vehicles over 7.5 tonnes on public holidays. This means that in addition to the seasonal effects of the series, they also show significant holiday effects. Ignoring these holiday effects can have very detrimental effects when interpreting the series.
[0199] Holiday effects
[0200] The holiday adjustment in the Chinese series only applies to the Chinese New Year. The Chinese New Year is an important family holiday in China, during which many workers take time off to visit their families. Therefore, industrial production, especially in human capital-intensive industries, decreases during this period. This should lead to a predictable drop in pollution during the Chinese New Year. Since the Chinese New Year marks the beginning of a new year in the Chinese lunisolar calendar, it does not correspond perfectly to a specific day in the Gregorian calendar. Therefore, it is a moveable holiday from the perspective of the Gregorian calendar.
[0201] For the adjustment of public holidays in Germany, three movable holidays must be taken into account: Easter, Ascension Day and Whit Monday. These are all public holidays in Germany and, as such, do not allow the driving of heavy vehicles. For both series, these holidays lead to significant drops in activity.
[0202] [Fig. 19] shows the holiday effects of the two series for the first two years.
[0203] Barnacle effects for Chinese New Year range from -2 to +14 days. Barnacle days for Easter, Ascension Day, and Whit Monday are (-3,+6), (0,+6), and (0,+6), respectively.
[0204] Seasonally adjusted series
[0205] Figures 20 and 21 show the first two years of the series after correction for seasonal variations, with and without removal of the holiday effect. The effect of The missing holiday effect in the German road transport series is much more apparent than in the Chinese pollution series. However, both effects are significant in the Barnacle procedure, indicating that they are clearly predictable. Just like predictable seasonal patterns, predictable holiday effects are not informative for nowcasting economic time series and should therefore be eliminated from the series.
[0206] We can also examine the variance by calendar day. Although the effect is smaller in the Chinese pollution series than in the German trucking series, we can clearly see a decline in the series variance during potential holiday periods in [Fig. 22], and to a lesser extent in [Fig. 23] for the Dongguan pollution series. In the German series, we can also observe that even regular holidays (such as Christmas) show higher volatility than "normal" periods.
[0207] If we observe the spectral density of the Dongguan series, the STAHL procedure eliminates the peak corresponding to the annual periodicity while preserving the shape of the spectrum for the other frequencies (see [Fig.30] in appendix A.4).
[0208] The spectral density of the STAHL-corrected German truck series shows that peaks with annual periodicity are suppressed, as are the multiple peaks that follow. These correspond in the frequency domain to the seasonal pattern and the Dirac comb-shaped holiday pattern, which are suppressed by the procedure (see [Fig.32] in Appendix A.4).
[0209] Conclusion
[0210] The STAHL (Seasonal Trend And Holiday decomposition based on LOESS) data processing method is an innovative framework for seasonal adjustment developed by the applicant. STAHL was designed to take into account the complex seasonal patterns observed in various daily data sets, such as satellite imagery, textual analysis, daily price records or web searches. The utility of STAHL is manifold, but its three main improvements stand out.
[0211] The first is the procedure's suitability for point estimation, which significantly enriches the "industrial output" of seasonally adjusted series. This advance paves the way for further exploratory research on different asymmetric filters to potentially improve the point trend extraction procedure. In addition, we introduce a "holiday loop" feature that effectively accounts for holiday influences, skillfully adjusting for time-varying calendar effects, including effects as diverse as Chinese New Year and the Hijri calendar. Finally, STAHL's improved ability to handle missing values represents a quantum leap in ahead in terms of robustness, facilitating the management of gaps in seasonal data.
[0212] STAHL's practical applications have been extensively tested, highlighting its ability to efficiently produce seasonally and trading-day adjusted data across a range of frequencies and series types, including traditional and alternative series. This sets a new benchmark for industrial-scale seasonal adjustment of high-frequency data sets and may broaden the horizon for future applications in this field.
[0213] As an example of STAHL implementation:
[0214] In electricity production companies: It becomes possible to measure electricity demand / consumption at high frequency (hourly, daily, weekly, etc.) and to break it down into trends (structural changes), predictable seasonal components, predictable and irregular holiday effects, which makes it possible to forecast demand and load on the local electricity network and to capture unexplained surprises or load anomalies in an industrial manner. It is then, for example, possible, in advance and precisely, from one or more indicators generated at the end of the data processing according to the invention, to adapt electricity production, to predict overloads and therefore to anticipate risks of breakdowns by preventively deploying countermeasure means, to plan maintenance interventions during lower demand, etc.
[0215] In cloud computing service providers or internet service providers: the intensity of high-frequency demand (hourly, daily, weekly, etc.) and server load are measured and the trend (structural evolution), the predictable seasonal component, and the predictable and irregular holiday effect are broken down industrially, making it possible to make forecasts, in advance and precisely, of demand and to capture unexplained surprises or load anomalies. It is then possible, in the event of a predicted threshold being exceeded (value and / or duration), to reallocate IT resources and / or schedule maintenance operations before overload to limit the risk of breakdown.
[0216] For storage companies and e-commerce platforms: we measure the intensity of demand for e-commerce products to break down, on an industrial scale, the demand into trends (structural evolution), predictable seasonal components, and the predictable and irregular holiday effect, which makes it possible to forecast demand and capture unexplained surprises on an industrial scale. It is then possible to take preparatory measures in advance and precisely, such as stock management and distribution, deliveries, anticipating fuel demands and scheduling its purchase and storage when prices fall, etc.
[0217] For air transport companies: measurement of the intensity of high-frequency reservation demand (hourly, daily, weekly, etc.) allowing the breakdown of the trend (structural evolution), the predictable seasonal component, and the predictable and irregular holiday effect, which makes it possible to carry out demand forecasts (occupancy rate and reservation rate) and to capture unexplained surprises or load anomalies in an industrial manner. It is then possible to take preparatory measures, such as flight management and distribution, anticipate fuel demands and schedule its purchase and storage when prices fall, schedule maintenance operations or equipment rental (engine, aircraft for example) in advance and precisely.
[0218] Thus, the invention finds application in any activity requiring the measurement of production or demand / consumption in real time regardless of the frequency (minutes, hourly, daily) which needs to extract a trend, seasonality effects and predictable public holidays and isolate surprises / anomalies allowing their activity to be better predicted with great ease (industrial production logic).
[0219] Appendix
[0220] A. 1 Sampling Procedure
[0221] Oversampling weekly data involves transforming our signal from a sampling rate of 365.25 / 7 ~ 52.18 to a sampling rate of exactly 53. In practice, given y a seasonal and weekly time series and "y its discrete Fourier transform, we obtain the modified Fourier transform by extracting "z:
[0222] [Equation 4]
[0223] A 7 “ ' l 0 sin(m J
[0224] where M = 52.18 / 53 and Z is the support of "y.
[0225] By taking the inverse Fourier transform or "z, we obtain a time series z that has exactly the same spectral signature as y, but sampled at a frequency that is suitable for the STAHL procedure. After seasonal adjustment of the oversampled data, with parameter np set to 53, we can resample the data to their weekly frequency by applying the same methodology with M' = 53 / 52.18.
[0226] A.2 LOESS, Confidence interval
[0227] In this part, we need to clarify some notions. (xi,yi ) refers to the set of (x,y) values observed in the kernel Ki at point i. The kernel simply applies the set of weights (as can be seen in Figure 2) to the points. A WLS (Weighted Least Square) regression is then applied to obtain the vector of estimates y.. N is the total number of data points (including missing values), q refers to the kernel width in LOESS, a is the significance level. We end with SEi which is the standard error for each estimation point in LOESS. Once this is calculated, it is easy to use a t-statistic (q) to find the confidence interval (CI) around the estimated regression line. The validity of the LOESS confidence procedure is confirmed by Monte Carlo simulations.
[0228] [Tables2] 1 For i < = N do 2 Forj <= n do 3 ' WL^y, x, H | / |. j) 4 ^[y] - (Ud - TDlf 5 SSTXj - . |;| - Wx,) 6 0 F2 7 8 SEi 1 vw(v.) ■ va^.v,) V 557¾ 9 Cf y.[0] ± (^significance, q) ■ SEi
[0229] A.3 HTL - Bamacle Procedure
[0230] The Bamacle procedure we developed here works particularly well for capturing flexible holiday effects. A particular challenge was ensuring that the number of false positives was minimized. To illustrate this point, we show how the Easter Bunny series responds to the inclusion of a Thanksgiving holiday dummy variable.
[0231] A.3.1 The Bamacle algorithm
[0232] Notation: L is an N x h matrix where N is the number of observations and h is the number of holidays evaluated. Lb is a shifted version of L to mark Bamacle days. In this example, h is equal to 2, Easter and Thanksgiving. L contains 0 for days that are not holidays and 1 for days that are holidays. H is a matrix with n x h values. n is the number of holidays observed in the series, for Easter or Thanksgiving, this means 1 holiday per year, so n « N. ycH is a vector of length N that contains all the estimated effects of public holidays.
[0233] [Tables3] Algorithm 2 Bamade Procedure 10 n 12 13 14 15 16 17 18 19 20 ] H y[I =1] I * LOESS < H )______ | if is important then ] î 1 ] then is important fair H i +1 H 2^ eJum^L> i ) HH^r[I=l] | LOESS (H) H if H is important then | ! L: J; 11 • >3 i I i 0 | while H is important fair H ? -1 HLK *— f ) HH <- r[Z = 1] H LOESS (H ) § ^if H is important then IL 1
[0234] A.3.2 Testing significance
[0235] Normal significance tests are performed by examining the portion of the confidence interval (CI) that contains the x-axis. When testing the significance of holidays, the main problem is that holidays tend to introduce a lot of extra volatility into the series. [Fig. 12] clearly shows that the series is very volatile around Easter events, while it is fairly stable the rest of the year. This poses a problem of false positives when testing the significance of other holidays occurring during "stable" periods. [Fig. 24] shows the variance of the Easter Bunny series for each day of the year. It is clear that the variance of searches is much higher during the period when Easter holidays are possible than during the rest of the year. During the period of possible Thanksgiving dates, the variance is close to zero.This poses a problem because the CI depends strongly on the variance of the true series.
[0236] [Fig.25] shows the subseries and smoothed regression with the Easter and Thanksgiving CIs next to each other. Both series pass the significance test of 80%. However, if we examine the scale and plot the two series ([Fig.26]), it is clear that Easter has a much larger effect on the series than Thanksgiving.
[0237] To counter these effects, we use an adjusted CI (aCI) to assess its significance. The aCI uses a multiplier to increase or decrease the CI of a series. To calculate the aCI, we take the average variance of the series and divide it by the variance of the holiday subseries.
[0238] [Equation 5]
[0239] average / m£ daily sub-series} aC! = Ct -.................................................. [going on vacation]
[0240] When we apply the ICA, we can see in [Fig.27] that the significance values have changed. The Easter sub-series now has a significance of 90%, while the significance of the Thanksgiving series is now only 10%. [Fig.28] shows that the width of the aCI of the two series is now very similar. The difference comes from the magnitude of the effect.
[0241] A.4 Spectral density plans
[0242] See Figures 29 to 32.
[0243] References: - Campante, Filipe and David Yanagizawa-Drott (2015). "Does Religion Affect Economic Growth and Happiness? Evidence from Ramadan". In: The Quarterly Journal of Economics 130(2), pp. 615-658. - Cleveland, Robert B. et al. (1990). "STL: A seasonal-trend decomposition". In: J.Off. Stat 6(1), pp. 3-73. - Ladiray, Dominique, Jean Palate, et al. (2018). "Deseasonalization of daily data." In: Proceedings of the 16th IAOS Conference, Paris, France, pp. 19-21. - Ollech, Daniel (2021). "Seasonal Adjustment of Daily Time Series". In: Journal of Time Series Econometrics 13(2), pp. 235-264. - Proietti, T and DJ Pedregal (2022). Seasonality in High Frequency Time Series. Econometrics & Statistics. - Taylor, Sean J and Benjamin Letham (2018). "Forecasting at scale". In: The American Statistician 72(1), pp. 37-45. - Webel, Karsten (2022). "A review of some recent developments in the modelling and seasonal adjustment of infra-monthly time series" In.
Claims
Claims
1. Method for processing at least one series of temporal data, such as daily data, representative of one or more physical quantity(ies) influenced by seasonal effects and implemented in a given industrial process, the method comprising the following steps: decomposition of the time series (Yt) into a trend-cycle component (Tt), a seasonal component (St), a calendar-holiday component (Ht) and an irregular component (It) as in STL using LOESS regressions and moving averages: [Equation 1] L= Tt+ St+ H,+ It, for t = 1 to N each value being regressed in LOESS regressions on a local neighborhood of a fitted linear or quadratic polynomial, the LOESS regressions combining the weights of the observations with a bi-weight function so that, for any point in time, the weight of the observation xi5 is given by: [Equation 2] 3 -with Xq (x) = lx; - xq I the distance between the q th x; the furthest and x; the treatment method being characterized in that it comprises the following steps: a) Adjustment of trends and holidays (S + I = Y- T- 0.5H) b) Preliminary smoothing of the sub-series by LOESS of bandwidth ns to obtain a preliminary factor including both irregular and seasonal components; c) Low-pass filter composed of moving average and asymmetric LOESS filters of bandwidth nt to capture low frequency from C*; d) First seasonal component: S^ = Cf - Lf; e) First series corrected for seasonal variations: SA1} = y* - S1} f) First trend series extracted from the SA1 series with a LOESS filter; g) First irregular component: = Yt - SAf - Tf;
2.
3.
4.
5.
6. h) Extraction of potential holiday dates from SA^; i) Estimation of vacancies, by LOESS regression of bandwidth nh; j) Selection of holidays with adjusted confidence interval; k) Adjustment of working days SA^ = SA^ - ; 1) SA^-based trend reestimation with nt bandwidth LOESS regression; m) Reconstruction of the irregular: — Yt - SA^ - ; n) Repeat steps a) to m) n0 times, n0 being an integer between 2 and 30, preferably between 4 and 15; o) After X iterations of steps a) to m), generate a trend indicator adapted to the determined industrial process, the indicator being accompanied by a predetermined threshold beyond which an action is programmed to be undertaken in said industrial process. Method according to the preceding claim, in which the threshold is representative of a duration and / or a critical value. Method according to any one of claims 1 or 2, further comprising a step p) of controlling a means for regulating the industrial process when the predetermined threshold of the trend indicator is exceeded. Method according to any one of claims 1 to 3, in which step n) further comprises a calculation of a convergence indicator on the trend-cycle (Tt), and the seasonal component (St). A method according to any one of claims 1 to 4, further comprising calculating a confidence interval (CI). The method of claim 5, wherein the calculation of the confidence interval is programmed as follows: 1 For i <= N do 2 For j <= n do 3 v U] WLS(y, x, w[z], j) 4 • ",1'1 - W 5 SSTx, E'U'M. / ] - Wx^ 6 7 8 Imrly ) ■ SEJ— y sstx,- 9 Cf >'([0] ± ( / (significance, qj ■ SE, Or : - (Xi,y; ) refers to the set of values (x,y) observed in a kernel K; at point i; - a WLS (Weighted Least Square) regression is applied to obtain a vector of estimates y; * / - N is the total number of data points (including missing values); - q is the kernel width in LOESS regression - a is the significance level; and - SEi is the standard error for each estimation point in the LOESS regression.
7. Method for producing electricity comprising an electrical production means, a distribution network and electricity consumption means, characterized in that the production means can be controlled by a computer programmed to implement the daily data processing method according to any one of claims 1 to 6.