Calculate the first order differencing of time series (SPMF documentation)

This example explains how to calculate the first order differencing of time series using the SPMF open-source data mining library.

How to run this example?

What is the calculation of the first order differencing for time series?

Calculating the first order differencing of a time series is useful for converting a non stationary time series to a stationary form. It is calculated as follows. The i-th data point Y_i of a time series is replaced by Y'_i = (Y_i - Y_(i-1). In other words, each point is replaced by the difference between its value and the value of the previous point.

What is the input of this algorithm?

The input is one or more time series. A time series is a sequence of floating-point decimal numbers (double values). A time-series can also have a name (a string).

Time series are used in many applications. An example of time series is the price of a stock on the stock market over time. Another example is a sequence of temperature readings collected using sensors.

For this example, consider the following time series:

Name Data points
ECG1 3,2,8,9,8,9,8,7,6,7,5,4,2,7,9,8,5

This example time series database is provided in the file contextMovingAverage.txt of the SPMF distribution.

In SPMF, to read a time-series file, it is necessary to indicate the "separator", which is the character used to separate data points in the input file. In this example, the "separator" is the comma ',' symbol.

What is the output?

The output is the first order differencing of the time series received as input. It is calculated as follows. The i-th data point Y_i of a time series is replaced by Y'_i = (Y_i - Y_(i-1). In other words, each point is replaced by the difference between its value and the value of the previous point.

For example, in the above example, the result is:

Name Data points
ECG1_FODIFF -1.0,6.0,1.0,-1.0,1.0,-1.0,-1.0,-1.0,1.0,-2.0,-1.0,-2.0,5.0,2.0,-1.0,-3.0

To see the result visually, it is possible to use the SPMF time series viewer, described in another example of this documentation. In the following figure, the original time series is displayed (top) and its first order differencing (bottom).

Input file format

The input file format is defined as follows. It is a text file. The text file contains one or more time series. Each time series is represented by two lines. The first line of the file contains optional metadata @FILETYPE="Time series database" indicating that this is a timeseries database file.
The second line contains optional metadata about the source of the file.
The third line contains the string "@NAME=" followed by the name of the time series. The fourth line is a list of data points, where data points are floating-point decimal numbers separated by a separator character (here the ',' symbol).
Then, the following lines may contain additional time series described in the same way.

For example, the input file of the previous example, named contextMovingAverage.txt is defined as follows:

@FILETYPE="Time series database"
@SOURCE="SPMF SOFTWARE https://philippe-fournier-viger.com/spmf/"
@NAME=ECG2
3,2,8,9,8,9,8,7,6,7,5,4,2,7,9,8,5

Consider the last two lines. It indicates that the first time series name is "ECG2" and that it consits of the data points: 3,2,8,9,8,9,8,7,6,7,5,4,2,7,9,8, and 5.

But note that it is possible to have more than one time series per file. For example, this is another input file called contextSax.txt, which contains 4 time series.

@FILETYPE="Time series database"
@SOURCE="SPMF SOFTWARE https://philippe-fournier-viger.com/spmf/"
@NAME=ECG1
1,2,3,4,5,6,7,8,9,10
@NAME=ECG2
1.5,2.5,10,9,8,7,6,5
@NAME=ECG3
-1,-2,-3,-4,-5
@NAME=ECG4
-2.0,-3.0,-4.0,-5.0,-6.0

Output file format

The output file format is the same as the input format. For example, there is the result of this example:

@NAME=ECG2_FODIFF
-1.0,6.0,1.0,-1.0,1.0,-1.0,-1.0,-1.0,1.0,-2.0,-1.0,-2.0,5.0,2.0,-1.0,-3.0

Implementation details

Beside the first order differencing, SPMF also offers second-order differencing and various other operations for time series.

Where can I get more information about the first order differencing?

The first order differencing is a basic operation for analyzing time series. It is described in many websites and books.