Support Vector Regression

From Coastal Wiki
Revision as of 14:46, 3 October 2026 by Dronkers J (talk | contribs)
Jump to: navigation, search
Definition of Support vector regression:
Support vector regression (SVR) is a regression method for fitting a curve to data in which the user chooses the type of curve; among all curves of that type, it finds the smoothest one while penalizing measurements that deviate from the curve by more than a chosen tolerance.
This is the common definition for Support vector regression, other definitions can be discussed in the article


Explanation

Glossary
Predictor variables: The variables from which the prediction is made, for example wind speed.
Target variable: The variable to be predicted, for example wave height.
Training data: A set of simultaneous measurements of the predictor variables and the target variable, for example wind speed and wave height measured at the same time. The curve is fitted to these data.
Standardization: Rescaling a predictor variable by subtracting its mean and dividing by its standard deviation.
Tolerance band: A band of chosen width above and below the fitted curve. Deviations of measurements inside the band are ignored.
Excess deviation: The distance by which a measurement lies outside the edge of the tolerance band.
Kernel function: The formula that gives the shape of the pieces from which the curve is built. For bell-shaped pieces (Gaussian kernel), it converts the difference between two predictor values into a number between 1 (no difference) and 0 (large difference).
Kernel width: The width of the bell-shaped pieces, chosen by the user, measured in units of the predictor variables.
Roughness: A measure of the complexity of the fitted curve, expressing how strongly it departs from a horizontal relationship by rising, falling and bending.
Penalty setting [math]C[/math]: A setting chosen by the user that determines how heavily excess deviations count compared with roughness.
Support vectors: The measurements on the edge of the tolerance band or outside it. Only these determine the fitted curve.
Overfitting: Following the training data so closely that random fluctuations and measurement errors are reproduced, making predictions for new situations unreliable.
Cross-validation: Testing a method repeatedly on parts of the data not used when fitting it.

Good predictions require a balance: the curve should fit the training data, but should avoid overfitting. SVR expresses this balance in a single score, which combines the deviations of the measurements from the curve with the roughness of the curve (the term roughness is used here for the regularization term that measures the complexity of the fitted function)[1]. The score is determined by:

  • Flexibility. This requires the curve to be smooth, without limiting the number of fitting parameters. A curve can be built from many pieces and still be smooth.
  • Tolerance band. Small deviations may reflect measurement noise rather than the true relationship. They can therefore be ignored by introducing a tolerance band around the curve.

Below an explanation of the support vector regression method is given by comparison with the ordinary least-squares fitting method.

  • SVR and the ordinary least-squares method are both based on parameter fitting to minimize a cost function. The cost functions are different, however.
  • In least squares, the user specifies the form of the fitting curve, for example a straight line or a polynomial of a chosen order. In SVR, the user specifies a kernel function, which transforms the data into a space with a large (possibly even infinite) number of dimensions; each dimension corresponds to a function of the predictor variables that follows from the chosen kernel. SVR fits a flat surface in this transformed space. Seen back in the original predictor variables, this flat surface becomes a curve that can take almost any shape.
  • The least-squares cost function increases with the squared deviation between fitted and observed values. The SVR cost function ignores training data that lie within a tolerance band around the fitted curve. For data outside the band, it increases with the excess deviation: the distance between the observed value and the nearest edge of the band. The width of the band is chosen by the user, usually in relation to the expected noise level of the training data.
  • The least-squares curve is kept smooth by limiting the number of fitting parameters. In SVR, smoothness is promoted by adding to the cost function a term that penalizes the complexity ('roughness') of the fitted function. This term is the sum of the squared slopes of the flat surface in the transformed space. For a straight-line fit without transformation, it is simply the sum of the squared slopes with respect to each predictor variable.


Analysis technique Strengths Limitations Application example
Curve fitting based on machine learning from training data * Handles nonlinear relationships
* Works well with small to medium-sized datasets
* Less sensitive to outliers than least squares fitting
* The optimization has no inferior local minima
* Predictions depend only on the support vectors
* The kernel function and its settings, together with the tolerance and penalty settings, must be chosen
* Computing time increases rapidly with the size of the dataset
* No uncertainty range for predictions; known measurement errors can only be used as weights
* Does not directly show the importance of predictor variables
* Poor at extrapolation beyond the range of the training data
Prediction of environmental variables from a limited number of measurements, examples
*estimating water depth, chlorophyll or turbidity from satellite data
*predicting wave height or water level at a coastal station


Appendix: The Gaussian Kernel SVR

In the case of a Gaussian kernel the SVR fitting curve is built from simple pieces, each representing a bell-shaped bump. A bump is associated with each of the [math]N[/math] training measurements. Each bump [math]i[/math] has:

  • a height [math]a_i[/math], which can be positive (a hump) or negative (a dip) and which is determined by SVR; [math]a_i[/math] is nonzero only for training measurements on or outside the tolerance band (the support vectors);
  • a width [math]\sigma[/math], the kernel width, which is the same for all bumps and is chosen by the user.

The curve is a horizontal line at a constant level plus the sum of all Gaussian functions. A Gaussian bump of zero height has no effect on the curve. With narrow bumps, the curve can follow small details in the data; with wide bumps, the curve can only change gradually. The kernel function converts the difference between predictor values into a number between 0 and 1. The number is 1 when the predictor values are equal and decreases towards 0 as the difference increases. For a single predictor variable, the fitting curve can be written as

[math]f(x)=\text{constant} + \sum_{i=1}^N a_i \, K(x,x_i) \, , \quad K(x,x_i)=\exp\Big(- \dfrac{(x-x_i)^2}{2 \sigma^2}\Big) \, ,[/math]

where [math]x_i[/math] is the predictor value of training measurement [math]i[/math] and the other symbols are defined above.

With several predictor variables, the Gaussian bumps become hills in a space with one axis for each predictor variable. The separation between two observations is then measured by their distance in this space. To prevent variables with large numerical scales from dominating this distance, predictor variables are generally standardized: from each variable its mean is subtracted, and the result is divided by its standard deviation.

Note that the kernel width and the width of the tolerance band are different settings. The tolerance band is measured vertically, in units of the target variable. The kernel width is measured horizontally, in units of the predictor variables, and determines how quickly the predicted value can change when the predictor values change.

The roughness is calculated from the bump heights. Each bump is paired with every bump, including itself; the two heights are multiplied by each other and by the kernel value for the distance between their predictor values, and all these products are added. Each bump is also paired with itself. In mathematical terms,

[math]\text{roughness} = \sum_{i=1}^N \sum_{j=1}^N a_i \, a_j \, K(x_i,x_j) .[/math]

Two nearby bumps of the same sign, which together form a large hump, increase the roughness. Two nearby bumps of opposite sign partly cancel each other and reduce it. The level of the horizontal line does not affect the roughness.

How SVR finds the smoothest curve

SVR gives every possible curve a score:

[math]\text{score} = \text{roughness} + C \times \text{sum of the excess deviations of all measurements} = \text{roughness} + C \, \sum_{i=1}^N \text{max} \big(0, |y_i-f(x_i)|-\epsilon \big) \, , [/math]

where [math]C[/math] is the penalty setting chosen by the user, [math]y_i[/math] is the observed target value, [math]f(x_i)[/math] is the fitted value and [math]\epsilon[/math] is the half-width of the tolerance band.

The curve with the lowest score is selected. With a large value of [math]C [/math], excess deviations are expensive, so the curve follows the measurements closely, with the risk of overfitting. With a small value of [math]C [/math], the curve stays smoother and more measurements are allowed outside the band.

The search starts with all bumps at zero height, so the curve is a horizontal line with zero roughness. If, at some level of this line, all measurements lie inside the tolerance band, this horizontal line is the answer: the data show no relationship larger than the tolerance. Otherwise, SVR adjusts the bump heights and the level of the line step by step.

Every measurement outside the tolerance band tends to pull the curve towards itself, corresponding to a positive or negative contribution of the Gaussian function centered on that measurement. Because excess deviations are not squared, each of these measurements pulls with the same strength, set by [math]C[/math], however far outside the band it lies. The roughness pulls the other way, towards a horizontal line. Measurements inside the band do not pull at all. The bump heights are adjusted until the pulls are in balance, so that any further change would increase the score.

At the balance point there are three kinds of measurements:

  • Measurements inside the band have a bump of zero height. They do not contribute to the curve.
  • Measurements exactly on the edge of the band pull just enough to hold the curve in place. Their bumps have an absolute height between zero and the maximum set by [math]C[/math].
  • Measurements outside the band all pull with full strength. Their bumps have the maximum absolute height set by [math]C[/math].

The measurements with a bump of nonzero height are called support vectors (in the mathematical description, each measurement is written as a list of numbers, called a vector). Only the support vectors determine the curve, and only these are needed to calculate predictions.

Because the pull of a measurement outside the band does not increase with its distance from the band, SVR is less sensitive to a strongly deviating measurement than least-squares fitting. In least squares fitting, by contrast, the pull of a measurement increases with its distance from the curve.

Why the search always finds the best curve

The score can be pictured as a landscape, with the bump heights and the level of the line as the position, and the score as the elevation. SVR searches for the lowest point of this landscape. The parts of the score each have a simple shape:

  • The roughness is shaped like a bowl: it is zero when all bumps have zero height and increases in every direction away from that point.
  • The excess deviation of each measurement is shaped like a trough with a flat bottom: it is zero as long as the curve passes within the tolerance band at that measurement, and increases steadily as the curve moves further away from the band, on either side.

Adding bowls and troughs together produces a convex landscape: it has no separate local dips in which the search can become trapped. Any minimum reached is therefore a global minimum. This is an important difference from neural networks, whose optimization landscape can contain many local minima.

Choosing the settings and testing the result

SVR has three settings: the width of the tolerance band, the penalty setting [math]C[/math] and the kernel width. The band width is often based on the measurement error of the target variable. [math]C[/math] and the kernel width are usually chosen by cross-validation: the training data are divided into several parts; the curve is fitted to all parts but one and tested on the part left out; this is repeated for each part. The settings giving the smallest prediction errors on the left-out parts are selected.

When measurements are close together in space or time, as is common with coastal and remote-sensing data, the left-out parts should consist of other locations or periods than the data used for fitting. Otherwise the method appears more reliable than it is [2]. A final test should be made with independent data that were not used in any way for fitting the curve or choosing the settings.


More detailed explanations

StatQuest: Support Vector Machines Part 1: Main Ideas by Josh Starmer
Wikipedia Support vector machine

Related articles

Data analysis techniques for the coastal zone
Random Forest Regression

References

  1. ↑ Smola, A.J. and Schölkopf, B. 2004. A tutorial on support vector regression. Statistics and Computing 14: 199–222
  2. ↑ Roberts, D.R. et al. 2017. Cross-validation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure. Ecography 40: 913–929