Skip to main content

Solar forecasts

Background​

Solar generation forecasts let Tensor Cloud run assets automatically: it trades on the electricity markets and submits generation, sales, and procurement plans to the grid operator based on them. For co-located battery storage, the forecasts are also an input to the economically optimized charge and discharge schedules.

Tensor Cloud forecasts solar generation for every eligible asset. The forecasting service runs at fixed times throughout the day, and how far ahead each run forecasts depends on the time of day (see Determining forecast horizon).

Eligible assets​

An asset is included in training and forecasting when all of these apply:

  1. it belongs to a paid workspace
  2. it has at least one solar array
  3. its balancing and trading feature is active, or scheduled to start within the next 14 days (training) or 5 days (forecasts)
  4. its grid interconnection date is no more than 14 days in the future

Forecasts start later than training so that there is time to review forecast quality before an asset starts trading.

Components and architecture​

Like the price forecasts, the solar forecast service has two components:

  1. the training service, which trains an individual forecast model for each asset
  2. the prediction service, which uses the trained models to predict solar generation for each asset

Training service​

The training service runs every day at 23:00 JST. It trains a new model for an asset when at least 7 days of new cleaned generation data have arrived since the last training, or when the asset's plant configuration has changed. If the asset's last model was rejected in validation, a correction to its historical generation data also triggers retraining.

Model architecture​

The models are trained on historical generation data provided by users and on historical weather forecasts. Each model is a multi-output Lasso linear regression, with its regularization strength chosen by cross-validation on the training data. In one pass, a model predicts solar generation for all 32 30-minute slots between 04:00 and 20:00 (16 hours). Generation outside this period is assumed to be zero.

The historical weather forecasts come from several national weather agencies. All assets use the same four weather sources, from which the model reads 20 weather inputs at hourly resolution, plus a day-of-year term for the seasonal cycle. The sources and variables are listed under Weather data.

Training data​

A model is created for an asset once at least 50 days of clean historical generation data are available. Before training, Tensor Cloud cleans the data:

  • Negative values and values above 110% of the asset's site AC capacity are treated as missing.
  • Days with any missing value while the sun is up at the asset's location are excluded. Daylight is calculated from the sun's position, so the window shifts with season and latitude. Missing values at night are set to zero.
  • Days with zero generation in any slot between 10:00 and 14:59 are excluded as likely curtailment.
  • Days on which OCCTO recorded curtailment in the asset's grid area are excluded, even if this asset was not curtailed. This filter is skipped if it would leave fewer than 50 days.

For assets with a battery or a load, only the solar generation data is used, because the meter also measures the battery or the load.

Before training, the generation data is normalized by the plant's expected module degradation. Predictions are scaled back to the degradation expected on the forecast day.

Model validation​

For each asset, we train a candidate model on a fixed combination of weather sources (US GFS, JMA MSM, ECMWF IFS, UKMO Global). The model is then evaluated on a test period made of the most recent 10% of the data, which training never saw. For the same period, we also run a simulation model based on our solar simulation engine. The simulation uses the per-slot mean of ECMWF IFS and Open-Meteo's best-match forecast for the location, falling back to GFS where neither is available.

If the machine learning model does not outperform the simulation, it is rejected and discarded, and the simulation is used for prediction. Otherwise the model is accepted and its RMSE on the test period is recorded. RMSE is calculated over every slot:

RMSE=∑t=1T(gt−gt^)2T\text{RMSE} = \sqrt{ \frac{\sum_{t=1}^T (g_t-\hat{g_t})^2}{T}}

Where TT is the total number of time steps in the validation period, gtg_t is the actual generation at time tt, and gt^\hat{g_t} is the prediction.

Because a model that performs poorly falls back to the simulation, faulty training data is unlikely to degrade the forecast.

Prediction service​

The prediction service runs 10 times a day, at 01:00, 04:00, 05:00, 06:00, 07:00, 10:00, 13:00, 16:00, 19:00, and 22:00 JST. Each run predicts the generation of each asset over a horizon set by the run time (see Determining forecast horizon).

For each asset, the prediction service loads the model created by the training service. It uses the physical simulation instead when:

  • there is no trained model, for example because there is no historical generation data for that asset, or because the asset has just been added to the platform
  • the trained model did not pass validation
  • the plant configuration has changed since the model was trained, until the next training run
  • the model fails to produce a prediction
  • the weather data for a day is incomplete, for that day only

Every predicted slot is then limited to between zero and the asset's capacity. Slots within an event with operational impact on generation are set to zero.

For an asset with a DC-coupled battery in service, the forecast is made on the DC side of the inverter and capped at the solar DC capacity. It therefore includes the energy that the inverter would clip and the battery can charge from. All other assets are forecast on the AC side and capped at the solar AC capacity.

The day-ahead forecast is the latest forecast created by 07:30 JST, normally the one from the 07:00 run.

Determining forecast horizon​

Forecasts reach up to 14 days ahead, but how far a run reaches depends on its time of day. Longer-range forecasts are updated less often: later runs replace them anyway, their weather inputs update less often and are less reliable, and refreshing them at every run would produce excessive data.

Each run forecasts from the run time to the end of the last day it reaches:

Run (JST)Forecast until the end of
07:00day 14
01:00, 13:00, 19:00day 7
05:00day 2
04:00, 06:00, 10:00, 16:00, 22:00day 1 (tomorrow)

Only the 07:00 run produces the full 14 days. As a result, forecasts are refreshed at these intervals:

  • Today and tomorrow: at every run, at most 3 hours apart
  • Day 2: at most 6 hours apart
  • Days 3 to 7: every 6 hours
  • Days 8 to 14: once a day, at 07:00

Weather data​

Sources​

Weather forecasts from several national weather providers are retrieved through the Open-Meteo API, with a Tensor-hosted Open-Meteo instance as a fallback. Training uses Open-Meteo's archive of past forecasts. All assets use the following sources:

National weather providerWeather modelSpatial resolutionTemporal resolutionUpdate frequency
JMAMSM0.05° (~5 km)1 hourEvery 3 hours
NOAA NCEPGFS0.11° (~13 km)1 hourEvery 6 hours
ECMWFIFS 0.25°0.25° (~25 km)3 hours (interpolated to hourly)Every 6 hours
UK Met OfficeUKMO Global0.09° (~10 km)1 hourEvery 6 hours

Each source covers a different horizon. Where a source has no value, including beyond its own forecast range (for example the short range of MSM), the GFS value for the same slot is used.

More sources will be added in the future.

Variables​

The models read these weather variables at hourly resolution, both in training and in prediction:

VariableUnitGFSMSMECMWF IFSUKMO Global
Global horizontal irradiance (GHI)W/m2YesYesYesNo
Diffuse horizontal irradiance (DHI)W/m2YesYesYesNo
Global tilted irradiance (GTI)W/m2YesYesYesNo
Direct normal irradiance (DNI)W/m2YesYesYesYes
Direct solar radiationW/m2YesYesYesYes
Cloud cover%YesNoNoNo
Relative humidity at 2 m%YesNoNoNo
Wind speedm/sYesNoNoNo

Other weather variables, such as air temperature, are not model inputs. The physical simulation uses its own weather inputs, described in Model validation.