<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.3 20210610//EN" "https://jats.nlm.nih.gov/publishing/1.3/JATS-journalpublishing1-3.dtd"><article xml:lang="en" dtd-version="1.3" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:ali="http://www.niso.org/schemas/ali/1.0/" article-type="research-article"><front><journal-meta><journal-id journal-id-type="issn">2583-6250</journal-id><journal-title-group><journal-title>International Journal of Data Informatics and Intelligent Computing </journal-title><abbrev-journal-title>International Journal of Data Informatics and Intelligent Computing </abbrev-journal-title></journal-title-group><issn pub-type="epub">2583-6250</issn><publisher><publisher-name>Prisma Publications</publisher-name><publisher-loc>India</publisher-loc></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.59461/ijdiic.v5i3.305</article-id><article-categories><subj-group><subject>Artificial Intelligence</subject></subj-group></article-categories><title-group><article-title>Optimized Machine Learning for Solar Energy Generation Forecasting: A Feature Selection and Hyperparameter Optimization Approach</article-title></title-group><contrib-group><contrib contrib-type="author"><contrib-id contrib-id-type="orcid">0009-0005-2383-3392</contrib-id><name><surname>Nawaz</surname><given-names>Farwa</given-names></name><xref ref-type="aff" rid="AFF-1"></xref><xref ref-type="corresp" rid="cor-0"></xref></contrib><contrib contrib-type="author"><contrib-id contrib-id-type="orcid">0009-0000-2326-2175</contrib-id><name><surname>Karamat</surname><given-names>Faria</given-names></name><xref ref-type="aff" rid="AFF-1"></xref></contrib></contrib-group><aff id="AFF-1"><institution content-type="dept">Faculty of Computer Science</institution><institution-wrap><institution>Riphah International University</institution><institution-id institution-id-type="ror">https://ror.org/02kdm5630</institution-id></institution-wrap><addr-line>Islamabad</addr-line><country country="PK">Pakistan</country></aff><author-notes><fn fn-type="coi-statement"><label>Conflict of Interest</label><p>The authors declare that they have no conflicts of interest to this work.</p></fn><corresp id="cor-0">Corresponding author: Farwa Nawaz, Faculty of Computer Science, Riphah International University, Islamabad, Pakistan. </corresp></author-notes><pub-date date-type="pub" iso-8601-date="2026-9-12" publication-format="electronic"><day>12</day><month>9</month><year>2026</year></pub-date><pub-date date-type="collection" iso-8601-date="2026-9-25" publication-format="electronic"><day>25</day><month>9</month><year>2026</year></pub-date><volume>3</volume><issue>5</issue><fpage>50</fpage><lpage>66</lpage><history><date date-type="received" iso-8601-date="2026-7-13"><day>13</day><month>7</month><year>2026</year></date><date date-type="rev-recd" iso-8601-date="2026-8-28"><day>28</day><month>8</month><year>2026</year></date><date date-type="accepted" iso-8601-date="2026-9-6"><day>6</day><month>9</month><year>2026</year></date></history><permissions><copyright-statement>© 2026 Farwa Nawaz, Faria Karamat</copyright-statement><copyright-year>2026</copyright-year><copyright-holder>The Author(s).</copyright-holder><license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-sa/4.0/"><ali:license_ref xmlns:ali="http://www.niso.org/schemas/ali/1.0/">https://creativecommons.org/licenses/by-sa/4.0/</ali:license_ref><license-p>This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.</license-p></license></permissions><self-uri xlink:href="https://doi.org/10.59461/ijdiic.v5i3.305" xlink:title="Optimized Machine Learning for Solar Energy Generation Forecasting: A Feature Selection and Hyperparameter Optimization Approach">Optimized Machine Learning for Solar Energy Generation Forecasting: A Feature Selection and Hyperparameter Optimization Approach</self-uri><abstract><p>Solar energy is among the most promising renewable resources for developing sustainable energy systems; however, the stochastic and weathersensitive nature of Solar Energy Generation (SEG) poses challenges for forecasting. The current study designs a machine learning framework for SEG forecasting by using feature selection methods in three stages. In the first stage, three feature selection methods, namely Pearson Correlation (PC), Recursive Feature Elimination (RFE), and Random Forest Feature Importance (RFFI), are compared systematically, each method selecting four predictors among ten weather and system variables, while avoiding the use of electrical telemetry variables to prevent data leakage. In the second stage, four regression methods, including Linear Regression, Random Forest (RF), Multilayer Perceptron (MLP), and Optimized Radial Basis Function Support Vector Regression (RBF-SVR), are tested using each feature set, and the results are verified in five independently repeated experiments. The feature selection method has a much bigger impact on accuracy than the regression method: RFFI-selected features allowed attaining a mean R² greater than 0.85 for three regression methods, whereas PC and RFE constrained all models to achieve R² less than 0.60. Among all models for RFFI features, Random Forest proved to have the best accuracy (mean R² = 0.8960 ± 0.0008), with MLP being second (R² = 0.8893 ± 0.0022), followed by the Optimized RBF-SVR model (R² = 0.8568 ± 0.0072); the Optimized RBF-SVR model took an average training time of 21.25 ± 1.12 seconds compared to RF's 35.78 ± 2.42 seconds. This shows the importance of selecting leakage-controlled features systematically in the process of forecasting SEG, and how a small number of 4 features is sufficient for achieving statistical stability.</p></abstract><kwd-group><kwd>Solar Energy Forecasting</kwd><kwd>Support Vector Regression</kwd><kwd>Renewable Energy</kwd><kwd>Machine Learning</kwd><kwd>Feature Selection</kwd><kwd>Smart Grid</kwd></kwd-group><custom-meta-group><custom-meta><meta-name>File created by JATS Editor</meta-name><meta-value><ext-link xlink:href="https://jatseditor.com" xlink:title="JATS Editor" ext-link-type="uri">JATS Editor</ext-link></meta-value></custom-meta><custom-meta><meta-name>issue-created-year</meta-name><meta-value>2026</meta-value></custom-meta></custom-meta-group></article-meta></front><body><sec><title>1. INTRODUCTION</title><p>Renewable energy is crucial for achieving sustainable and low-carbon energy systems worldwide <xref ref-type="bibr" rid="BIBR-1">[1]</xref>. Solar photovoltaics continue to be one of the fastest-growing renewable resources <xref ref-type="bibr" rid="BIBR-2">[2]</xref>, with an expected 507 GW of renewable energy generated in 2023, a 50% increase from 2022 <xref rid="BIBR-3" ref-type="bibr">[3]</xref>. Recent developments have expedited the integration of solar systems into contemporary power grids, and their growing deployment aids international efforts toward carbon neutrality and energy security <xref ref-type="bibr" rid="BIBR-4">[4]</xref>. Clean, plentiful, and ecologically friendly electricity may be produced with solar energy <xref ref-type="bibr" rid="BIBR-5">[5]</xref><xref ref-type="bibr" rid="BIBR-6">[6]</xref>. Nevertheless, weather fluctuations continue to have a significant impact on solar energy production (SEG) <xref ref-type="bibr" rid="BIBR-7">[7]</xref>. Photovoltaic output is directly impacted by variations in temperature, humidity, wind speed, cloud cover, and solar irradiation. These variations also complicate grid management and lower forecasting accuracy <xref rid="BIBR-8" ref-type="bibr">[8]</xref>. Thus, reliable energy scheduling and smart grid functioning depend on precise SEG predictions <xref rid="BIBR-9" ref-type="bibr">[9]</xref><xref ref-type="bibr" rid="BIBR-10">[10]</xref>.</p><p>For SEG prediction, a number of forecasting methods have been developed. Initially, statistical models such as autoregressive, moving average, and time-series forecasting techniques were used due to their ease of use and low processing requirements <xref rid="BIBR-11" ref-type="bibr">[11]</xref><xref ref-type="bibr" rid="BIBR-12">[12]</xref>. However, extremely nonlinear correlations found in meteorological data are difficult to model using statistical methods. Later, heuristic and optimization-based approaches were developed to enhance forecasting performance <xref rid="BIBR-13" ref-type="bibr">[13]</xref><xref rid="BIBR-14" ref-type="bibr">[14]</xref>. However, these approaches frequently need significant processing resources and parameter tuning, which limits their usefulness for real-time forecasting systems.</p><p>Recently, machine learning (ML) has become a viable substitute for SEG forecasting, successfully capturing nonlinear correlations between solar power generation and meteorological variables <xref ref-type="bibr" rid="BIBR-4">[4]</xref>. Promising forecasting performance has been shown by a number of regression algorithms, such as Random Forest, Support Vector Regression, and Neural Networks; deep learning models have further increased prediction accuracy by utilizing intricate hierarchical feature representations <xref ref-type="bibr" rid="BIBR-15">[15]</xref><xref ref-type="bibr" rid="BIBR-16">[16]</xref>. Despite these developments, the quality of the input features has a significant impact on prediction accuracy. Selecting informative features continues to be a crucial difficulty for reliable SEG forecasting since redundant and irrelevant variables frequently increase computational complexity and degrade forecasting performance <xref ref-type="bibr" rid="BIBR-17">[17]</xref>.</p><p>To achieve high prediction performance, model generalization, and computational efficiency, feature selection (FS) techniques reduce data dimensionality without losing vital predictive data <xref ref-type="bibr" rid="BIBR-18">[18]</xref><xref ref-type="bibr" rid="BIBR-19">[19]</xref>. Popular FS methods include Random Forest Feature Importance, Pearson Correlation, and Recursive Feature Elimination <xref ref-type="bibr" rid="BIBR-17">[17]</xref><xref ref-type="bibr" rid="BIBR-20">[20]</xref>. Nevertheless, few studies have thoroughly examined different methods for SEG forecasting under the same circumstances, and earlier research frequently prioritizes prediction accuracy at the expense of computational efficiency <xref ref-type="bibr" rid="BIBR-21">[21]</xref>. Seldom are advanced evaluation criteria like Huber Loss, Explained Variance Score, and Median Absolute Error taken into account. Thus, an effective forecasting framework that combines reliable machine learning models with ideal feature selection is still required.</p><p>This paper suggests an enhanced machine learning approach for SEG forecasting in order to overcome these constraints. The most instructive meteorological features are first determined using three feature selection methods. Second, the same experimental conditions are used to assess four machine learning regression models. Third, to increase forecasting precision and computational efficiency, an improved Radial Basis Function Support Vector Regression (RBF-SVR) model is created. The publicly accessible Desert Knowledge Australia Solar Centre dataset is used in extensive tests, and the experimental findings show better forecasting performance than current machine learning techniques.</p><p>The following is a summary of this study's primary contributions:</p><list list-type="bullet"><list-item><p>Three feature selection methods are integrated with machine learning regression models to create a comprehensive SEG forecasting system.</p></list-item><list-item><p>For the best feature selection under the same experimental settings, Pearson Correlation, Recursive Feature Elimination, and Random Forest Feature Importance are compared.</p></list-item><list-item><p>The proposed optimized model of RBF-SVR with optimal hyperparameters is achieved, which provides high prediction accuracy but has significantly reduced computational complexity compared to the best baseline model.</p></list-item><list-item><p>Experiments, with leakage prevention measures, are performed for the comparison of linear regression, random forest, MLP, and the proposed optimized RBF-SVR, using three feature selection methods, verified by conducting the experiment five times independently with standard deviation provided.</p></list-item><list-item><p>The proposed method is tested using seven regression performance measures – MSE, RMSE, MAE, MedAE, EVS, Huber loss, and R², proving that the feature selection method impacts the forecasting accuracy more than model selection does.</p></list-item></list><p>The remainder of this paper is organized as follows. Section 2 reviews recent SEG forecasting approaches. Section 3 presents the proposed forecasting framework, dataset, preprocessing, feature selection, machine learning models, experimental setup, and evaluation metrics. Section 4 presents the experimental results and comparative analysis. Finally, Section 5 concludes the paper and outlines future research directions.</p></sec><sec><title>2. LITERATURE REVIEW</title><sec><title>2.1. Machine Learning and Optimization-Based SEG Forecasting</title><p>SEG forecasting has attracted significant research efforts owing to the rise of PV systems and smart grid installations. The application of various machine learning techniques such as Support Vector Regression (SVR), Random Forest (RF), Artificial Neural Networks (ANN), Gaussian Process Regression (GPR), Decision Tree (DT), and Gradient Boosting for forecasting has shown great success in accurately modeling the nonlinear dependency between meteorological features and power production <xref ref-type="bibr" rid="BIBR-22">[22]</xref><xref rid="BIBR-23" ref-type="bibr">[23]</xref><xref ref-type="bibr" rid="BIBR-17">[17]</xref><xref ref-type="bibr" rid="BIBR-4">[4]</xref>. Particularly, SVR is still one of the most popular forecasting methods due to its good generalization ability, whereas RF and ANN yield high forecasting accuracy regardless of climate <xref ref-type="bibr" rid="BIBR-17">[17]</xref>.</p><p>Machine learning models and optimization algorithms have been combined to increase prediction accuracy. To optimize model hyperparameters and lower forecasting errors, methods like Particle Swarm Optimization (PSO) and a Hybrid Improved Multi-Verse Optimizer (HIMVO) <xref ref-type="bibr" rid="BIBR-22">[22]</xref>, Genetic Algorithms (GA/GASVM) <xref ref-type="bibr" rid="BIBR-23">[23]</xref>, Ant Colony Optimization (ACO) <xref ref-type="bibr" rid="BIBR-24">[24]</xref>, the Whale Optimization Algorithm (WOA) <xref ref-type="bibr" rid="BIBR-25">[25]</xref>, and Bayesian Optimization <xref ref-type="bibr" rid="BIBR-21">[21]</xref> have been used. Despite the fact that these hybrid approaches increase prediction accuracy, most research ignores the impact of feature selection on forecasting efficiency in favor of parameter tuning.</p></sec><sec><title>2.2. Deep Learning and Feature Selection</title><p>Deep learning models, comprising Long Short-Term Memory (LSTM) networks <xref ref-type="bibr" rid="BIBR-26">[26]</xref>, Gated Recurrent Units (GRU) <xref ref-type="bibr" rid="BIBR-27">[27]</xref>, Convolutional Neural Networks (CNN) <xref ref-type="bibr" rid="BIBR-28">[28]</xref>, and Transformer architectures <xref ref-type="bibr" rid="BIBR-29">[29]</xref>, have recently demonstrated excellent performance in solar forecasting by effectively learning temporal and nonlinear patterns from large datasets. However, these models' application in computationally constrained situations is limited since they often demand large amounts of training data, higher computational resources, and longer training times <xref ref-type="bibr" rid="BIBR-30">[30]</xref>.</p><p>An efficient preprocessing method for increasing forecasting efficiency is feature selection. Data dimensionality is decreased while model generalization is enhanced by techniques including Pearson Correlation (PC) <xref ref-type="bibr" rid="BIBR-17">[17]</xref>, Recursive Feature Elimination (RFE) <xref ref-type="bibr" rid="BIBR-31">[31]</xref>, Random Forest Feature Importance (RFFI) <xref ref-type="bibr" rid="BIBR-19">[19]</xref>, Mutual Information <xref rid="BIBR-32" ref-type="bibr">[32]</xref>, and wrapper-based approaches <xref ref-type="bibr" rid="BIBR-20">[20]</xref>. However, current research usually assesses a single feature selection method or combines it with a small number of forecasting models, offering little information on the relative efficacy of several feature selection approaches.</p></sec><sec><title>2.3. Research Gap and Motivation</title><p>SEG forecasting has come a long way, but there are still a number of obstacles. The majority of current research focuses on optimizing models or hyperparameters while paying little attention to systematic feature selection. There are also few comparative analyses of various feature selection methods under the same experimental conditions. Additionally, traditional criteria are frequently used to evaluate predictive ability, while advanced regression measures are rarely taken into account. These constraints drive the creation of computationally effective forecasting frameworks that use small feature subsets to retain high predicted accuracy.</p><p>This study suggests a feature-selection-driven forecasting paradigm that compares Random Forest Feature Importance (RFFI), Recursive Feature Elimination (RFE), and Pearson Correlation (PC) in order to address these issues. An optimal Radial Basis Function Support Vector Regression (RBF-SVR) model is created after four machine learning regression models are trained with similar feature subsets. Both traditional and sophisticated regression performance indicators are used to thoroughly assess the suggested architecture. The current work is positioned in relation to representative recent investigations, which are summarized in Table <xref ref-type="table" rid="table-1">1</xref>.</p><table-wrap id="table-1" ignoredToc=""><label>Table 1</label><caption><p>Summary of recent SEG forecasting studies</p></caption><table frame="box" rules="all"><thead><tr><th align="left" colspan="1" valign="top"><bold>Reference</bold></th><th align="left" colspan="1" valign="top"><bold>Method</bold></th><th colspan="1" valign="top" align="left"><bold>Feature Selection</bold></th><th valign="top" align="left" colspan="1"><bold>Limitation</bold></th></tr></thead><tbody><tr><td valign="top" align="left" colspan="1">33[]</td><td valign="top" align="left" colspan="1">RF (ensemble component)</td><td valign="top" align="left" colspan="1">No</td><td align="left" colspan="1" valign="top">Large feature set</td></tr><tr><td valign="top" align="left" colspan="1">34[]</td><td colspan="1" valign="top" align="left">SVR</td><td align="left" colspan="1" valign="top">No</td><td valign="top" align="left" colspan="1">Hyperparameter sensitive</td></tr><tr><td align="left" colspan="1" valign="top">22[]</td><td align="left" colspan="1" valign="top">PSO-SVR / HIMVO-SVM</td><td valign="top" align="left" colspan="1">No</td><td align="left" colspan="1" valign="top">No feature selection</td></tr><tr><td align="left" colspan="1" valign="top">24[]</td><td valign="top" align="left" colspan="1">ACO-SVR</td><td colspan="1" valign="top" align="left">No</td><td valign="top" align="left" colspan="1">Computationally expensive</td></tr><tr><td valign="top" align="left" colspan="1">26[]</td><td valign="top" align="left" colspan="1">LSTM (WPD-LSTM)</td><td colspan="1" valign="top" align="left">No</td><td align="left" colspan="1" valign="top">Long training time</td></tr><tr><td colspan="1" valign="top" align="left">27[]</td><td colspan="1" valign="top" align="left">GRU (TCN-ECANet-GRU)</td><td valign="top" align="left" colspan="1">No</td><td align="left" colspan="1" valign="top">Large data requirement</td></tr><tr><td valign="top" align="left" colspan="1">17[]</td><td valign="top" align="left" colspan="1">Pearson Correlation + RF</td><td valign="top" align="left" colspan="1">Pearson Correlation</td><td valign="top" align="left" colspan="1">Single feature selection technique</td></tr><tr><td align="left" colspan="1" valign="top">20[]</td><td align="left" colspan="1" valign="top">Wrapper + SVR</td><td valign="top" align="left" colspan="1">Wrapper</td><td valign="top" align="left" colspan="1">Higher feature dimensionality</td></tr></tbody></table></table-wrap></sec></sec><sec><title>3. METHOD</title><p>The general structure of the suggested Solar Energy Generation (SEG) forecasting system is shown in Figure  <xref ref-type="fig" rid="figure-6">1</xref>. Five steps make up the framework: (1) data collection and preprocessing; (2) feature selection; (3) machine learning model building; (4) model optimization; and (5) performance evaluation. To enhance data quality, raw meteorological and electricity-generating data are first cleaned and preprocessed. The most informative predictors of solar energy generation are then found using three feature selection strategies. Four popular machine learning regression models are then trained using the chosen feature subsets. Lastly, a Radial Basis Function (RBF) Support Vector Regression (SVR) framework is used to further refine the top-performing model, and different regression metrics are used to assess the model's performance.</p><fig id="figure-6" ignoredToc=""><label>Figure 1</label><graphic mime-subtype="png" mimetype="image" xlink:href="https://ijdiic.com/research/article/download/305/version/306/228/2212/International_Journal_of_Data_Informatics_and_Intelligent_Computing-3-5-50-g1.png"><alt-text>Image</alt-text></graphic></fig><p>The suggested framework focuses on identifying a small subset of useful meteorological characteristics, in contrast to traditional forecasting methods that depend on the entire feature set. This method preserves good predictive accuracy while lowering computational complexity. To determine the best forecasting framework, a thorough comparison of feature selection methods and machine learning models is carried out under identical experimental conditions.</p><sec><title>3.1. Dataset Description</title><p>The experiments were performed based on the photovoltaic data set from the Desert Knowledge Australia Solar Centre (DKASC), which is an open-source data set available for research use. This data set includes actual measurements from Site 28 at the DKASC Alice Springs site, which is a First Solar CdTe (thin-film) solar power plant. Table <xref ref-type="table" rid="table-2">2</xref> presents the specifications of Site 28.</p><table-wrap id="table-2" ignoredToc=""><label>Table 2</label><caption><p>Details of the solar panel installation</p></caption><table frame="box" rules="all"><thead><tr><th valign="top" align="left" colspan="1"><bold>Parameter</bold></th><th valign="top" align="left" colspan="1"><bold>Value</bold></th></tr></thead><tbody><tr><td valign="top" align="left" colspan="1">Manufacturer</td><td colspan="1" valign="top" align="left">First Solar</td></tr><tr><td align="left" colspan="1" valign="top">Array Rating</td><td valign="top" align="left" colspan="1">5.6 kW</td></tr><tr><td colspan="1" valign="top" align="left">PV Technology</td><td colspan="1" valign="top" align="left">CdTe (thin film)</td></tr><tr><td valign="top" align="left" colspan="1">Array Structure</td><td valign="top" align="left" colspan="1">Fixed: Ground Mount</td></tr><tr><td colspan="1" valign="top" align="left">Panel Rating</td><td valign="top" align="left" colspan="1">87.5 W</td></tr><tr><td valign="top" align="left" colspan="1">Number of Panels</td><td valign="top" align="left" colspan="1">64</td></tr><tr><td colspan="1" valign="top" align="left">Panel Type</td><td valign="top" align="left" colspan="1">First Solar FS-387</td></tr><tr><td colspan="1" valign="top" align="left">Array Area</td><td valign="top" align="left" colspan="1">46.08 m²</td></tr><tr><td valign="top" align="left" colspan="1">Inverter Size / Type</td><td valign="top" align="left" colspan="1">6 kW, SMA SMC 6000</td></tr><tr><td valign="top" align="left" colspan="1">Installation Completed</td><td valign="top" align="left" colspan="1">February 06 2013</td></tr><tr><td valign="top" align="left" colspan="1">Array Tilt / Azimuth</td><td colspan="1" valign="top" align="left">Tilt = 20°, Azimuth = 0° (Solar North)</td></tr></tbody></table></table-wrap><p>The dataset includes 13 predictor variables, a target variable (Active Power), and a timestamp field, recorded at an interval of five minutes over several years, resulting in 1,281,034 total observations after cleaning. However, two out of the 13 predictor variables recorded (Current Phase Average and Active Energy Delivered/Received) have been excluded from the candidate feature set because they represent electrical equivalents of the target variable and thus do not qualify as predictor variables, nor will they be available at the time of forecasting. The ten remaining candidates are listed in Table <xref ref-type="table" rid="table-3">3</xref> below.</p><table-wrap id="table-3" ignoredToc=""><label>Table 3</label><caption><p>Features in the dataset</p></caption><table frame="box" rules="all"><thead><tr><th align="left" colspan="1" valign="top"><bold>No</bold></th><th align="left" colspan="1" valign="top"><bold>Feature</bold></th><th align="left" colspan="1" valign="top"><bold>Description</bold></th></tr></thead><tbody><tr><td valign="top" align="left" colspan="1">1</td><td align="left" colspan="1" valign="top">Active Power</td><td valign="top" align="left" colspan="1">Solar energy being generated at a given time (target variable)</td></tr><tr><td align="left" colspan="1" valign="top">2</td><td align="left" colspan="1" valign="top">Active Energy Delivered/Received</td><td valign="top" align="left" colspan="1">Total energy delivered to or received from the grid per unit time [Excluded from candidates — leakage risk]</td></tr><tr><td align="left" colspan="1" valign="top">3</td><td valign="top" align="left" colspan="1">Current Phase Average</td><td colspan="1" valign="top" align="left">Average current flowing through a particular phase [Excluded from candidates — leakage risk]</td></tr><tr><td valign="top" align="left" colspan="1">4</td><td colspan="1" valign="top" align="left">Active Power Performance Ratio</td><td align="left" colspan="1" valign="top">Measure of efficiency of the solar power system</td></tr><tr><td valign="top" align="left" colspan="1">5</td><td align="left" colspan="1" valign="top">Wind Speed</td><td align="left" colspan="1" valign="top">The speed of wind per unit time</td></tr><tr><td valign="top" align="left" colspan="1">6</td><td align="left" colspan="1" valign="top">Weather Temperature (Celsius)</td><td valign="top" align="left" colspan="1">The ambient temperature per unit time</td></tr><tr><td valign="top" align="left" colspan="1">7</td><td valign="top" align="left" colspan="1">Weather Relative Humidity</td><td valign="top" align="left" colspan="1">The amount of moisture in the air</td></tr><tr><td colspan="1" valign="top" align="left">8</td><td align="left" colspan="1" valign="top">Global Horizontal Radiation (GHI)</td><td colspan="1" valign="top" align="left">Total radiation received by a horizontal surface</td></tr><tr><td colspan="1" valign="top" align="left">9</td><td align="left" colspan="1" valign="top">Diffuse Horizontal Radiation (DHI)</td><td valign="top" align="left" colspan="1">Diffuse portion of solar radiation on a horizontal surface</td></tr><tr><td colspan="1" valign="top" align="left">10</td><td valign="top" align="left" colspan="1">Wind Direction</td><td valign="top" align="left" colspan="1">The direction of wind</td></tr><tr><td valign="top" align="left" colspan="1">11</td><td valign="top" align="left" colspan="1">Weather Daily Rainfall</td><td valign="top" align="left" colspan="1">The amount of rainfall at the solar system site</td></tr><tr><td valign="top" align="left" colspan="1">12</td><td valign="top" align="left" colspan="1">Radiation Global Tilted</td><td align="left" colspan="1" valign="top">Total radiation received by a tilted surface</td></tr><tr><td align="left" colspan="1" valign="top">13</td><td valign="top" align="left" colspan="1">Radiation Diffuse Tilted</td><td align="left" colspan="1" valign="top">Diffuse portion of radiation on a tilted surface</td></tr></tbody></table></table-wrap></sec><sec><title>3.2. Data Preprocessing</title><p>Reliable forecasting depends on high-quality input data. Three preprocessing steps were performed before model development.</p><p>Handling Missing Data: The missing data, which occurred due to disruptions in sensors and/or communication, were filled through median imputation calculated for both training and testing datasets separately so that no information was leaked from the testing dataset to the training one.</p><p>Data Filtering: Occasionally, anomalous values in sensor measurements arise from hardware constraints or environmental disruptions. Prior to model training, filtering techniques were used to eliminate noisy and inconsistent readings.</p><p>Feature Normalization: Meteorological variables possess different numerical ranges, which may negatively influence machine learning algorithms. Min–Max normalization was applied to scale each feature into the interval [0,1], accelerating model convergence and preventing variables with larger magnitudes from dominating the learning process, as defined in Equation (1):</p><p>x_scaled = (x - min(x)) / (max(x) - min(x))            <italic>(1)</italic></p><p>The dataset is sorted chronologically, and in the case of large-scale model comparisons presented in Section 4, a random sampling of 60,000 rows is used, where a 70:30 split ratio is used between train and test sets (42,000 training and 18,000 testing samples) in order to maintain computability and validity of the results on such large-scale datasets. This method is being used for all of the reported experiments.</p></sec><sec><title>3.3. Feature Selection</title><p>The most informative meteorological variables for SEG forecasting were determined through feature selection. The four top-ranked characteristics from the initial set of thirteen were chosen by each of the three complementary methods that were examined, each of which assessed feature significance from a different angle.</p><sec><title>3.3.1. Pearson Correlation (PC)</title><p>Pearson Correlation (PC) measures the strength of the linear relationship between each predictor variable and the target variable, as given in Equation (2):</p><p><inline-formula><tex-math id="math-1"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle r\ = \ \Sigma(xᵢ\ - \ x\bar{})(yᵢ\ - \ ȳ)\ /\ \lbrack\sqrt{}\Sigma(xᵢ\ - \ x\bar{})²\ \cdot \ \sqrt{}\Sigma(yᵢ\ - \ ȳ)²\rbrack \end{document} ]]></tex-math></inline-formula>      (2)</p><p>Table <xref ref-type="table" rid="table-4">4</xref> reports the correlation coefficients obtained for each candidate feature against the target (Active Power).</p><table-wrap id="table-4" ignoredToc=""><label>Table 4</label><caption><p>Correlation of each feature with the target feature (Active Power)</p></caption><table frame="box" rules="all"><thead><tr><th colspan="1" valign="top" align="left"><bold>No</bold></th><th colspan="1" valign="top" align="left"><bold>Feature</bold></th><th valign="top" align="left" colspan="1"><bold>Correlation (r)</bold></th></tr></thead><tbody><tr><td valign="top" align="left" colspan="1">1</td><td valign="top" align="left" colspan="1">Global_Horizontal_Radiation</td><td valign="top" align="left" colspan="1">0.6947</td></tr><tr><td valign="top" align="left" colspan="1">2</td><td align="left" colspan="1" valign="top">Radiation_Global_Tilted</td><td valign="top" align="left" colspan="1">0.6092</td></tr><tr><td valign="top" align="left" colspan="1">3</td><td valign="top" align="left" colspan="1">Diffuse_Horizontal_Radiation</td><td align="left" colspan="1" valign="top">0.3945</td></tr><tr><td valign="top" align="left" colspan="1">4</td><td colspan="1" valign="top" align="left">Radiation_Diffuse_Tilted</td><td align="left" colspan="1" valign="top">0.3460</td></tr><tr><td valign="top" align="left" colspan="1">5</td><td colspan="1" valign="top" align="left">Weather_Temperature_Celsius</td><td align="left" colspan="1" valign="top">0.3240</td></tr><tr><td colspan="1" valign="top" align="left">6</td><td colspan="1" valign="top" align="left">Weather_Relative_Humidity</td><td align="left" colspan="1" valign="top">0.3194</td></tr><tr><td align="left" colspan="1" valign="top">7</td><td valign="top" align="left" colspan="1">Wind_Speed</td><td valign="top" align="left" colspan="1">0.2769</td></tr><tr><td align="left" colspan="1" valign="top">8</td><td valign="top" align="left" colspan="1">Performance_Ratio</td><td align="left" colspan="1" valign="top">0.0988</td></tr><tr><td align="left" colspan="1" valign="top">9</td><td valign="top" align="left" colspan="1">Wind_Direction</td><td valign="top" align="left" colspan="1">0.0669</td></tr><tr><td valign="top" align="left" colspan="1">10</td><td valign="top" align="left" colspan="1">Weather_Daily_Rainfall</td><td valign="top" align="left" colspan="1">0.0392</td></tr></tbody></table></table-wrap><p>The four features with the highest absolute correlation, Global Horizontal Radiation, Diffuse Horizontal Radiation, Weather Temperature, and Wind Speed, were selected for model development.</p></sec><sec><title>3.3.2. Recursive Feature Elimination (RFE)</title><p>RFE is a wrapper-based technique that iteratively removes the least important predictor according to the performance of a regression estimator, continuing until the required number of features remains, as shown in Table <xref ref-type="table" rid="table-5">5</xref>. Unlike correlation analysis, RFE evaluates feature importance during model training and therefore captures interactions among predictors.</p><table-wrap id="table-5" ignoredToc=""><label>Table 5</label><caption><p>Feature ranking using Recursive Feature Elimination</p></caption><table frame="box" rules="all"><thead><tr><th align="left" colspan="1" valign="top"><bold>No</bold></th><th valign="top" align="left" colspan="1"><bold>Feature</bold></th><th align="left" colspan="1" valign="top"><bold>Rank</bold></th></tr></thead><tbody><tr><td colspan="1" valign="top" align="left">1</td><td align="left" colspan="1" valign="top">Wind_Speed</td><td colspan="1" valign="top" align="left">1</td></tr><tr><td align="left" colspan="1" valign="top">2</td><td align="left" colspan="1" valign="top">Weather_Temperature_Celsius</td><td valign="top" align="left" colspan="1">1</td></tr><tr><td align="left" colspan="1" valign="top">3</td><td valign="top" align="left" colspan="1">Weather_Daily_Rainfall</td><td valign="top" align="left" colspan="1">1</td></tr><tr><td align="left" colspan="1" valign="top">4</td><td align="left" colspan="1" valign="top">Global_Horizontal_Radiation</td><td valign="top" align="left" colspan="1">1</td></tr><tr><td valign="top" align="left" colspan="1">5</td><td valign="top" align="left" colspan="1">Weather_Relative_Humidity</td><td align="left" colspan="1" valign="top">2</td></tr><tr><td valign="top" align="left" colspan="1">6</td><td align="left" colspan="1" valign="top">Radiation_Diffuse_Tilted</td><td valign="top" align="left" colspan="1">3</td></tr><tr><td align="left" colspan="1" valign="top">7</td><td valign="top" align="left" colspan="1">Diffuse_Horizontal_Radiation</td><td colspan="1" valign="top" align="left">4</td></tr><tr><td valign="top" align="left" colspan="1">8</td><td valign="top" align="left" colspan="1">Radiation_Global_Tilted</td><td align="left" colspan="1" valign="top">5</td></tr><tr><td valign="top" align="left" colspan="1">9</td><td align="left" colspan="1" valign="top">Wind_Direction</td><td colspan="1" valign="top" align="left">6</td></tr><tr><td valign="top" align="left" colspan="1">10</td><td align="left" colspan="1" valign="top">Performance_Ratio</td><td valign="top" align="left" colspan="1">7</td></tr></tbody></table></table-wrap><p>A lower rank indicates greater importance. The four top-ranked features — Wind Speed, Global Horizontal Radiation, Weather Temperature, and Weather Relative Humidity — were selected for model development.</p></sec><sec><title>3.3.3. Random Forest Feature Importance (RFFI)</title><p>The relevance of predictors is evaluated in terms of the decrease in impurity in RFFI based on the mean decrease in impurity (MDI), defined by Equation (3):</p><p><tex-math>MDI(Xⱼ) = (1/T) Σₜ Σₙ ∈ Nodes(t) ΔI(n) ⋅ 𝟙{split_feature(n) = Xⱼ}       <italic>(3)</italic></tex-math></p><table-wrap id="table-6" ignoredToc=""><label>Table 6</label><caption><p>Feature importance scores using Random Forest Feature Importance</p></caption><table frame="box" rules="all"><thead><tr><th valign="top" align="left" colspan="1"><bold>No</bold></th><th align="left" colspan="1" valign="top"><bold>Feature</bold></th><th align="left" colspan="1" valign="top"><bold>Importance</bold></th></tr></thead><tbody><tr><td align="left" colspan="1" valign="top">1</td><td colspan="1" valign="top" align="left">Global_Horizontal_Radiation</td><td valign="top" align="left" colspan="1">0.4482</td></tr><tr><td valign="top" align="left" colspan="1">2</td><td align="left" colspan="1" valign="top">Performance_Ratio</td><td align="left" colspan="1" valign="top">0.2225</td></tr><tr><td colspan="1" valign="top" align="left">3</td><td align="left" colspan="1" valign="top">Wind_Direction</td><td colspan="1" valign="top" align="left">0.1628</td></tr><tr><td align="left" colspan="1" valign="top">4</td><td valign="top" align="left" colspan="1">Radiation_Global_Tilted</td><td align="left" colspan="1" valign="top">0.0875</td></tr><tr><td align="left" colspan="1" valign="top">5</td><td align="left" colspan="1" valign="top">Radiation_Diffuse_Tilted</td><td valign="top" align="left" colspan="1">0.0229</td></tr><tr><td valign="top" align="left" colspan="1">6</td><td align="left" colspan="1" valign="top">Wind_Speed</td><td align="left" colspan="1" valign="top">0.0178</td></tr><tr><td colspan="1" valign="top" align="left">7</td><td valign="top" align="left" colspan="1">Weather_Temperature_Celsius</td><td valign="top" align="left" colspan="1">0.0146</td></tr><tr><td align="left" colspan="1" valign="top">8</td><td align="left" colspan="1" valign="top">Weather_Relative_Humidity</td><td valign="top" align="left" colspan="1">0.0119</td></tr><tr><td align="left" colspan="1" valign="top">9</td><td align="left" colspan="1" valign="top">Diffuse_Horizontal_Radiation</td><td align="left" colspan="1" valign="top">0.0113</td></tr><tr><td align="left" colspan="1" valign="top">10</td><td valign="top" align="left" colspan="1">Weather_Daily_Rainfall</td><td colspan="1" valign="top" align="left">0.0005</td></tr></tbody></table></table-wrap><p>The four features with the highest MDI, Global Horizontal Radiation, Performance Ratio, Wind Speed, and Weather Temperature, were selected for model development as shown in Table <xref ref-type="table" rid="table-6">6</xref>. Compared with the statistical techniques, RFFI effectively captures nonlinear relationships and complex feature interactions.</p></sec><sec><title>3.3.4. Comparative Summary</title><p>The three feature selection techniques generated different forecasting subsets because each evaluates relevance using a different criterion: RFFI uses ensemble learning to estimate importance, RFE assesses predictor contribution during model formation, and PC stresses linear dependency. All three feature subsets were assessed using similar machine learning models under identical experimental conditions, and the results were compared using MSE, RMSE, MAE, MAPE, MedAE, EVS, Huber Loss, and R2, rather than assuming that one approach is always better (Section 4).</p><fig id="figure-2" ignoredToc=""><label>Figure 2</label><caption><p>Overlap of top 4 selected features across PC, RFE, and RFFI</p></caption><graphic mime-subtype="png" mimetype="image" xlink:href="https://ijdiic.com/research/article/download/305/version/306/228/2213/International_Journal_of_Data_Informatics_and_Intelligent_Computing-3-5-50-g2.png"><alt-text>Image</alt-text></graphic></fig><p>Figure <xref ref-type="fig" rid="figure-2">2</xref> provides an illustration of the intersection among the three subsets of features described in Tables 4 to 6. Global Horizontal Radiation, Wind Speed, and Weather Temperature are chosen by the three methods, revealing a high level of agreement that such features are the main factors influencing the Active Power generation. Each technique has chosen one feature that the other two techniques have not chosen: Diffuse Horizontal Radiation for PC, Weather Relative Humidity for RFE, and Performance Ratio for RFFI. In doing so, it is clear that the three methods have reached an agreement about the set of important predictors but only differ on the choice of the secondary predictor, which can be explained by the superior forecasting performance achieved by RFFI in Section 4.3.</p></sec></sec><sec><title>3.4. Machine Learning Models</title><p>After feature selection, each subset of features was assessed using four popular regression algorithms, which include Linear Regression (LR), Random Forest (RF), Support Vector Regression (SVR), and Multilayer Perceptron (MLP). The four baseline models are briefly introduced below; their performances are compared in Section 4, where Random Forest is found to be the best among them. The Optimized RBF-SVR algorithm proposed in the paper is described in detail in Section 3.5.</p><sec><title>3.4.1. Linear Regression (LR)</title><p>In LR, the model assumes that the predictor variables are linked to the PV power generation through a linear model. The parameters of the model are found through least squares estimation, which minimizes the sum of squared residuals.</p></sec><sec><title>3.4.2. Random Forest (RF)</title><p>RF is an example of an ensemble method where decision trees are combined through randomly sampling observations and predictors in such a way that the final prediction will be a result of the averaging of all these trees. This helps RF to model nonlinear dependencies without overfitting, since, as can be seen from Section 4, it performs better than other baselines in terms of accuracy.</p></sec><sec><title>3.4.3. Support Vector Regression (SVR)</title><p>The SVR model generalizes Support Vector Machines to regression through the generation of a function that will stay inside an ε-insensitive boundary of the training objective values, in such a way that it reduces the complexity of the model. In the current model, we apply a Radial Basis Function (RBF) kernel; the fully optimized model can be found in Section 3.5.</p></sec><sec><title>3.4.4. Multilayer Perceptron (MLP)</title><p>MLP is a feedforward neural network that consists of one or multiple layers, which are trained using backpropagation to learn the nonlinear relationship between inputs and outputs. MLP needs proper parameter tuning to prevent overfitting and usually has longer training times than LR and RF.</p></sec></sec><sec><title>3.5. Proposed Optimized RBF-SVR</title><p>While traditional SVR has shown competitive performance in solar energy forecasting, feature quality and hyperparameter choice have a significant impact on prediction accuracy. This section combines feature selection, polynomial feature transformation, feature normalization, and randomized hyperparameter optimization into a single forecasting pipeline.</p><p>The reason why RBF-SVR was preferred over deep architectures (such as LSTM, GRU, Transformer-based models; see Section 2.2) is two-fold. Firstly, kernel-based models such as SVR usually need considerably fewer training samples to perform well compared to deep architectures. This is due to the fact that this research evaluates models on a small 4-5 feature set for each feature selection algorithm, rather than large multivariate sequences as input. Secondly, training deep architectures usually takes considerably more computational resources and time to be fine-tuned. The training times for the Optimized RBF-SVR presented in Section 4.4 demonstrate that, on average, it needs only 21.25 seconds to complete the training, while in the literature (see Section 2.2) LSTM- and Transformer-based solar forecasting models require much more time for training. The speed of training is explained by the fact that the objective of this research is resource-efficient forecasting, not absolute accuracy regardless of the computational time needed for that.</p><sec><title>3.5.1. Radial Basis Function Kernel</title><p>The RBF kernel was selected for its ability to model the highly nonlinear relationship between meteorological variables and photovoltaic power generation, as given in Equation (4). The resulting SVR decision function is given in Equation (5):</p><p><tex-math>K(xᵢ, xⱼ) = exp( − γ∥xᵢ − xⱼ∥²)         <italic>(4)</italic></tex-math></p><p><tex-math>f(x) = Σᵢ (αᵢ − αᵢ * ) K(xᵢ, x) + b        <italic>(5)</italic></tex-math></p><p>where γ controls the kernel width; its appropriate selection balances model flexibility and generalization performance.</p></sec><sec><title>3.5.2. Polynomial Feature Transformation</title><p>In order to enhance the nonlinear relationship between the chosen meteorological variables, the polynomial features transformation step was used prior to regression modeling according to Equation (6):</p><p><inline-formula><tex-math id="math-2"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \Phi(x)\ = \ \{ 1,\ x,\ x²,\ \ldots,\ x\hat{}\delta\} \end{document} ]]></tex-math></inline-formula>         (6)</p></sec><sec><title>3.5.3. Hyperparameter Optimization</title><p>The SVR prediction power depends mainly on the three hyperparameters: the regularization parameter C, kernel gamma (γ), and insensitive loss ε. Instead of selecting them manually, our framework uses RandomizedSearchCV with an extensive search space P = {C, γ, ε}. An SVR model is trained and evaluated for each combination of hyperparameters, and the one giving the best result is kept. On the selected feature set with the RFFI method for our proposed model, the search resulted in C = 0.4642, γ = 0.0492, and ε = 0.1041 with polynomial feature transformation with degree 2 before using the RBF kernel (Section 3.5.2).</p></sec><sec><title>3.5.4. Optimization Algorithm</title><p>Algorithm 1: Optimized Radial Basis Function Support Vector Regression (RBF-SVR)</p><p>Input:</p><p>Dataset D = {d1, d2, ..., dn}</p><p>Feature set F = {f1, f2, f3, f4}</p><p>Target variable y</p><p>SVR hyperparameter space P = {p1, p2, ..., pk}</p><p>Training/testing split ratio alpha</p><p>Polynomial degree delta</p><p>Evaluation metrics M = {MSE, RMSE, MAE, MedAE, EVS, Huber Loss, R^2}</p><p>Number of search iterations T</p><p>Training Process:</p><p>1. Load dataset D containing features and target variable.</p><p>2. Select the four most relevant features F (Section 3.3).</p><p>3. Split D into training set D_train and testing set D_test based on ratio alpha.</p><p>4. Handle missing data using median imputation, fit on D_train only and apply to both D_train and D_test.</p><p>5. Apply polynomial feature transformation of degree delta to F.</p><p>6. Standardize features F using z-score standardization (fit on D_train only) to produce F_hat.</p><p>7. Generate T random hyperparameter combinations from P.</p><p>8. For each combination p_t, train SVR(p_t) on D_train with features F_hat_train and target y_train.</p><p>9. Predict y_hat_test on the testing set using the trained SVR(p_t).</p><p>10. Evaluate each SVR(p_t) using metric set M.</p><p>11. Select the best model SVR* with the optimal hyperparameter combination p* based on M.</p><p>Testing Process:</p><p>Input the testing data F_hat_test into SVR* to obtain the final prediction y_hat* = SVR*(F_hat_test).</p><p>Output:</p><p>The final predicted target y_hat* for the testing data.</p></sec><sec><title>3.5.5. Computational Complexity</title><p>Reducing the number of predictor variables from thirteen to four substantially decreases computational complexity and accelerates model training. Assuming N training samples and d predictor variables, the computational complexity of kernel construction is approximately O(N²d). By reducing d from 13 to 4, the proposed framework significantly lowers computational cost while preserving forecasting performance, as confirmed experimentally in Section 4.4.</p></sec></sec><sec><title>3.6. Experimental Setup</title><p>The proposed forecasting framework was implemented in Python 3.11 using the Google Colab environment, with additional validation performed in a Kaggle Notebook. Data preprocessing, feature selection, machine learning implementation, and performance evaluation were carried out using the Scikit-learn, NumPy, Pandas, and Matplotlib libraries. Hyperparameter optimization for the proposed RBF-SVR model was performed using RandomizedSearchCV on the training data only. Table <xref ref-type="table" rid="table-7">7</xref> summarizes the experimental configuration.</p><table-wrap ignoredToc="" id="table-7"><label>Table 7</label><caption><p>Experimental settings</p></caption><table frame="box" rules="all"><thead><tr><th colspan="1" valign="top" align="left"><bold>Parameter</bold></th><th valign="top" align="left" colspan="1"><bold>Value</bold></th></tr></thead><tbody><tr><td valign="top" align="left" colspan="1">Programming Language</td><td align="left" colspan="1" valign="top">Python 3.11</td></tr><tr><td valign="top" align="left" colspan="1">Environment</td><td align="left" colspan="1" valign="top">Google Colab, Kaggle Notebook</td></tr><tr><td align="left" colspan="1" valign="top">Libraries</td><td colspan="1" valign="top" align="left">Scikit-learn, NumPy, Pandas, Matplotlib</td></tr><tr><td align="left" colspan="1" valign="top">Dataset</td><td align="left" colspan="1" valign="top">DKASC (Site 28, Alice Springs)</td></tr><tr><td valign="top" align="left" colspan="1">Sampling Interval</td><td align="left" colspan="1" valign="top">5 minutes</td></tr><tr><td align="left" colspan="1" valign="top">Study Period</td><td align="left" colspan="1" valign="top">Jan–Dec 2023</td></tr><tr><td valign="top" align="left" colspan="1">Total Features</td><td align="left" colspan="1" valign="top">13</td></tr><tr><td valign="top" align="left" colspan="1">Selected Features</td><td align="left" colspan="1" valign="top">4</td></tr><tr><td align="left" colspan="1" valign="top">Feature Selection Techniques</td><td valign="top" align="left" colspan="1">PC, RFE, RFFI</td></tr><tr><td align="left" colspan="1" valign="top">Baseline Models</td><td align="left" colspan="1" valign="top">LR, RF, SVR, MLP</td></tr><tr><td align="left" colspan="1" valign="top">Proposed Model</td><td align="left" colspan="1" valign="top">Optimized RBF-SVR</td></tr><tr><td valign="top" align="left" colspan="1">Train/Test Split</td><td align="left" colspan="1" valign="top">80:20 (chronological)</td></tr><tr><td colspan="1" valign="top" align="left">Hyperparameter Optimization</td><td align="left" colspan="1" valign="top">Randomized Search</td></tr></tbody></table></table-wrap></sec><sec><title>3.7. Evaluation Metrics</title><p>Model performance was assessed using eight complementary regression metrics: MSE, RMSE, MAE, MAPE, MedAE, EVS, Huber Loss, and R². Error-based metrics quantify prediction accuracy, whereas EVS and R² evaluate the goodness-of-fit between predicted and observed photovoltaic power output. Their mathematical definitions are given below.</p><p><tex-math>MSE = (1/n) Σ(yᵢ − ŷᵢ)²       <italic>(7)</italic></tex-math></p><p><inline-formula><tex-math id="math-3"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle RMSE\ = \ \sqrt{}\lbrack(1/n)\ \Sigma(yᵢ\ - \ ŷᵢ)²\rbrack \end{document} ]]></tex-math></inline-formula><italic>(8)</italic></p><p><tex-math>MAE = (1/n) Σ|yᵢ − ŷᵢ|       <italic>(9)</italic></tex-math></p><p><tex-math>MAPE = (1/n) Σ|(yᵢ − ŷᵢ)/yᵢ| × 100       <italic>(10)</italic></tex-math></p><p><tex-math>MedAE = median(|y₁ − ŷ₁|, |y₂ − ŷ₂|, …, |yₙ − ŷₙ|)     <italic>(11)</italic></tex-math></p><p><tex-math>EVS = 1 − Var(y − ŷ) / Var(y)         <italic>(12)</italic></tex-math></p><p><tex-math>R² = 1 − Σ(y − ŷ)² / Σ(y − ȳ)²        <italic>(13)</italic></tex-math></p><p><tex-math>Lδ(a) = ½a²  if |a| ≤ δ;   Lδ(a) = δ(|a| − ½δ)        <italic>(14)</italic></tex-math></p><p>where yᵢ is the actual value, ŷᵢ is the predicted value, n is the number of observations, ȳ is the mean of the observed values, and δ is the Huber threshold parameter. Lower values of MSE, RMSE, MAE, MAPE, MedAE, and Huber Loss indicate better prediction accuracy, whereas higher values of EVS and R² represent superior model performance, as defined in Equations (7) to (14) above. Table <xref ref-type="table" rid="table-8">8</xref> summarizes the objective of each metric.</p><p>MAPE is not included in the results in Tables <xref rid="table-9" ref-type="table">9</xref>, <xref ref-type="table" rid="table-10">10</xref>, and <xref ref-type="table" rid="table-11">11</xref> due to its being undefined for yᵢ = 0 and mathematically unstable when yᵢ ≈ 0, which is common for Active Power in the evening and dawn/dusk transitions in the dataset at hand. Even upon restriction of MAPE calculation to daytime samples (Active Power &gt; 0.05), the MAPE in this experiment continued to be in the hundreds of percent, owing to small denominators that still occurred during the sunrise/set ramp. It is a known flaw of MAPE for signals that pass through or approach zero, which is the reason why this paper uses MedAE and Huber Loss (which behave fine regardless of proximity to zero) as its main robust error measures.</p><table-wrap id="table-8" ignoredToc=""><label>Table 8</label><caption><p>Performance metrics</p></caption><table rules="all" frame="box"><thead><tr><th align="left" colspan="1" valign="top"><bold>Metric</bold></th><th align="left" colspan="1" valign="top"><bold>Objective</bold></th><th align="left" colspan="1" valign="top"><bold>Better Value</bold></th></tr></thead><tbody><tr><td valign="top" align="left" colspan="1">MSE</td><td align="left" colspan="1" valign="top">Average squared error</td><td valign="top" align="left" colspan="1">Lower</td></tr><tr><td colspan="1" valign="top" align="left">RMSE</td><td valign="top" align="left" colspan="1">Prediction error (same units as target)</td><td valign="top" align="left" colspan="1">Lower</td></tr><tr><td align="left" colspan="1" valign="top">MAE</td><td valign="top" align="left" colspan="1">Average absolute error</td><td valign="top" align="left" colspan="1">Lower</td></tr><tr><td align="left" colspan="1" valign="top">MAPE</td><td align="left" colspan="1" valign="top">Average percentage error (not reported in results — unstable near zero)</td><td align="left" colspan="1" valign="top">Lower</td></tr><tr><td valign="top" align="left" colspan="1">MedAE</td><td align="left" colspan="1" valign="top">Median prediction error (robust to outliers)</td><td valign="top" align="left" colspan="1">Lower</td></tr><tr><td valign="top" align="left" colspan="1">EVS</td><td align="left" colspan="1" valign="top">Explained variance</td><td colspan="1" valign="top" align="left">Higher</td></tr><tr><td valign="top" align="left" colspan="1">Huber Loss</td><td align="left" colspan="1" valign="top">Robust regression loss</td><td align="left" colspan="1" valign="top">Lower</td></tr><tr><td valign="top" align="left" colspan="1">R²</td><td align="left" colspan="1" valign="top">Goodness-of-fit</td><td valign="top" align="left" colspan="1">Higher</td></tr></tbody></table></table-wrap></sec></sec><sec><title>4. RESULTS AND DISCUSSION</title><p>Three independent experiments were conducted using the feature subsets selected by PC, RFE, and RFFI. Each feature subset was evaluated using LR, RF, SVR, MLP, and the proposed Optimized RBF-SVR under identical experimental conditions. Note that the target variable retains its original measurement scale in the tables below (it is not the [0,1]-normalized predictor scale used internally for training), so error values are directly comparable across models within this study.</p><sec><title>4.1. Results Using Features Selected by PC</title><p>Figure <xref ref-type="fig" rid="figure-3">3</xref> contains all seven metrics listed in Table <xref ref-type="table" rid="table-9">9</xref>, separated into two panels to keep metrics with conflicting meanings on one coordinate axis separate: error metrics (MSE, RMSE, MAE, MedAE, Huber Loss) for which low values are good, and goodness-of-fit metrics (R-squared, EVS) for which high values are good. There is no model surpassing R² = 0.57 for the PC-selected features, and the Optimized RBF-SVR performs about equally well to the Linear Regression using this particular set of features (MSE = 1.134 compared to 1.152) – falling behind the RF and MLP. The absence of a clear winner is by itself instructive: it suggests that PC's four selected features, being physically correlated radiation measurements (Table <xref ref-type="table" rid="table-4">4</xref>), are not enough of a signal for any of the four models.</p><table-wrap id="table-9" ignoredToc=""><label>Table 9</label><caption><p>Evaluation results of ML models on features selected by PC</p></caption><table frame="box" rules="all"><thead><tr><th valign="top" align="left" colspan="1"><bold>Model</bold></th><th valign="top" align="left" colspan="1"><bold>MSE</bold></th><th align="left" colspan="1" valign="top"><bold>RMSE</bold></th><th align="left" colspan="1" valign="top"><bold>MAE</bold></th><th valign="top" align="left" colspan="1"><bold>R²</bold></th><th valign="top" align="left" colspan="1"><bold>MedAE</bold></th><th valign="top" align="left" colspan="1"><bold>EVS</bold></th><th align="left" colspan="1" valign="top"><bold>Huber Loss</bold></th></tr></thead><tbody><tr><td valign="top" align="left" colspan="1">LR</td><td valign="top" align="left" colspan="1">1.1516</td><td align="left" colspan="1" valign="top">1.0731</td><td valign="top" align="left" colspan="1">0.6273</td><td align="left" colspan="1" valign="top">0.4917</td><td align="left" colspan="1" valign="top">0.0432</td><td valign="top" align="left" colspan="1">0.4917</td><td align="left" colspan="1" valign="top">0.4134</td></tr><tr><td valign="top" align="left" colspan="1">RF</td><td align="left" colspan="1" valign="top">0.9889</td><td align="left" colspan="1" valign="top">0.9944</td><td align="left" colspan="1" valign="top">0.5055</td><td valign="top" align="left" colspan="1">0.5635</td><td valign="top" align="left" colspan="1">0.0007</td><td valign="top" align="left" colspan="1">0.5635</td><td align="left" colspan="1" valign="top">0.3321</td></tr><tr><td align="left" colspan="1" valign="top">MLP</td><td valign="top" align="left" colspan="1">0.9971</td><td align="left" colspan="1" valign="top">0.9986</td><td align="left" colspan="1" valign="top">0.5587</td><td colspan="1" valign="top" align="left">0.5599</td><td align="left" colspan="1" valign="top">0.0258</td><td valign="top" align="left" colspan="1">0.5599</td><td valign="top" align="left" colspan="1">0.3623</td></tr><tr><td valign="top" align="left" colspan="1">Optimized RBF-SVR</td><td valign="top" align="left" colspan="1">1.1341</td><td valign="top" align="left" colspan="1">1.0650</td><td valign="top" align="left" colspan="1">0.5292</td><td align="left" colspan="1" valign="top">0.4994</td><td align="left" colspan="1" valign="top">0.0370</td><td align="left" colspan="1" valign="top">0.5091</td><td valign="top" align="left" colspan="1">0.3518</td></tr></tbody></table></table-wrap><fig id="figure-3" ignoredToc=""><label>Figure 3</label><caption><p>Comparison of classifier performances with PC feature selection in terms of all metrics</p></caption><graphic mime-subtype="png" mimetype="image" xlink:href="https://ijdiic.com/research/article/download/305/version/306/228/2214/International_Journal_of_Data_Informatics_and_Intelligent_Computing-3-5-50-g3.png"><alt-text>Image</alt-text></graphic></fig><p>With PC-chosen features, all models showed similar performance, with LR yielding an R² of 0.4917, RF giving an R² of 0.5635, MLP yielding an R² of 0.5599, and OptRBF-SVR yielding an R² of 0.4994. No model achieved an R² greater than 0.57 using this set of features, suggesting that the four PC chosen features, which are different radiations that are physically related to each other (see Table <xref ref-type="table" rid="table-4">4</xref>), were not diverse enough for accurate predictions.</p></sec><sec><title>4.2. Results Using Features Selected by RFE</title><p>Figure <xref ref-type="fig" rid="figure-4">4</xref> is consistent with Figure <xref ref-type="fig" rid="figure-3">3</xref> in terms of structure, being divided into error metrics (on the left side) and goodness-of-fit metrics (on the right side) in order to prevent combining metrics having opposite meanings. The results obtained by using the feature subset chosen by the RFE algorithm were very similar to those for PC; all the models performed with R² less than 0.60, with the highest value of R² belonging to MLP (R² = 0.5957), slightly exceeding RF (R² = 0.5814), as shown in Table <xref ref-type="table" rid="table-10">10</xref>. Once again, the Optimized RBF-SVR model performed similarly to Linear Regression (MSE = 1.128 vs. 1.148), showing that neither PC nor the RFE algorithm helped to find a suitable feature subset.</p><table-wrap id="table-10" ignoredToc=""><label>Table 10</label><caption><p>Evaluation results of ML models on features selected by RFE</p></caption><table frame="box" rules="all"><thead><tr><th valign="top" align="left" colspan="1"><bold>Model</bold></th><th valign="top" align="left" colspan="1"><bold>MSE</bold></th><th align="left" colspan="1" valign="top"><bold>RMSE</bold></th><th valign="top" align="left" colspan="1"><bold>MAE</bold></th><th align="left" colspan="1" valign="top"><bold>R²</bold></th><th align="left" colspan="1" valign="top"><bold>MedAE</bold></th><th valign="top" align="left" colspan="1"><bold>EVS</bold></th><th valign="top" align="left" colspan="1"><bold>Huber Loss</bold></th></tr></thead><tbody><tr><td valign="top" align="left" colspan="1">LR</td><td align="left" colspan="1" valign="top">1.1478</td><td align="left" colspan="1" valign="top">1.0713</td><td align="left" colspan="1" valign="top">0.6485</td><td align="left" colspan="1" valign="top">0.4934</td><td valign="top" align="left" colspan="1">0.1161</td><td align="left" colspan="1" valign="top">0.4934</td><td align="left" colspan="1" valign="top">0.4164</td></tr><tr><td valign="top" align="left" colspan="1">RF</td><td valign="top" align="left" colspan="1">0.9484</td><td align="left" colspan="1" valign="top">0.9739</td><td valign="top" align="left" colspan="1">0.4859</td><td valign="top" align="left" colspan="1">0.5814</td><td align="left" colspan="1" valign="top">0.0027</td><td align="left" colspan="1" valign="top">0.5814</td><td align="left" colspan="1" valign="top">0.3159</td></tr><tr><td colspan="1" valign="top" align="left">MLP</td><td colspan="1" valign="top" align="left">0.9159</td><td align="left" colspan="1" valign="top">0.9570</td><td align="left" colspan="1" valign="top">0.5412</td><td valign="top" align="left" colspan="1">0.5957</td><td align="left" colspan="1" valign="top">0.0851</td><td valign="top" align="left" colspan="1">0.5959</td><td align="left" colspan="1" valign="top">0.3322</td></tr><tr><td valign="top" align="left" colspan="1">Optimized RBF-SVR</td><td valign="top" align="left" colspan="1">1.1277</td><td valign="top" align="left" colspan="1">1.0619</td><td align="left" colspan="1" valign="top">0.5481</td><td align="left" colspan="1" valign="top">0.5022</td><td valign="top" align="left" colspan="1">0.1239</td><td valign="top" align="left" colspan="1">0.5092</td><td align="left" colspan="1" valign="top">0.3413</td></tr></tbody></table></table-wrap><fig ignoredToc="" id="figure-4"><label>Figure 4</label><caption><p>Classifier performance for all metrics using RFE feature selection</p></caption><graphic xlink:href="https://ijdiic.com/research/article/download/305/version/306/228/2215/International_Journal_of_Data_Informatics_and_Intelligent_Computing-3-5-50-g4.png" mime-subtype="png" mimetype="image"><alt-text>Image</alt-text></graphic></fig><p>The features selected through RFE gave similar performances across all models, with R² ranging from 0.49 to 0.60, which was in line with PC. The MLP model performed better than other models for these features (R²=0.5957), slightly beating the RF (R²=0.5814). There is no feature subset obtained from either PC or RFE that made any of the five models perform well.</p></sec><sec><title>4.3. Results Using Features Selected by RFFI</title><p>Figure <xref rid="figure-5" ref-type="fig">5</xref> also uses a two-plot design similar to Figures <xref ref-type="fig" rid="figure-3">3</xref> and <xref ref-type="fig" rid="figure-4">4</xref>, where the error bars indicate ± one standard deviation calculated over the five repetitions listed in Table <xref ref-type="table" rid="table-11">11</xref>. Unlike the PC and RFE techniques, the RFFI technique leads to a clear distinction between the models: the Random Forest is the top performer on all measures (R² = 0.896 ± 0.001), followed closely by MLP (R² = 0.889 ± 0.002), the Optimized RBF-SVR (R² = 0.857 ± 0.007), all clearly ahead of Linear Regression (R² = 0.532 ± 0.007). The error bars in the plots for RF and MLP appear visually almost invisible, which proves that the results presented are indeed very consistent over multiple independent runs instead of being dependent on one particular training/testing data split. It should also be noted that the MedAE of the Random Forest is now represented correctly close to zero (0.000 ± 0.000) instead of the erroneous number 4.0 in the earlier single-run pass of the experiment; the number 4.0 was obviously a mistake in transcription (a calculated value of 4.0 × 10⁻⁶ was transcribed as 4.0000).</p><p>RFFI’s selected subset of features provided a clear advantage over PC and RFE for all models and was statistically validated in 5 independent repeat executions of the experiment: Random Forest provided the best results (mean R² = 0.8960 ± 0.0008, mean MSE = 0.2295 ± 0.0036), followed by MLP (mean R² = 0.8893 ± 0.0022) and Opt-RBF-SVR (mean R² = 0.8568 ± 0.0072). Low standard deviations in repeats (all less than 0.011 for R²) indicate that results are consistent and not influenced by any lucky train-test set partitioning. This difference between RFFI and PC/RFE is much greater than in the case of any model on PC and RFE, which confirms the selection of physically relevant features (Global Horizontal Radiation, Performance Ratio, Wind Direction, and Tilted Global Radiation; Table <xref ref-type="table" rid="table-6">6</xref>) by RFFI.</p><table-wrap id="table-11" ignoredToc=""><label>Table 11</label><caption><p>Final model comparison using the best-performing feature selection technique (RFFI)</p></caption><table frame="box" rules="all"><thead><tr><th align="left" colspan="1" valign="top"><bold>Model</bold></th><th align="left" colspan="1" valign="top"><bold>MSE</bold></th><th align="left" colspan="1" valign="top"><bold>RMSE</bold></th><th valign="top" align="left" colspan="1"><bold>MAE</bold></th><th valign="top" align="left" colspan="1"><bold>R²</bold></th><th align="left" colspan="1" valign="top"><bold>MedAE</bold></th><th valign="top" align="left" colspan="1"><bold>EVS</bold></th><th align="left" colspan="1" valign="top"><bold>Huber Loss</bold></th></tr></thead><tbody><tr><td align="left" colspan="1" valign="top">LR</td><td valign="top" align="left" colspan="1">1.0322±0.0062</td><td align="left" colspan="1" valign="top">1.0159±0.0031</td><td valign="top" align="left" colspan="1">0.6979±0.0041</td><td valign="top" align="left" colspan="1">0.5322±0.0069</td><td valign="top" align="left" colspan="1">0.3388±0.0095</td><td valign="top" align="left" colspan="1">0.5322±0.0068</td><td align="left" colspan="1" valign="top">0.3897±0.0029</td></tr><tr><td valign="top" align="left" colspan="1">RF</td><td align="left" colspan="1" valign="top">0.2295±0.0036</td><td align="left" colspan="1" valign="top">0.4791±0.0038</td><td align="left" colspan="1" valign="top">0.1163±0.0021</td><td valign="top" align="left" colspan="1">0.8960±0.0008</td><td valign="top" align="left" colspan="1">0.0000±0.0000</td><td valign="top" align="left" colspan="1">0.8960±0.0008</td><td valign="top" align="left" colspan="1">0.0717±0.0012</td></tr><tr><td valign="top" align="left" colspan="1">MLP</td><td valign="top" align="left" colspan="1">0.2443±0.0053</td><td colspan="1" valign="top" align="left">0.4942±0.0054</td><td valign="top" align="left" colspan="1">0.1835±0.0063</td><td valign="top" align="left" colspan="1">0.8893±0.0022</td><td valign="top" align="left" colspan="1">0.0410±0.0106</td><td colspan="1" valign="top" align="left">0.8894±0.0023</td><td valign="top" align="left" colspan="1">0.0859±0.0017</td></tr><tr><td align="left" colspan="1" valign="top">RBF-SVR</td><td align="left" colspan="1" valign="top">0.3160±0.0100</td><td valign="top" align="left" colspan="1">0.5621±0.0089</td><td align="left" colspan="1" valign="top">0.1880±0.0048</td><td align="left" colspan="1" valign="top">0.8568±0.0072</td><td valign="top" align="left" colspan="1">0.0683±0.0000</td><td valign="top" align="left" colspan="1">0.8592±0.0051</td><td valign="top" align="left" colspan="1">0.0899±0.0015</td></tr></tbody></table></table-wrap><p>Note: an earlier single-run pass of this experiment reported an anomalous Random Forest MedAE value; across 5 independent repeats with different random samples and train/test splits (reported above), Random Forest's true MedAE is consistently near zero (mean 0.0000, SD 0.0000), confirming the earlier value was not representative and that Random Forest's median error on this feature set is in fact the lowest of all four models evaluated.</p><fig id="figure-5" ignoredToc=""><label>Figure 5</label><caption><p>Comparison of classifier results for all metrics using RFFI feature selection</p></caption><graphic mime-subtype="png" mimetype="image" xlink:href="https://ijdiic.com/research/article/download/305/version/306/228/2216/International_Journal_of_Data_Informatics_and_Intelligent_Computing-3-5-50-g5.png"><alt-text>Image</alt-text></graphic></fig></sec><sec><title>4.4. Training Time Analysis</title><p>Table <xref ref-type="table" rid="table-12">12</xref> and Figure  <xref ref-type="fig" rid="figure-1">6</xref> show average measured time (± SD in 5 experiments) for training of all models with RFFI feature set (60,000 rows in each experiment's training set). Linear regression was the fastest-trained model (less than 0.01 seconds), although with the worst accuracy performance. The Optimized RBF-SVR was the second-fastest model trained (21.25 ± 1.12 seconds), although faster than the Random Forest (average training time – 35.78 ± 2.42 seconds), with slightly better accuracy. The fastest nonlinear model was MLP (12.76 ± 2.25 seconds) with accuracy near Random Forest. That was a real compromise between accuracy and speed of calculations: RF is the most accurate model, and Optimized RBF-SVR provides the best compromise in terms of accuracy (R² difference near 4%) and speed.</p><p>Figure </p><p> 6</p><p> illustrates the exact measured training time for all four models using all three feature selection methods. The training time of the Optimized RBF-SVR varies greatly based on which feature set it uses. It is the slowest among all models when used with PC features (259.8 seconds due to the hyperparameters selected for it in that feature set), of moderate speed with RFE features (45.3 seconds), and relatively fast with RFFI features (21.3 seconds), where it reaches its highest accuracy (Section 4.3). This means that the training time of the Optimized RBF-SVR is not fast in general, but is dependent on the combination of features and their corresponding hyperparameters; therefore, the good balance of accuracy and speed of training achieved with the RFFI feature set (Section 4.4) is more relevant for this research.</p><fig id="figure-1" ignoredToc=""><label>Figure 6</label><caption><p>Approximate training time comparison across models and feature selection techniques</p></caption><graphic mimetype="image" xlink:href="https://ijdiic.com/research/article/download/305/version/306/228/2217/International_Journal_of_Data_Informatics_and_Intelligent_Computing-3-5-50-g6.png" mime-subtype="png"><alt-text>Image</alt-text></graphic></fig><table-wrap id="table-12" ignoredToc=""><label>Table 12</label><caption><p>Approximate training time by model (seconds)</p></caption><table frame="box" rules="all"><thead><tr><th colspan="1" valign="top" align="left"><bold>Model</bold></th><th align="left" colspan="1" valign="top"><bold>PC features</bold></th><th align="left" colspan="1" valign="top"><bold>RFE features</bold></th><th align="left" colspan="1" valign="top"><bold>RFFI features</bold></th></tr></thead><tbody><tr><td valign="top" align="left" colspan="1">LR</td><td align="left" colspan="1" valign="top">0.008</td><td colspan="1" valign="top" align="left">0.007</td><td align="left" colspan="1" valign="top">0.01±0.00</td></tr><tr><td align="left" colspan="1" valign="top">RF</td><td valign="top" align="left" colspan="1">39.73</td><td valign="top" align="left" colspan="1">20.91</td><td align="left" colspan="1" valign="top">35.78±2.42</td></tr><tr><td align="left" colspan="1" valign="top">MLP</td><td valign="top" align="left" colspan="1">11.34</td><td valign="top" align="left" colspan="1">17.11</td><td align="left" colspan="1" valign="top">12.76±2.25</td></tr><tr><td align="left" colspan="1" valign="top">Optimized RBF-SVR</td><td valign="top" align="left" colspan="1">259.77</td><td align="left" colspan="1" valign="top">45.32</td><td valign="top" align="left" colspan="1">21.25±1.12</td></tr></tbody></table></table-wrap></sec><sec><title>4.5. Comparative Analysis of Feature Selection Techniques</title><p>The feature selection approach had a major impact on forecasting accuracy. PC and RFE generated feature subsets that limited all models to having an R² value below 0.60, whereas the use of RFFI allowed for R² greater than 0.85 for all three models, RF, MLP, and Optimized RBF-SVR alike. This difference is significantly greater than the one found in previous research, where the same single feature selection approach was used in isolation, showing that the choice of feature selection approach may matter even more than the choice of regression model when performing this kind of forecasting. The benefit of RFFI probably derives from the fact that this approach includes consideration of nonlinear, tree-based feature importance (which includes cyclical variables, such as Wind Direction, where sine/cosine encoding preserves the cyclical nature of the data).</p></sec><sec><title>4.6. Comparative Analysis of ML Models</title><p>Random Forest had the best performance among the five methods analyzed, marginally beating MLP and Optimized RBF-SVR when trained on RFFI-selected features. All three nonlinear methods showed a significant improvement over Linear Regression, clearly indicating that the photovoltaic power generation cannot be well modeled using a linear relationship between the predictors chosen and the response. However, Optimized RBF-SVR did not have the best raw performance among the five methods but had the second-best performance while taking only about half the time taken by Random Forest to train, hence giving it an advantage wherever training efficiency is a concern (Section 4.4).</p></sec><sec><title>4.7. Comparison with State-of-the-Art Methods</title><p>To contextualize the proposed framework, Table <xref ref-type="table" rid="table-13">13</xref> compares it against prior studies that used the same DKASC dataset and therefore report errors on a comparable measurement scale. It is worth mentioning that the studies to be compared <xref ref-type="bibr" rid="BIBR-22">[22]</xref><xref ref-type="bibr" rid="BIBR-24">[24]</xref><xref ref-type="bibr" rid="BIBR-26">[26]</xref> provide results from other DKASC installations than the one analyzed in this study and, at times, use other measurement scales or normalization approaches. Thus, the results provided in Table <xref ref-type="table" rid="table-13">13</xref> should be considered in this context and cannot be used as a strictly apples-to-apples accuracy comparison. Previous works also utilized a greater number of features (7–13) and did not compute any advanced metrics such as MedAE, EVS, and Huber Loss. What makes the current research valuable is not outperforming previous works with higher error values, but demonstrating the possibility of achieving R² greater than 0.85 using a compact set of four features selected with the RFFI algorithm, and using three model types (RF, MLP, and Optimized RBF-SVR).</p><table-wrap id="table-13" ignoredToc=""><label>Table 13</label><caption><p>Comparison with prior studies using the DKASC dataset</p></caption><table frame="box" rules="all"><thead><tr><th valign="top" align="left" colspan="1"><bold>Reference</bold></th><th valign="top" align="left" colspan="1"><bold>Method</bold></th><th valign="top" align="left" colspan="1"><bold>MSE</bold></th><th align="left" colspan="1" valign="top"><bold>RMSE</bold></th><th align="left" colspan="1" valign="top"><bold>MAE</bold></th><th align="left" colspan="1" valign="top"><bold>R²</bold></th></tr></thead><tbody><tr><td valign="top" align="left" colspan="1"><xref rid="BIBR-22" ref-type="bibr">[22]</xref></td><td valign="top" align="left" colspan="1">HIMVO-SVM</td><td colspan="1" valign="top" align="left">0.0025</td><td align="left" colspan="1" valign="top">0.0500</td><td valign="top" align="left" colspan="1">—</td><td align="left" colspan="1" valign="top">—</td></tr><tr><td colspan="1" valign="top" align="left"><xref ref-type="bibr" rid="BIBR-24">[24]</xref></td><td align="left" colspan="1" valign="top">ACO-SVM</td><td valign="top" align="left" colspan="1">0.0349</td><td valign="top" align="left" colspan="1">0.1868</td><td valign="top" align="left" colspan="1">0.1569</td><td colspan="1" valign="top" align="left">0.9970</td></tr><tr><td align="left" colspan="1" valign="top"><xref ref-type="bibr" rid="BIBR-26">[26]</xref></td><td align="left" colspan="1" valign="top">WPD-LSTM</td><td valign="top" align="left" colspan="1">0.0372</td><td align="left" colspan="1" valign="top">0.1929</td><td align="left" colspan="1" valign="top">—</td><td align="left" colspan="1" valign="top">—</td></tr><tr><td colspan="1" valign="top" align="left">Proposed Model</td><td align="left" colspan="1" valign="top">RFFI + Random Forest</td><td align="left" colspan="1" valign="top">0.2295±0.0036</td><td valign="top" align="left" colspan="1">0.4791±0.0038</td><td colspan="1" valign="top" align="left">0.1163±0.0021</td><td valign="top" align="left" colspan="1">0.8960±0.0008</td></tr></tbody></table></table-wrap></sec><sec><title>4.8. Discussion</title><p>The obtained results, which have been consistently reproduced in 5 independent repeated experiments with low standard deviations throughout, show that feature selection is significantly more important for the accuracy of predictions of this problem compared to model selection: the set of features selected by RFFI allowed achieving a mean R² above 0.85 for RF, MLP, and Optimized RBF-SVR models, while PC and RFE could not bring any model to R² above 0.60 irrespective of the algorithm. Such a gap is much wider compared to those presented in previous papers analyzing a single feature selection method separately, and it shows how much the choice of feature selection method can be important for this particular forecasting task rather than the choice of a regression algorithm itself. The RFFI's advantage can be attributed to the ability of the method to detect nonlinear and tree-based feature relevance (including the cyclic Wind Direction feature, which has been represented via sine/cosine transform); PC finds linear correlation only and selects four highly correlated radiation measurement features, while RFE, being greedy, selects the Weather Daily Rainfall feature with little variance.</p></sec></sec><sec><title>5. CONCLUSION</title><p>The feature selection approach was proposed for the short-term prediction of solar energy production. The efficiency of three feature selection algorithms - Pearson Correlation, Recursive Feature Elimination, and Random Forest Feature Importance – was compared based on four machine learning algorithms within the framework of a large-scale, leakage-controlled experimental pipeline consisting of five repeats. It is shown that the feature selection method has much more influence on forecast accuracy than the choice of the machine learning algorithm. Features selected by the RFFI algorithm (Global Horizontal Radiation, Performance Ratio, Wind Direction, and Tilted Global Radiation) allowed reaching a mean R² of more than 0.85 for Random Forest, MLP, and Optimized RBF-SVR models, while PC and RFE resulted in R² lower than 0.60 for all models. As far as model efficiency is concerned, Random Forest showed the best result (mean R²=0.8960 ± 0.0008), slightly better than MLP and Optimized RBF-SVR. The latter algorithm proved to be significantly faster in training (mean 21.25 sec) than Random Forest (mean 35.78 sec).</p><p>This shows that the use of appropriate feature selection methods becomes an important parameter for obtaining accurate forecasting rather than the selection of an appropriate regression model, and that the use of a rigorous and leakage-free feature selection process becomes necessary in order to achieve good forecasting results that can be generalized to real-world application scenarios where there are no electrical telemetry data concurrent with the forecasting target. The suggested framework represents an efficient way of performing short-term forecasting of solar energy output through the use of a limited number of easy-to-obtain variables related to weather and system conditions. Further research work would be devoted to testing the effectiveness of this framework on the basis of other DKASC installations and types of photovoltaic systems using repeated or cross-validated evaluation with confidence intervals, development of a multi-step forecast using this framework, and combination of Random Forest and kernel approaches.</p></sec></body><back><sec sec-type="data-availability"><title>Data Availability</title><p>The dataset used in this study is publicly available from the Desert Knowledge Australia Solar Centre (http://dkasolarcentre.com.au). Processed data and code will be made available by the corresponding author upon reasonable request.</p></sec><bio><title>Biography</title><p><bold>Farwa Nawaz</bold> is a lecturer in the Faculty of Computing at Riphah International University, Islamabad, Pakistan. Her research interests include machine learning, deep learning, uncertainty-aware artificial intelligence, and prediction modeling for renewable energy and biomedicine, specifically medical image analysis and risk prediction. She can be contacted at email: farwa.nawaz@riphah.edu.pk.</p><p><bold>Faria Karamat</bold> is currently serving as a Lecturer at the Faculty of Computing of Riphah International University, Islamabad, Pakistan. She is interested in the field of data science, machine learning, federated learning, privacy-preserving artificial intelligence, and distributed intelligent systems and their application in healthcare and biomedical informatics. She can be contacted at email: faria.karamat@riphah.edu.pk.</p></bio><ref-list><title>References</title><ref id="BIBR-1"><element-citation publication-type="journal"><article-title>Global energy transformation: A roadmap to 2050</article-title><source>IRENA</source><person-group person-group-type="author"><name><surname>Gielen</surname><given-names>D.</given-names></name><name><surname>Gorini</surname><given-names>R.</given-names></name><name><surname>Wagner</surname><given-names>N.</given-names></name><name><surname>Leme</surname><given-names>R.</given-names></name><name><surname>Gutierrez</surname><given-names>L.</given-names></name><name><surname>Prakash</surname><given-names>G.</given-names></name><name><surname>Asmelash</surname><given-names>E.</given-names></name><name><surname>Janeiro</surname><given-names>L.</given-names></name><name><surname>Gallina</surname><given-names>G.</given-names></name><name><surname>Vale</surname><given-names>G.</given-names></name></person-group><year>2019</year><page-range>978-92-9260-121-8</page-range></element-citation></ref><ref id="BIBR-2"><element-citation publication-type="book"><article-title>Green software engineering development paradigm: an approach to a sustainable renewable energy future</article-title><source>Advancing software engineering through AI, federated learning, and large language models: IGI Global</source><person-group person-group-type="author"><name><surname>Matthew</surname><given-names>U.O.</given-names></name><name><surname>Asuni</surname><given-names>O.</given-names></name><name><surname>Fatai</surname><given-names>L.O.</given-names></name></person-group><year>2024</year><page-range>281-294,</page-range><pub-id pub-id-type="doi">10.4018/979-8-3693-3502-4.ch018</pub-id></element-citation></ref><ref id="BIBR-3"><element-citation publication-type="journal"><article-title>Feasibility of future transition to 100% renewable energy: Recent progress, policies, challenges, and perspectives</article-title><source>Journal of Cleaner Production</source><issue>143942</issue><person-group person-group-type="author"><name><surname>Al-Shetwi</surname><given-names>A.Q.</given-names></name><name><surname>Abidin</surname><given-names>I.Z.</given-names></name><name><surname>Mahafzah</surname><given-names>K.A.</given-names></name><name><surname>Hannan</surname><given-names>M.A.</given-names></name></person-group><year>2024</year><pub-id pub-id-type="doi">10.1016/j.jclepro.2024.143942</pub-id></element-citation></ref><ref id="BIBR-4"><element-citation publication-type="journal"><article-title>Machine learning based solar photovoltaic power forecasting: A review and comparison</article-title><volume>11</volume><person-group person-group-type="author"><name><surname>Gaboitaolelwe</surname><given-names>J.</given-names></name><name><surname>Zungeru</surname><given-names>A.M.</given-names></name><name><surname>Yahya</surname><given-names>A.</given-names></name><name><surname>Lebekwe</surname><given-names>C.K.</given-names></name><name><surname>Vinod</surname><given-names>D.N.</given-names></name><name><surname>Salau</surname><given-names>A.O.J.I.A.</given-names></name></person-group><year>2023</year><page-range>40820-40845,</page-range><pub-id pub-id-type="doi">10.1109/ACCESS.2023.3270041</pub-id></element-citation></ref><ref id="BIBR-5"><element-citation publication-type="journal"><article-title>Solar energy potential in Pakistan: A review</article-title><source>Proceedings of the Pakistan Academy of Sciences: B. Life and Environmental Sciences</source><volume>61</volume><issue>1</issue><person-group person-group-type="author"><name><surname>Muhammadi</surname><given-names>A.</given-names></name><name><surname>Wasib</surname><given-names>M.</given-names></name><name><surname>Muhammadi</surname><given-names>S.</given-names></name><name><surname>Ahmed</surname><given-names>S.Riaz</given-names></name><name><surname>Lahori</surname><given-names>A.H.</given-names></name><name><surname>Vambol</surname><given-names>S.</given-names></name><name><surname>Trush</surname><given-names>O.</given-names></name></person-group><year>2024</year><page-range>1-10,</page-range><pub-id pub-id-type="doi">10.53560/PPASB(61-1)931</pub-id></element-citation></ref><ref id="BIBR-6"><element-citation publication-type="journal"><article-title>Recent developments in PV/wind hybrid renewable energy systems: A review</article-title><source>Energy Systems</source><volume>17</volume><person-group person-group-type="author"><name><surname>Bhimaraju</surname><given-names>A.</given-names></name><name><surname>Mahesh</surname><given-names>A.</given-names></name></person-group><year>2026</year><page-range>1303-1345,</page-range><pub-id pub-id-type="doi">10.1007/s12667-024-00679-3</pub-id></element-citation></ref><ref id="BIBR-7"><element-citation publication-type="journal"><article-title>Feasibility of future transition to 100% renewable energy: Recent progress, policies, challenges, and perspectives</article-title><source>Journal of Cleaner Production</source><issue>143942</issue><person-group person-group-type="author"><name><surname>Al-Shetwi</surname><given-names>A.Q.</given-names></name><name><surname>Abidin</surname><given-names>I.Z.</given-names></name><name><surname>Mahafzah</surname><given-names>K.A.</given-names></name><name><surname>Hannan</surname><given-names>M.A.</given-names></name></person-group><year>2024</year><pub-id pub-id-type="doi">10.1016/j.jclepro.2024.143942</pub-id></element-citation></ref><ref id="BIBR-8"><element-citation publication-type="journal"><article-title>A review of deep learning for renewable energy forecasting</article-title><volume>198</volume><person-group person-group-type="author"><name><surname>Wang</surname><given-names>H.</given-names></name><name><surname>Lei</surname><given-names>Z.</given-names></name><name><surname>Zhang</surname><given-names>X.</given-names></name><name><surname>Zhou</surname><given-names>B.</given-names></name><name><surname>Peng</surname><given-names>J.J.E.C.</given-names></name><name name-style="given-only"><given-names>Management</given-names></name></person-group><year>2019</year><page-range>111799,</page-range><pub-id pub-id-type="doi">10.1016/j.enconman.2019.111799</pub-id></element-citation></ref><ref id="BIBR-9"><element-citation publication-type="journal"><article-title>Quantifying the accelerated diffusion and cost savings of global solar photovoltaic supply chains</article-title><source>iScience</source><volume>28</volume><issue>1, Art. no. 111610</issue><person-group person-group-type="author"><name><surname>Chen</surname><given-names>Z.</given-names></name><name><surname>Gu</surname><given-names>B.</given-names></name><name><surname>Yu</surname><given-names>D.</given-names></name><name><surname>Wang</surname><given-names>C.</given-names></name></person-group><year>2025</year><pub-id pub-id-type="doi">10.1016/j.isci.2024.111610</pub-id></element-citation></ref><ref id="BIBR-10"><element-citation publication-type="journal"><article-title>The analysis on photovoltaic electricity generation status, potential and policies of the leading countries in solar energy</article-title><volume>15</volume><issue>1</issue><person-group person-group-type="author"><name><surname>Dincer</surname><given-names>F.J.R.</given-names></name><name><surname>reviews</surname></name></person-group><year>2011</year><page-range>713-720,</page-range><pub-id pub-id-type="doi">10.1016/j.rser.2010.09.026</pub-id></element-citation></ref><ref id="BIBR-11"><element-citation publication-type="journal"><article-title>Forecasting energy consumption with a novel ensemble deep learning framework</article-title><source>Journal of Building Engineering</source><issue>110452</issue><person-group person-group-type="author"><name><surname>Shojaei</surname><given-names>T.</given-names></name><name><surname>Mokhtar</surname><given-names>A.</given-names></name></person-group><year>2024</year><pub-id pub-id-type="doi">10.1016/j.jobe.2024.110452</pub-id></element-citation></ref><ref id="BIBR-12"><element-citation publication-type="journal"><article-title>A review on the prediction of building energy consumption</article-title><volume>16</volume><issue>6</issue><person-group person-group-type="author"><name><surname>Zhao</surname><given-names>H.-x</given-names></name><name><surname>Magoulès</surname><given-names>F.J.R.</given-names></name><name><surname>Reviews</surname><given-names>S.E.</given-names></name></person-group><year>2012</year><page-range>3586-3592,</page-range><pub-id pub-id-type="doi">10.1016/j.rser.2012.02.049</pub-id></element-citation></ref><ref id="BIBR-13"><element-citation publication-type="journal"><article-title>Analysis on intelligent machine learning enabled with meta-heuristic algorithms for solar irradiance prediction</article-title><volume>15</volume><issue>1</issue><person-group person-group-type="author"><name><surname>Vaisakh</surname><given-names>T.</given-names></name><name><surname>Jayabarathi</surname><given-names>R.J.E.I.</given-names></name></person-group><year>2022</year><page-range>235-254,</page-range><pub-id pub-id-type="doi">10.1007/s12065-020-00505-6</pub-id></element-citation></ref><ref id="BIBR-14"><element-citation publication-type="journal"><article-title>Optimal operation of distributed generations in micro‐grids under uncertainties in load and renewable power generation using heuristic algorithm</article-title><volume>9</volume><issue>8</issue><person-group person-group-type="author"><name><surname>Nikmehr</surname><given-names>N.</given-names></name><name><surname>Najafi‐Ravadanegh</surname><given-names>S.J.I.</given-names></name></person-group><year>2015</year><page-range>982-990,</page-range><pub-id pub-id-type="doi">10.1049/iet-rpg.2014.0357</pub-id></element-citation></ref><ref id="BIBR-15"><element-citation publication-type="journal"><article-title>Solar power generation prediction based on deep learning</article-title><source>Sustain. Energy Technol. Assess</source><volume>47</volume><person-group person-group-type="author"><name><surname>Chang</surname><given-names>R.</given-names></name><name><surname>Bai</surname><given-names>L.</given-names></name><name><surname>Hsu</surname><given-names>C.H.</given-names></name></person-group><year>2021</year><page-range>101354,</page-range><pub-id pub-id-type="doi">10.1016/j.seta.2021.101354</pub-id></element-citation></ref><ref id="BIBR-16"><element-citation publication-type="journal"><article-title>Prediction of solar irradiance and photovoltaic solar energy product based on cloud coverage estimation using machine learning methods</article-title><source>Atmosphere</source><volume>12</volume><person-group person-group-type="author"><name><surname>Park</surname><given-names>S.</given-names></name><name><surname>Kim</surname><given-names>Y.</given-names></name><name><surname>Ferrier</surname><given-names>N.J.</given-names></name><name><surname>Collis</surname><given-names>S.M.</given-names></name><name><surname>Sankaran</surname><given-names>R.</given-names></name><name><surname>Beckman</surname><given-names>P.H.</given-names></name></person-group><year>2021</year><page-range>395,</page-range><pub-id pub-id-type="doi">10.3390/atmos12030395</pub-id></element-citation></ref><ref id="BIBR-17"><element-citation publication-type="journal"><article-title>Prediction of solar energy guided by Pearson correlation using machine learning</article-title><source>Energy</source><volume>224</volume><person-group person-group-type="author"><name><surname>Jebli</surname><given-names>I.</given-names></name><name><surname>Belouadha</surname><given-names>F.Z.</given-names></name><name><surname>Kabbaj</surname><given-names>M.I.</given-names></name><name><surname>Tilioua</surname><given-names>A.</given-names></name></person-group><year>2021</year><page-range>120109,</page-range><pub-id pub-id-type="doi">10.1016/j.energy.2021.120109</pub-id></element-citation></ref><ref id="BIBR-18"><element-citation publication-type="journal"><article-title>Feature selection for energy system modeling: Identification of relevant time series information</article-title><source>Energy AI</source><volume>4</volume><person-group person-group-type="author"><name><surname>Müller</surname><given-names>I.M.</given-names></name></person-group><year>2021</year><page-range>100057,</page-range><pub-id pub-id-type="doi">10.1016/j.egyai.2021.100057</pub-id></element-citation></ref><ref id="BIBR-19"><element-citation publication-type="journal"><article-title>Feature-selective ensemble learning-based long-term regional PV generation forecasting</article-title><source>IEEE Access</source><volume>8</volume><person-group person-group-type="author"><name><surname>Eom</surname><given-names>H.</given-names></name><name><surname>Son</surname><given-names>Y.</given-names></name><name><surname>Choi</surname><given-names>S.</given-names></name></person-group><year>2020</year><page-range>54620-54630,</page-range><pub-id pub-id-type="doi">10.1109/ACCESS.2020.2981819</pub-id></element-citation></ref><ref id="BIBR-20"><element-citation publication-type="journal"><article-title>Innovative framework for accurate and transparent forecasting of energy consumption: A fusion of feature selection and interpretable machine learning</article-title><source>Appl. Energy</source><volume>366</volume><person-group person-group-type="author"><name><surname>Eskandari</surname><given-names>H.</given-names></name><name><surname>Saadatmand</surname><given-names>H.</given-names></name><name><surname>Ramzan</surname><given-names>M.</given-names></name><name><surname>Mousapour</surname><given-names>M.</given-names></name></person-group><year>2024</year><page-range>123314,</page-range><pub-id pub-id-type="doi">10.1016/j.apenergy.2024.123314</pub-id></element-citation></ref><ref id="BIBR-21"><element-citation publication-type="journal"><article-title>Forecasting solar irradiance with geographical considerations: integrating feature selection and learning algorithms</article-title><volume>8</volume><issue>5</issue><person-group person-group-type="author"><name><surname>Soleymani</surname><given-names>S.</given-names></name><name><surname>S. S. Talebi</surname><given-names>A.J.A.J.</given-names></name></person-group><year>2024</year></element-citation></ref><ref id="BIBR-22"><element-citation publication-type="journal"><article-title>Renewable energy prediction: A novel short-term prediction model of photovoltaic output power</article-title><source>J. Clean. Prod</source><volume>228</volume><person-group person-group-type="author"><name><surname>Li</surname><given-names>L.L.</given-names></name><name><surname>Wen</surname><given-names>S.Y.</given-names></name><name><surname>Tseng</surname><given-names>M.L.</given-names></name><name><surname>Wang</surname><given-names>C.S.</given-names></name></person-group><year>2019</year><page-range>359-375,</page-range><pub-id pub-id-type="doi">10.1016/j.jclepro.2019.04.331</pub-id></element-citation></ref><ref id="BIBR-23"><element-citation publication-type="journal"><article-title>Short-term PV power forecasting using hybrid GASVM technique</article-title><source>Renew. Energy</source><volume>140</volume><person-group person-group-type="author"><name><surname>VanDeventer</surname><given-names>W.</given-names></name><name><surname>Jamei</surname><given-names>E.</given-names></name><name><surname>Thirunavukkarasu</surname><given-names>G.S.</given-names></name><name><surname>Seyedmahmoudian</surname><given-names>M.</given-names></name><name><surname>Soon</surname><given-names>T.K.</given-names></name><name><surname>Horan</surname><given-names>B.</given-names></name><name><surname>Mekhilef</surname><given-names>S.</given-names></name><name><surname>Stojcevski</surname><given-names>A.</given-names></name></person-group><year>2019</year><page-range>367-379,</page-range><pub-id pub-id-type="doi">10.1016/j.renene.2019.02.087</pub-id></element-citation></ref><ref id="BIBR-24"><element-citation publication-type="journal"><article-title>Photovoltaic power forecasting based on a support vector machine with improved ant colony optimization</article-title><source>J. Clean. Prod</source><volume>277</volume><person-group person-group-type="author"><name><surname>Pan</surname><given-names>M.</given-names></name><name><surname>Li</surname><given-names>C.</given-names></name><name><surname>Gao</surname><given-names>R.</given-names></name><name><surname>Huang</surname><given-names>Y.</given-names></name><name><surname>You</surname><given-names>H.</given-names></name><name><surname>Gu</surname><given-names>T.</given-names></name><name><surname>Qin</surname><given-names>F.</given-names></name></person-group><year>2020</year><page-range>123948,</page-range><pub-id pub-id-type="doi">10.1016/j.jclepro.2020.123948</pub-id></element-citation></ref><ref id="BIBR-25"><element-citation publication-type="journal"><article-title>A hybrid RBF neural network based model for day-ahead prediction of photovoltaic plant power output</article-title><source>Front. Energy Res</source><volume>11</volume><person-group person-group-type="author"><name><surname>Zhang</surname><given-names>Q.</given-names></name><name><surname>Tang</surname><given-names>N.</given-names></name><name><surname>Lu</surname><given-names>J.</given-names></name><name><surname>Wang</surname><given-names>W.</given-names></name><name><surname>Wu</surname><given-names>L.</given-names></name><name><surname>Kuang</surname><given-names>W.</given-names></name></person-group><year>2024</year><page-range>1338195,</page-range><pub-id pub-id-type="doi">10.3389/fenrg.2023.1338195</pub-id></element-citation></ref><ref id="BIBR-26"><element-citation publication-type="journal"><article-title>A hybrid deep learning model for short-term PV power forecasting</article-title><source>Applied Energy</source><issue>114216</issue><person-group person-group-type="author"><name><surname>Li</surname><given-names>P.</given-names></name><name><surname>Zhou</surname><given-names>K.</given-names></name><name><surname>Lu</surname><given-names>X.</given-names></name><name><surname>Yang</surname><given-names>S.</given-names></name></person-group><year>2020</year><pub-id pub-id-type="doi">10.1016/j.apenergy.2019.114216</pub-id></element-citation></ref><ref id="BIBR-27"><element-citation publication-type="journal"><article-title>A short-term forecasting method for photovoltaic power generation based on the TCN-ECANet-GRU hybrid model</article-title><source>Scientific Reports</source><volume>14</volume><issue>1, Art. no. 6744</issue><person-group person-group-type="author"><name><surname>Xiang</surname><given-names>X.</given-names></name><name><surname>Li</surname><given-names>X.</given-names></name><name><surname>Zhang</surname><given-names>Y.</given-names></name><name><surname>Hu</surname><given-names>J.</given-names></name></person-group><year>2024</year><pub-id pub-id-type="doi">10.1038/s41598-024-56751-6</pub-id></element-citation></ref><ref id="BIBR-28"><element-citation publication-type="journal"><article-title>A novel convolutional neural net architecture based on incorporating meteorological variable inputs into ultra-short-term photovoltaic power forecasting</article-title><source>Sustainability</source><volume>16</volume><issue>7, Art. no. 2786</issue><person-group person-group-type="author"><name><surname>Ren</surname><given-names>X.</given-names></name><name><surname>Zhang</surname><given-names>F.</given-names></name><name><surname>Yan</surname><given-names>J.</given-names></name><name><surname>Liu</surname><given-names>Y.</given-names></name></person-group><year>2024</year><pub-id pub-id-type="doi">10.3390/su16072786</pub-id></element-citation></ref><ref id="BIBR-29"><element-citation publication-type="journal"><article-title>Solar irradiance forecasting using temporal fusion transformers</article-title><source>International Journal of Energy Research</source><issue>3534500</issue><person-group person-group-type="author"><name><surname>Alorf</surname><given-names>A.</given-names></name><name><surname>Khan</surname><given-names>M.U.G.</given-names></name></person-group><year>2025</year><pub-id pub-id-type="doi">10.1155/er/3534500</pub-id></element-citation></ref><ref id="BIBR-30"><element-citation publication-type="journal"><article-title>Solar power generation forecasting by a new hybrid cascaded extreme learning method with maximum relevance interaction gain feature selection</article-title><source>Energy Convers. Manage</source><volume>298</volume><person-group person-group-type="author"><name><surname>Memarzadeh</surname><given-names>G.</given-names></name><name><surname>Keynia</surname><given-names>F.</given-names></name></person-group><year>2023</year><page-range>117763,</page-range><pub-id pub-id-type="doi">10.1016/j.enconman.2023.117763</pub-id></element-citation></ref><ref id="BIBR-31"><element-citation publication-type="journal"><article-title>Predicting the energy output of hybrid PV–wind renewable energy system using feature selection technique for smart grids</article-title><source>Energy Rep</source><volume>7</volume><person-group person-group-type="author"><name><surname>Qadir</surname><given-names>Z.</given-names></name><name><surname>Khan</surname><given-names>S.I.</given-names></name><name><surname>Khalaji</surname><given-names>E.</given-names></name><name><surname>Munawar</surname><given-names>H.S.</given-names></name><name><surname>Al-Turjman</surname><given-names>F.</given-names></name><name><surname>Mahmud</surname><given-names>M.P.</given-names></name><name><surname>Kouzani</surname><given-names>A.Z.</given-names></name><name><surname>Le</surname><given-names>K.</given-names></name></person-group><year>2021</year><page-range>8465-8475,</page-range><pub-id pub-id-type="doi">10.1016/j.egyr.2021.01.018</pub-id></element-citation></ref><ref id="BIBR-32"><element-citation publication-type="journal"><article-title>Parallel boosting neural network with mutual information for day-ahead solar irradiance forecasting</article-title><source>Sci. Rep</source><volume>15</volume><issue>1</issue><person-group person-group-type="author"><name><surname>Ahmed</surname><given-names>U.</given-names></name><name><surname>Mahmood</surname><given-names>A.</given-names></name><name><surname>Khan</surname><given-names>A.R.</given-names></name><name><surname>Kuhlmann</surname><given-names>L.</given-names></name><name><surname>Alimgeer</surname><given-names>K.S.</given-names></name><name><surname>Razzaq</surname><given-names>S.</given-names></name><name><surname>Aziz</surname><given-names>I.</given-names></name><name><surname>Hammad</surname><given-names>A.</given-names></name></person-group><year>2025</year><page-range>11642,</page-range><pub-id pub-id-type="doi">10.1038/s41598-025-95891-1</pub-id></element-citation></ref><ref id="BIBR-33"><element-citation publication-type="journal"><article-title>Solar irradiance forecasting using dynamic ensemble selection</article-title><source>Applied Sciences</source><volume>12</volume><issue>7, Art. no. 3510</issue><person-group person-group-type="author"><name><surname>O. Santos</surname><given-names>P.S.G.de Mattos Neto</given-names></name><name><surname>Oliveira</surname><given-names>J.F.L.</given-names></name><name><surname>Siqueira</surname><given-names>H.V.</given-names></name><name><surname>Barchi</surname><given-names>T.M.</given-names></name><name><surname>Lima</surname><given-names>A.R.</given-names></name><name><surname>Madeiro</surname><given-names>F.</given-names></name><name><surname>Dantas</surname><given-names>D.A.P.</given-names></name><name><surname>Converti</surname><given-names>A.</given-names></name><name><surname>Pereira</surname><given-names>A.C.</given-names></name><name><surname>Melo Filho</surname><given-names>J.B.</given-names></name><name><surname>Marinho</surname><given-names>M.H.N.</given-names></name></person-group><year>2022</year><pub-id pub-id-type="doi">10.3390/app12073510</pub-id></element-citation></ref><ref id="BIBR-34"><element-citation publication-type="journal"><article-title>Machine learning techniques for daily solar energy prediction and interpolation using numerical weather models</article-title><source>Concurrency and Computation: Practice and Experience</source><volume>28</volume><issue>4</issue><person-group person-group-type="author"><name><surname>Martin</surname><given-names>R.</given-names></name><name><surname>Aler</surname><given-names>R.</given-names></name><name><surname>Valls</surname><given-names>J.M.</given-names></name><name><surname>Galván</surname><given-names>I.M.</given-names></name></person-group><year>2016</year><page-range>1261-1274,</page-range><pub-id pub-id-type="doi">10.1002/cpe.3631</pub-id></element-citation></ref></ref-list></back></article>