Research interests

Bayesian system availability with variable hardware component in presence of human and system errors
P Mostert, A Bekker (University of Pretoria)
Availability analysis of a various multi-unit systems of two or more configurations is normally subjected to unit failures, common cause failures and human error. Various distributions are normally assumed for repair and hardware failures. A Bayesian approach is adopted with different types of priors assumed for the unknown parameters to evaluate the availability of the system. MCMC methods are used to derive realisations of the posterior distribution for the steady-state availability. The performance of the estimators is evaluated in terms of bias, mean squared error and frequentist coverage by means of a simulation study to highlight the importance of the study.
Biplot visualisation for symbolic data
Unlike traditional data analysis, where each observation is described by a single value per variable, Symbolic Data Analysis accommodates more complex data structures such as intervals and histograms, representing intrinsic variability within each unit. New methods are being developed to construct biplots for this richer type of data.
Biplots for compositional data
S Lubbe, P Sebatjane (UNISA), NJ le Roux (Emeritus Professor)
Since compositional data is not observed in a Euclidean space but in the simplex space, ordinary multivariate data analysis methods cannot be applied directly to this data. This project investigates how to construct biplots, analogous to biplots for multivariate data, but taking the specific nature of compositional data into account.
Campanometry
T de Wet, JL Teugels (KU Leuven), PJU van Deventer (Retired)
Campanology refers to the study of bells, bell-casting, and bell-ringing and has been well studied. Because the sound of a bell is central to this field, campanology has traditionally been closely linked to physics, particularly acoustics. By contrast, the quantitative and statistical study of bells and their properties, referred to as campanometry, is relatively recent. Of particular interest is the measurement and statistical analysis of the different partials of bells and carillons. Since 2008, physical and acoustical data, relevant historical information and photographs on bells in the Western Cape Province have been collected and stored in the Stellenbosch University Digital Library at https://digital.lib.sun.ac.za/handle/10019.2/3975. This database serves as an important resource for both statistical and historical research.
Classification of rot in wine grapes
S Lubbe, NJ le Roux (Emeritus Professor), M Cornelissen, Namaqua Wines, H Nieuwoudt (South African Grape and Wine Research Institute at SU)
Since grape rot can have a severe effect on the quality of wine, it is important to assess the rot level in wine grapes delivered to a cellar. This project make use of infrared spectra and advanced multivariate data analysis methods to classify the extent of rot present in grapes.
Digital soil mapping
Digital soil mapping focuses on the spatial prediction and analysis of soil properties using statistical, geostatistical, and machine learning approaches together with environmental covariates and remotely sensed data. Research in this area investigates predictive modelling of soil characteristics such as soil texture, soil depth, carbon content, and other environmental variables across multiple spatial scales. Particular emphasis is placed on machine learning, uncertainty quantification, spatial validation, compositional data analysis, and interpretable modelling approaches to support sustainable land management, agriculture, and environmental decision-making.
Fairness in machine learning
T Sandrock, J Nienkemper-Swanepoel, A Bhatti
The use of machine learning algorithms to automate decisions is a prevalent practice. Consequently, studying the fairness of these algorithms before implementation is crucial. The effect of missing data mechanisms and imputation methods on the fairness of such algorithms is currently being studied.
Functional uniform priors in a setting of nonlinear mixed effects models
P Mostert, E Lesaffre (KU Leuven), F Mercier (Roche), D Dejardin (Roche)
This research investigates Bayesian inference for non-linear and non-linear mixed models, with particular emphasis on the choice of non-informative priors. Traditional uniform priors may unintentionally favour certain non-linear model shapes and introduce bias. We therefore study functional uniform priors (FUPs), which are parametrisation invariant and satisfy the likelihood principle. Applications include Emax and tumour growth inhibition models in pharmaceutical and oncology research, with collaborations involving Roche Pharmaceuticals in Basel, Switzerland.
Interpretable machine learning
Interpretable machine learning focuses on developing and applying methods that improve the transparency, understanding, and trustworthiness of machine learning models. Research in this area investigates how complex predictive models can be explained both globally and locally using approaches such as partial dependence plots, accumulated local effects, surrogate models, feature importance measures, LIME, and SHAP. Applications include environmental science, digital soil mapping, finance, and other domains where understanding model behaviour, uncertainty, and decision-making processes is essential.
Local Scale Estimation from LULU Residuals under Gaussian and Impulsive Noise
T de Wet, R Buys, CH Rohwer (Retired, SU Mathematical Sciences)
The present research investigates an estimator arising from the theory of LULU smoothers. These were developed as nonlinear smoothers for sequences, motivated by the need to remove impulsive components while preserving local structural features. The statistical properties of these LULU-based scale estimators are developed. Consistency is shown using the finite-dependence structure of the residual sequence and the corresponding asymptotic normal approximation is obtained. The theoretical analysis shows that the first LULU residual contains enough information to estimate the Gaussian scale, despite the nonlinear and dependent form of the residual amplitudes. The resulting estimators are compared with established scale estimators through simulations and application to real data.
Model-Based Statistical Frameworks for Complex Directional Data
This research develops advanced statistical frameworks for data observed as directions, angles, or orientations on curved, non-Euclidean spaces such as circles, spheres, and tori. Combining two complementary methodological threads — directional statistics and model-based clustering — the project constructs flexible distribution families, regression models, and mixture model-based techniques to uncover hidden structure in complex data. A particular emphasis is placed on datasets exhibiting heterogeneity, asymmetry, heavy tails, missing values, or atypical observations, where standard linear methods are inadequate. Applications are drawn from environmental sciences, meteorology, ecology, biomechanics, and health sciences, where directional measurements such as wind direction, ocean currents, and animal movement naturally co-occur with these statistical irregularities.
Multi-label classification
T Sandrock, S Steel, A Muller
Multi-label classification problems arise in scenarios where every observation can be associated with multiple labels simultaneously. Such data possess unique characteristics which result in additional challenges when analysing the data. Recent studies that have been undertaken include the simulation of multi-label classification data, and a proposal for a new tree-based ensemble method for multi-label classification, which exploits label correlations in an efficient way.
Nonparametric estimation of extreme expected shortfall for compound Poisson models
T de Wet, PJ de Jongh (NWU), H Raubenheimer (NWU), J Beirlant (KU Leuven)
This research investigates the nonparametric estimation of extreme Expected Shortfall (ES) for compound Poisson aggregate loss distributions using a multiplier methodology. Compound Poisson models are widely used to model aggregate losses in operational risk, insurance and credit risk, where accurate estimation of extreme tail risk is essential for economic and regulatory capital determination. Building on our previously developed nonparametric multiplier methodology for extreme Value-at-Risk, the research employs techniques from Extreme Value Theory to derive first- and second-order multiplier estimators for extreme ES. The estimators will be evaluated through simulation studies and applications to operational risk loss data. The research aims to develop a practical, robust, and distribution-free framework for estimating extreme ES when conventional parametric approaches are unreliable.
Predictive modelling of compositional data
Predictive modelling of compositional data focuses on statistical and machine learning methods for responses and predictors that represent relative information and proportions. Applications include soil texture, geochemical compositions, microbiome data, financial portfolios, and other multivariate systems constrained to a constant sum. Research in this area investigates compositional data analysis principles, log-ratio transformations, simplex-aware machine learning models, interpretable machine learning, uncertainty quantification, and spatial prediction methodologies. This project has recently published an R package called DirichletRF, a novel machine learning model designed specifically for compositional data.
Software for animated biplots
J Nienkemper-Swanepoel, R Ganey (University of the Witwatersrand)
This project is focused on continuous development of the R package, moveEZ (pronounced move easy). In its current form it allows the dynamic visualisation of principal component analysis (PCA) and canonical variate analysis (CVA) biplots when sequential measurements are available.
Statistical Modelling of Natural Selection Between Populations with Genetic Data
Modern biological evolutionary questions are becoming increasingly complex. This is in part due to the unending disease outbreaks associated with new pathogen strains and the tremendous speed with which easily accessible molecular data are generated. Meaningful analyses of these molecular data in ways that provide answers to important questions, often require creative adaptations of established statistical techniques. This research involves interrogating natural selection as an important theme driving these complexities and creativity, with focus on inter-population analyses. The primary objective is to lead rigorous research into the use of existing and the development of novel statistical models for diverse computational molecular biological studies.
Strategies for missing data
J Nienkemper-Swanepoel, D Cook (Monash University)
Multiple imputation replaces missing observations with multiple plausible values which results in multiple completed data sets for standard analyses. Visualisation of multiple data sets representing the same original set can become cumbersome. This project develops visualisation tools by bridging biplot methodology and data tours with a focus on projections to understand the variation between imputations.
Visualisation of multi-dimensional data and biplots
In 2021 the Centre for Multi-dimensional Data Visualisation (MuViSU) was established within the department. The centre provides a support infrastructure for researchers within the department focussing on multi-dimensional data visualisation.
Visualising emission profiles
J Nienkemper-Swanepoel, R Coetzer (North West University), H Neomagus (North West University), R Ganey (University of the Witwatersrand)
This project considers the emission profiles of heat sources such as coal and wood to quantify cooking efficiency as well as understanding the possible exposure to certain emissions.
Actuarial Science Education
G Slattery
Investigating the success of Stellenbosch University students in terms of progression through the actuarial science programme and exemptions from the examinations of the Actuarial Society of South Africa. Comparing the success of students from different universities in the profession’s examinations.
Alternative investments and fine-wine finance
This research studies fine wine as an investable asset class, focusing on return dynamics, liquidity, and diversification benefits relative to traditional portfolios. It includes constructing and analysing wine indices (such as the SAFW10 South African Fine Wine Index) and assessing performance under different economic regimes. Methods include time-series econometrics, correlation and connectedness analysis, and portfolio optimisation under realistic trading and liquidity constraints. Applications include alternative-asset allocation, performance benchmarking, and understanding how emerging wine regions evolve as financial markets.
Benchmark reform and basis risk
This research focuses on quantitative modelling and risk analysis for benchmark transitions in interest-rate markets, with emphasis on the South African move from JIBAR toward alternative risk-free rates such as ZARONIA. It investigates how liquidity, market microstructure, and policy or event-driven shocks affect the spread and its dynamics over time. Methods include multi-curve term-structure modelling, jump processes, and simulation-based risk measures used in valuation, hedging, and exposure management. Applications include curve construction, discounting/forecasting consistency, and the implications of benchmark reform for banks, asset managers, and regulators.
Model validation, model risk, and governance for quantitative finance
Model validation research studies how quantitative models should be tested, challenged, and documented to ensure they are fit for purpose in production risk and pricing systems. It includes independent replication, sensitivity analysis, benchmarking, back-testing, and assessing numerical robustness and implementation risk. The work emphasizes transparent assumptions, limitations, and controls, aligning technical validation with governance requirements. Applications include validation of structured products tools, curve construction frameworks, and exotic option models, helping institutions manage model risk while maintaining analytical integrity.
Multi-curve term structure modelling and curve construction
Multi-curve modelling studies the joint behaviour of discounting and forwarding curves across tenors, especially in markets where collateral, liquidity, and credit effects matter. Research in this area develops coherent frameworks for building and stress-testing yield curves (including OIS-style curves) and analysing roll-over and funding components embedded in observed rates. Methodologies include affine and factor-based term-structure models, calibration to swap/IRS markets, and diagnostics for curve quality and stability. Applications range from pricing interest-rate derivatives and structured products to risk management, scenario analysis, and regulatory readiness for benchmark transitions.
Realised volatility forecasting and risk forecasting for emerging markets
This research investigates forecasting volatility and risk in emerging markets, where structural breaks, liquidity constraints, and regime shifts often challenge standard models. It compares and improves prominent volatility forecasting approaches - such as heterogeneous autoregressive models, realised GARCH-type frameworks, and related specifications - using rigorous evaluation metrics and robust out-of-sample testing. The work also considers how volatility forecasts translate into practical risk measures and portfolio decisions. Applications include risk budgeting, derivatives trading inputs, and stress-testing in markets with episodic shocks and evolving microstructure.
Rough volatility and short-maturity options (including 0DTE)
Rough-volatility research studies models where volatility exhibits fractal-like, rough behaviour that better matches short-maturity option data than classical diffusions. This area develops parsimonious rough-volatility specifications and efficient calibration procedures suitable for very short maturities, where microstructure effects and steep implied-volatility skews are prominent. Methods include rough Heston-type dynamics, fast numerical schemes for implied-volatility surfaces, and empirical diagnostics using high-frequency or dense option panels. Applications include understanding short-dated risk, explaining extreme skew behaviour, and building tractable models for fast pricing and hedging.
Stochastic volatility and jump-diffusion option pricing
This research investigates option pricing and risk measurement under stochastic volatility and jump dynamics, where markets exhibit volatility clustering, skew, and sudden discontinuities. It focuses on building models that are both theoretically sound and numerically robust, including calibration to market data and sensitivity (Greeks) computation. Techniques include PDE methods, Methods of Lines (MOL), Fourier/cosine expansions, and Monte Carlo simulation with variance reduction. Applications include equity and FX derivatives, exotic options (e.g., barriers), and model risk assessment in trading and banking environments.