Document Type : Original Article
Authors
1
Department of Economics, Faculty of Islamic Studies and Economics, Imam Sadiq University
2
Department of Financial Economics, Faculty of Islamic Studies and Economics, Imam Sadiq University, Tehran, Iran
Abstract
Introduction: The main objective of this study is to test weak-form efficiency in capital markets, using the Tehran Stock Exchange composite index and the S&P 500. The main question is whether returns in these markets are consistent with weak-form efficiency when tested against past price and volume information. Out-of-sample return prediction provides the empirical testing framework, with a zero-return random walk as the statistical benchmark. Further analyses ask whether nonlinear models improve on linear regression and whether Tehran trading-condition indicators add information. These supplementary comparisons support the efficiency test; the indicators neither measure an order book nor isolate the effect of market rules. The design therefore treats statistical predictability as evidence about the information contained in market history, while keeping economic efficiency, risk adjustment, transaction costs, and implementability conceptually separate.
Methods: Daily index data are used to predict the next day’s log return. The standard features include current returns, three return lags, ten-day volatility and five-day mean volume. For Tehran, a second set includes distance from a hypothetical band, its five-day mean, proximity to the band, positive or negative returns with low relative volume, and a period indicator. The combined set contains both groups. Linear regression and LightGBM are compared with the zero-return benchmark in five expanding-window folds. Training and validation come before each test block, with one-row gaps at the selection boundaries. Preprocessing uses training data only. Performance is assessed with forecast errors, test-sample R-squared and direction accuracy. Paired squared-error losses are compared by a Diebold–Mariano test with a standard error robust to changing variance and serial dependence. Checks of result stability vary the volume threshold and initial training share and use a subsample before the period boundary. A separate analysis compares tuned and default LightGBM settings to assess sensitivity to hyperparameters. Models are re-estimated as the training window expands. Performance is summarized both by averaging the five test-fold metrics and by pooling test errors, so fold averages are distinguished from aggregate results.
Findings: For Tehran with standard features, mean fold RMSE is 0.011300 for the random walk, 0.010504 for linear regression and 0.010709 for LightGBM. Both fitted models improve on the benchmark, with HAC p-values of 0.0002 and 0.0014, respectively. Their difference is not significant in the full-sample baseline. Neither model shows a significant gain over the benchmark for the S&P 500. In Tehran, gains from tuned models remain under the tested volume thresholds and in the subsample before the boundary. However, default LightGBM settings lead to worse performance. This comparison demonstrates sensitivity to hyperparameters, not model robustness. Stability under the tested volume thresholds and subsample choices does not establish stability across model settings. Feature-attribution results are used to describe model behavior rather than to establish causal effects or incremental forecasting value.
Conclusions: The zero-return random walk is a statistical forecasting benchmark, not a complete test of weak-form market efficiency. Market efficiency does not require a zero expected return. The absence of significant gains for the S&P 500 therefore does not prove efficiency, while lower forecast errors in Tehran do not by themselves establish inefficiency or net trading profits. LightGBM performance is sensitive to hyperparameters; the default-setting comparison does not establish model robustness. No consistent advantage of nonlinear complexity is found, and a nonsignificant difference does not establish model equivalence. The cross-market comparison is descriptive and cannot identify causal effects of market rules or structure. Accordingly, the evidence supports a cautious, benchmark-dependent interpretation: historical data contain useful predictive content for Tehran under selected specifications, whereas the S&P 500 results do not provide comparable evidence within this sample and research design.
Keywords