aub logo
  • AUB Admission is ongoing for Fall - 2026 (September to December)  | To Apply Click Here  *** www.aub.ac.bd is our only website. All other websites in the name of AUB are fake. So everyone is warned not to be deceived. 
aub logo white
Apply Now

Contact us

+8801678664413-19

aub_admin July 23, 2026 23 Views

GridTradeRL: a multi-agent deep reinforcement learning framework for decentralized peer-to-peer energy trading in smart grid prosumer networks

Authors & Affiliations

Md. Irfan Hossain1 & Mithila Arman1
1 Department of Computer Science and Engineering, BRAC University, Dhaka, 1212, Bangladesh

Md Shahedul Islam2
2 Department of Electrical and Electronic Engineering, University of Dhaka, Dhaka, Bangladesh

Hridya Sudeepon Roy3
3 Department of Electrical and Electronic Engineering, Bangladesh University of Engineering and Technology, Dhaka, Bangladesh

Sagor Ahmed4
4 Department of Computer Science and Engineering, United International University, Dhaka, Bangladesh

Abdullah Rakib Akand5
5 Department of Computer Science and Engineering, Asian University of Bangladesh, Dhaka, Bangladesh

Sayed Esmat Mirja6
6 Department of Computer Science and Engineering, University of Chittagong, Chittagong, Bangladesh

Samuel Kim7
7 Department of Computer Science, Postdoctoral Fellow, Stanford University, Stanford, CA, USA

Haewon Byeon8
8 Department Of Artificial Intelligence, European Institute for Materials, AI & Technology (EIMATEL), Madrid, Spain

Journal Information

Publisher: Springer Nature

Editorial History

Received 15 May 2026
Accepted 16 June 2026
Published 23 July 2026

Digital Object Identifier (DOI)

https://doi.org/10.1038/s41598-026-58710-9

Abstract

The accelerating penetration of distributed photovoltaic (PV) generation and battery storage at the residential edge of the distribution network is converting passive consumers into active prosumers, exposing the limits of the unidirectional buy-from-utility / sell-to-utility settlement model. Peer-to-peer (P2P) energy trading promises to internalize the bid-ask spread of time-of-use (ToU) tariffs within local communities, but its decentralized, non-stationary, and partially observable nature makes hand-engineered control strategies brittle. This manuscript presents GridTradeRL, a multi-agent deep reinforcement learning framework for decentralized P2P energy trading in smart-grid prosumer networks. We formulate the trading problem as a partially observable stochastic game over a heterogeneous community of N prosumers and solve it with a multi-agent extension of Twin-Delayed Deep Deterministic Policy Gradient (MA-TD3) under the centralized-training, decentralized-execution (CTDE) paradigm. Each agent learns an independent actor that selects a continuous battery-dispatch and bid-aggressiveness action from local observations alone, while training is coordinated by a centralized twin critic with full access to the joint observation-action space. The local market clears via a uniform-price double auction at each hour. On a simulated 8-prosumer community with time-of-use tariffs, the trained policy reduces net daily community electricity expenditure by 42.7% versus a no-trade baseline and by 24.2% versus a rule-based heuristic, while routing 18.3% of all energy through bilateral peer-to-peer trades. Against learned baselines, it improves on independent DDPG by 29.7% and on MADDPG by 10.4%, and it attains performance within 11.5% of a perfect-foresight model-predictive-control oracle without requiring forecasts or disclosure of private prosumer data. We further establish convergence stability over a 2000-episode, five-seed training protocol; characterize sensitivity to the reward-shaping coefficients, robustness to observational noise and communication dropout, and scalability to communities of up to 64 prosumers; and prove that the midpoint clearing rule distributes surplus equitably (Gini coefficient 0.039). A fast 350-episode configuration that attains 97.8% of converged performance completes in 224 s on a single CPU core in pure NumPy without any GPU acceleration, demonstrating the suitability of the framework for nightly retraining on edge gateways. The full source code, trained policies, and an interactive simulation dashboard accompany this manuscript.

Keywords: Peer-to-peer energy trading, Multi-agent reinforcement learning, Smart grids, Prosumers, Deep learning, GridTradeRL.

SCIENTIFIC REPORTS | NATURE PORTFOLIO | 2026