Resume

Resume

Sreejeet Maity
Ph.D. Student in Electrical Engineering · North Carolina State University · Raleigh, NC, U.S.A
I develop provably robust finite-sample guarantees for reinforcement learning (RL) under uncertainty and adversarial corruption. My current research interests include corruption-tolerant reinforcement learning, robust policy evaluation, distributed and federated reinforcement learning, and minimax lower bounds that characterize the fundamental limits of robust learning.
Reinforcement Learning Statistical Learning Theory Control Theory Applied Probability Stochastic Approximation Algorithms Optimization Robust Statistics
Collaboration Note. I am always happy to engage with researchers working on broad areas of reinforcement learning, control theory, federated learning, and trustworthy machine learning. If you are interested in discussing possible collaborations or exchanging research ideas, please reach out to .
Prospective Opportunities. I will be entering the academic and industry research job market next year (2027), and I would be grateful to hear about opportunities aligned with my interests in broad areas of robust and safe RL, and reliable decision-making and control. I am also open to postdoctoral opportunities beginning in Fall 2027, as well as opportunities in subsequent academic cycles. Prospective recruiters, search committees, and researchers are warmly welcome to reach out.

Education

North Carolina State University
Advisor: Dr. Aritra Mitra.
Indian Institute of Science, Bangalore
Jadavpur University

Research Experience

Graduate Research / Teaching Assistant
  • Robust optimal policy learning from corrupted and correlated observations. Showed that vanilla Q-Learning is provably fragile under reward corruption, designed robust Bellman-update methods, and established finite-time convergence with matching minimax lower bounds. Results disseminated across ICML 2026, NeurIPS 2025, and IEEE CDC 2024, CDC 2026.
  • Robust policy evaluation under adversarial influences and Markovian data. Developed finite-time theory for robust temporal-difference learning with Markovian noise and function approximation, including upper bounds and near-tight lower bounds. This research is published in AISTATS 2025.
  • Robust federated and multi-agent reinforcement learning. Developed adversarially robust and communication-efficient reinforcement learning algorithms for federated multi-agent settings, including Byzantine-resilient methods with collaborative speedups. Two papers published at ACC 2026.
Research Collaboration · Byzantine-Robust Representation Learning

Developing Byzantine-robust representation learning methods for heterogeneous reinforcement learning agents operating in distinct MDPs with temporally dependent data. The approach decomposes each agent's value function into a shared low-dimensional representation and a personalized local parameter, enabling collaborative learning of common structure while accommodating differences in transitions, rewards, policies, and value functions without persistent heterogeneity bias.

Research Collaboration · Robust Curriculum Learning for LLM Post-Training

Developing adaptive adversarial curricula for LLM reasoning post-training by extending static GDRO-style group reweighting to a sequential curriculum-learning framework, where a curriculum controller uses on-policy self-distillation and single-trajectory, corruption-tolerant Q-learning to prioritize difficult task groups despite unreliable verifier, teacher, or state feedback.

Representative Publications

[1] S. Maity and A. Mitra, “Corruption-Tolerant Optimal Asynchronous Q-Learning”, in
[2] S. Maity and A. Mitra, “Adversarially-Robust TD Learning with Markovian Data”, in
[3] S. Maity and A. Mitra, “Robust Asynchronous Q-Learning under Reward and State Corruption via Batching”, in
[4] S. Maity and A. Mitra, “Robust Q-Learning under Corrupted Rewards”, in
[5] S. Maity and A. Mitra, “Robust Federated Q-Learning with Almost No Communication”, in
[6] S. Maity, F. Zhu, R. Heath, and A. Mitra, “Variance-Reduced Q-Learning over Static and Time-Varying Networks”, in
[7] S. Maity and A. Mitra, “Learning from Unreliable Trajectories: Adversarially-Robust Federated Q-Learning”,
[8] S. Maity, F. Zhu, R. Heath, and A. Mitra, “Decentralized Q-learning with Asynchronous Sampling and Partial Coverage”,
[9] S. Maity, L.Toso, J.Anderson, and A.Mitra, “Value Function Representation for Adversarially Robust Federated TD Learning”,

Projects

Federated MARL-GYM
We introduce a custom multi-agent reinforcement learning environment built with Gymnasium and Pygame, designed for evaluating federated RL (FRL) algorithms. The environment models a grid world where multiple agents navigate to accomplish spatially distributed tasks, like reaching delivery points.

Academic and Professional Service

  • Head Teaching Assistant for ECE 516: Systems and Control Engineering and ECE 308: Elements of Control Systems, Department of Electrical and Computer Engineering, NC State.
  • Served as a reviewer for 40+ papers in multiple flagship control/ ML venues, including the American Control Conference (ACC), the IEEE Conference on Decision and Control (CDC), Learning for Dynamics and Control (L4DC), Annual Conference on Neural Information Processing Systems(NeuRIPS), Journal of Machine Learning Research (JMLR), Transactions in Machine Learning Reserach (TMLR), IEEE Transactions in Automatic Control (TACON), Transactions in Signal and Information Processing over Networks (TSIPN), and Transactions in Signal Processing (TSP).

Awards

  • ACC 2026 Travel Award May 2026
  • L4DC Student Support Grant May 2025
  • NESCW 2025 Student Support Grant May 2025
  • IEEE CDC 2024 Student Support Award August 2024
  • NC State ECE Student Research Support Award August 2024
  • College of Engineering Graduate Merit Award 2023--24, 2024--25

Skills

Mathematics
Linear Algebra, Matrix Theory, Probability Theory, Random Process, Randomized Algorithms, Graph Theory, Analysis, Causal Inference, Bayesian Analysis, High-Dimensional Statistics, Stochastic Optimization.
Research
Reinforcement Learning, Statistical Learning Theory, Bandit Algorithms, Optimization, Information Theory, Control Theory, Stochastic Approximation, Robust Statistics, Distributed Networks.
Programming
Python, MATLAB, Simulink, C/C++.
Software and Libraries
PyTorch, TensorFlow, NumPy, Pandas, scikit-learn, Stable-Baselines3, CVXPY, CasADi, CleanRL, Gymnasium, MuJoCo.

Relevant Coursework

  • Learning Theory: Theoretical Foundations of Large-Scale Machine Learning, Reinforcement Learning, Machine Learning for Signal Processing, Bayesian Learning, Physics Modelling with Neural Networks, Deep Learning and Neural Networks.
  • Mathematics: Analysis I, Probability and Stochastic Processes, Stochastic Models and Applications, Convex Optimization for Data Science, Detection and Estimation Theory.
  • Control Theory: Dynamics of Linear Systems, Networked and Distributed Control, Safety-Critical Control for Robotic Systems, Non-Linear Control Theory, Formal Analysis for Control Theory.