From classical to quantum reinforcement learning and its applications in quantum control: A beginner’s tutorial
DOI (Low Temperature Physics):
https://doi.org/10.1063/10.0044783Ключові слова:
reinforcement learning, artificial intelligence, quantum control, Monte Carlo methods, probability theoryАнотація
Посібник розроблено, щоб зробити навчання з підкріпленням (RL) доступнішим для студентів та дослідників, пропонуючи чіткі пояснення на прикладах. Він зосереджений на подоланні розриву між теорією RL та практичним кодуванням, вирішуючи поширені проблеми, з якими стикаються під час переходу від концептуального розуміння до впровадження. За допомогою практичних прикладів та доступних пояснень посібник має на меті надати читачеві базові навички, необхідні для впевненого застосування методів RL у реальних сценаріях.
Доступність коду: https://github.com/asen009/Reinforcement_Tutorial_Undergraduates
Посилання
J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel, T. Lillicrap, and D. Silver, “Mastering atari, Go, chess and shogi by planning with a learned model,” Nature 588, 604 (2020). https://doi.org/10.1038/s41586-020-03051-4
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, T. Lillicrap, K. Simonyan, and D. Hassabis, “A general reinforcement learning algorithm that masters chess, shogi, and go through self-play,” Science 362, 1140 (2018). https://doi.org/10.1126/science.aar6404
M. G. Bellemare, S. Candido, P. Samuel Castro, J. Gong, M. C. Machado, S. Moitra, S. S. Ponda, and Z. Wang, “Autonomous navigation of stratospheric balloons using reinforcement learning,” Nature 588, 77 (2020). https://doi.org/10.1038/s41586-020-2939-8
A. Irshayyid, J. Chen, and G. Xiong, “A review on reinforcement learning-based highway autonomous vehicle control,” Green Energy Intell. Transp. 3, 100156 (2024). https://doi.org/10.1016/j.geits.2024.100156
Md. Al-M. Khan, Md. R. J. Khan, A. Tooshil, N. Sikder, M. A. P. Mahmud, A. Z. Kouzani, and A.-Al Nahid, “A systematic review on reinforcement learning-based robotics within the last decade,” IEEE Access 8, 176598 (2020). https://doi.org/10.1109/ACCESS.2020.3027152
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe, “Training language models to follow instructions with human feedback,” NeurIPS 35, 27730 (2022). https://doi.org/10.52202/068431-2011
A. Sen and S. Panda, Reinforcement tutorial for undergraduates—booklets, https://github.com/asen009/Reinforcement_Tutorial_Undergraduates/tree/main/Booklets, GitHub repository (2025).
B. Jaeger, and A. Geiger, “An invitation to deep reinforcement learning,” Found. Trends Optim. 7, 1 (2024). https://doi.org/10.1561/2400000049
A. Gosavi, “Reinforcement learning: A tutorial survey and recent advances,” INFORMS J. Comput. 21, 178 (2009). https://doi.org/10.1287/ijoc.1080.0305
V. Heidrich-Meisner, M. Lauer, C. Igel, and M. Riedmiller, Reinforcement learning in a nutshell, in Proceedings of the 15th European Symposium on Artificial Neural Networks (ESANN 2007), Bruges, Belgium (2007), p. 277.
M. E. Harmon and S. S. Harmon, Reinforcement Learning: A Tutorial, Technical Report, (Wright Laboratory, WL/AACF, Wright-Patterson Air Force Base, OH (1996). Available at https://sites.socsci.uci.edu/∼lpearl/courses/readings/HarmonHarmon1996.pdf.
M. Naeem, G. De Pietro, and A. Coronato, “Application of reinforcement learning and deep learning in multiple-input and multiple-output (MIMO) systems,” Sensors 22, 309 (2022). https://doi.org/10.3390/s22010309
Please consult the official documentation https://docs.python.org/3/tutorial/datastructures.html#dictionaries for details about Python dictionaries.
Note there is a separate data type in Python for ordered dictionaries https://docs.python.org/3/library/collections.html#ordereddict-objects.
R. S. Sutton, and A. G. Barto, Reinforcement Learning, Adaptive Computation and Machine Learning Series (Bradford Books, Cambridge, MA, 2018).
https://docs.python.org/3/tutorial/datastructures.html #more-on-lists.
https://docs.python.org/3/tutorial/datastructures.html #tuples-and-sequences.
M. Ghasemi, A. H. Moosavi, I. Sorkhoh, A. Agrawal, F. Alzhouri, and D. Ebrahimi, An introduction to reinforcement learning: Fundamental concepts and practical applications, arXiv.2408.07712 (2024).
R. S. Sutton, “Learning to predict by the methods of temporal differences,” Mach. Learn. 3, 9 (1988). https://doi.org/10.1023/A:1022633531479
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Mach. Learn. 8, 229 (1992). https://doi.org/10.1023/A:1022672621406
E. Greensmith, P. L. Bartlett, and J. Baxter, “Variance reduction techniques for gradient estimates in reinforcement learning,” J. Mach. Learn. Res. 5, 1471 (2004).
W. van Heeswijk, Three Fundamental Flaws in Common Reinforcement Learning Algorithms (And How To Fix Them), Towards Data Science (2023). Available at https://towardsdatascience.com/three-fundamental-flaws-in-common-reinforcement-learning-algorithms-and-how-to-fix-them-951160b7a207
J. Peters and S. Schaal, “Natural actor-critic, proceedings of the 25th international conference on machine learning,” Neurocomputing. 71, 1180 (2008). https://doi.org/10.1016/j.neucom.2007.11.026
R. Wang, S. S. Du, L. F. Yang, and S. M. Kakade, “Is long horizon reinforcement learning more difficult than short horizon reinforcement learning?,” in Advances in Neural Information Processing Systems (NeurIPS 2020), edited by H. Larochelle et al (Curran Associates, Inc., 2020), Vol. 33. Available at https://proceedings.neurips.cc/paper/2020/file/6734fa703f6633ab896eecbdfad8953a-Paper.pdf.
J. Zhang, J. Kim, B. O’Donoghue, and S. Boyd, “Sample efficient reinforcement learning with REINFORCE,” Proceedings of the AAAI Conference on Artificial Intelligence 35, 10887 (2021). https://doi.org/10.1609/aaai.v35i12.17300
A. G. Barto, R. S. Sutton, and C. W. Anderson, “Neuronlike adaptive elements that can solve difficult learning control problems,” IEEE Transactions on Systems, Man, and Cybernetics SMC-13, 834 (1983). https://doi.org/10.1109/TSMC.1983.6313077
D. J. Tannor and S. A. Rice, “Control of selectivity of chemical reaction via control of wave packet evolution,” J. Chem. Phys. 83, 5013 (1985).
P. Brumer, and M. Shapiro, “Control of unimolecular reactions using coherent light,” Chem. Phys. Lett. 126, 541 (1986). https://doi.org/10.1016/S0009-2614(86)80171-3
R. S. Judson and H. Rabitz, “Teaching lasers to control molecules,” Phys. Rev. Lett. 68, 1500 (1992). https://doi.org/10.1103/PhysRevLett.68.1500
W. S. Warren, H. Rabitz, and M. Dahleh, “Coherent control of quantum dynamics,” Science 259, 1581 (1993). https://doi.org/10.1126/science.259.5101.1581
A. Assion, T. Baumert, M. Bergt, T. Brixner, B. Kiefer, V. Seyfried, M. Strehle, and G. Gerber, “Control of chemical reactions by feedback-optimized phase-shaped femtosecond laser pulses,” Science 282, 919 (1998). https://doi.org/10.1126/science.282.5390.919
J. L. Herek, W. Wohlleben, R. J. Cogdell, D. Zeidler, and M. Motzkus, “Quantum control of energy flow in light-harvesting complexes,” Nature 417, 533 (2002). https://doi.org/10.1038/417533a
A. P. Peirce, M. A. Dahleh, and H. Rabitz, “Optimal control of quantum-mechanical systems: Existence, numerical approximation, and applications,” Phys. Rev. A 37, 4950 (1988). https://doi.org/10.1103/PhysRevA.37.4950
N. Khaneja, T. Reiss, C. Kehlet, T. Schulte-Herbrüggen, and S. J. Glaser, “Optimal control of coupled spin dynamics: Design of NMR pulse sequences by gradient ascent algorithms,” J. Magn. Reson. 172, 296 (2005). https://doi.org/10.1016/j.jmr.2004.11.004
V. F. Krotov, Global Methods in Optimal Control Theory (Marcel Dekker, New York, 1996).
T. Caneva, T. Calarco, and S. Montangero, “Chopped random-basis quantum optimization,” Phys. Rev. A 84, 022326 (2011). https://doi.org/10.1103/PhysRevA.84.022326
D. J. Egger, and F. K. Wilhelm, “Adaptive hybrid optimal quantum control for imprecisely characterized systems,” Phys. Rev. Lett. 112, 240503 (2014). https://doi.org/10.1103/PhysRevLett.112.240503
S. Machnes, U. Sander, S. J. Glaser, P. de Fouquieres, A. Gruslys, S. Schirmer, and T. Schulte-Herbrüggen, “Comparing, optimizing, and benchmarking quantum-control algorithms in a unifying programming framework,” Phys. Rev. A 84, 022305 (2011). https://doi.org/10.1103/PhysRevA.84.022305
D. I. Bondar, L. Balada Gaggioli, G. Korpas, J. Mareček, J. Vala, and K. Jacobs, “Globally optimal control of quantum dynamics,” Phys. Rev. Res. 7, 043202 (2025). https://doi.org/10.1103/g4fb-xm13
L. B. Gaggioli, D. I. Bondar, J. Vala, R. Ovsiannikov, and J. Mareček, Unitary gate synthesis via polynomial optimization, arXiv:2508.01356 (2025).
S. J. Glaser, U. Boscain, T. Calarco, C. P. Koch, W. Köckenberger, R. Kosloff, I. Kuprov, B. Luy, S. Schirmer, T. Schulte-Herbrüggen, D. Sugny, and F. K. Wilhelm, “Training schrödinger’s cat: Quantum optimal control. strategic report on current status, visions and goals for research in Europe,” Eur. Phys. J. D 69, 279 (2015). https://doi.org/10.1140/epjd/e2015-60464-1
É Genois, N. J. Stevenson, N. Goss, I. Siddiqi, and A. Blais, “Quantum optimal control of superconducting qubits based on machine-learning characterization,” Phys. Rev. Appl. 24, 034073 (2025). https://doi.org/10.1103/d9yg-d3qr
D. Bukov, A. G. R. Day, P. Weinberg, A. Polkovnikov, and P. Mehta, “Reinforcement learning in different phases of quantum control,” Phys. Rev. X 8, 031086 (2018). https://doi.org/10.1103/PhysRevX.8.031086
Z. An, and D. L. Zhou, “Deep reinforcement learning for quantum gate control,” Europhys. Lett. 126, 60002 (2019). https://doi.org/10.1209/0295-5075/126/60002
L. Giannelli, S. Sgroi, J. Brown, G. S. Paraoanu, M. Paternostro, E. Paladino, and G. Falci, “A tutorial on optimal control and reinforcement learning methods for quantum technologies,” Phys. Lett. A 434, 128054 (2022). https://doi.org/10.1016/j.physleta.2022.128054
Y. Gao, X. Wang, N. Yu, and B. M. Wong, “Harnessing deep reinforcement learning to construct time-dependent optimal fields for quantum control dynamics,” Phys. Chem. Chem. Phys. 24, 24208 (2022).
D. Koutromanos, D. Stefanatos, and E. Paspalakis, “Control of qubit dynamics using reinforcement learning,” Information 15, 272 (2024). https://doi.org/10.3390/info15050272
J. O. Ernst, A. Chatterjee, T. Franzmeyer, and A. Kuhn, Reinforcement learning for quantum control under physical constraints, arXiv:2501.14372 (2025).
S. Li, Y. Fan, X. Li, X. Ruan, Q. Zhao, Z. Peng, R.-B. Wu, J. Zhang, and P. Song, “Robust quantum control using reinforcement learning from demonstration,” npj Quantum Inf. 11, 124 (2025). https://doi.org/10.1038/s41534-025-01065-2
R. Porotti, A. Essig, B. Huard, and F. Marquardt, “Deep reinforcement learning for quantum state preparation with weak nonlinear measurements,” Quantum 6, 747 (2022). https://doi.org/10.22331/q-2022-06-28-747
M. Yuezhen Niu, S. Boixo, V. N. Smelyanskiy, and H. Neven, “Universal quantum control through deep reinforcement learning,” npj quantum Inf. 5, 33 (2019). https://doi.org/10.1038/s41534-019-0141-3
T. Fösel, P. Tighineanu, T. Weiss, and F. Marquardt, “Reinforcement learning with neural networks for quantum feedback,” Phys. Rev. X 8, 031084 (2018). https://doi.org/10.1103/PhysRevX.8.031084
F. Shuang and H. Rabitz, “Control of quantum observables by tracking control fields,” J. Chem. Phys. 121, 9270 (2004). https://doi.org/10.1063/1.1799591
D. Dong and I. R. Petersen, “Quantum control theory and applications: A survey,” IET Control Theory Appl. 4, 2651 (2010). https://doi.org/10.1049/iet-cta.2009.0508
A. Isidori, Nonlinear Control Systems, Springer, 1995).
R. Marino, and P. Tomei, Nonlinear Control Design: Geometric, Adaptive And Robust (Prentice Hall, London, New York, 1995).
A. G. Campos, D. I. Bondar, R. Cabrera, and H. A. Rabitz, “How to make distinct dynamical systems appear spectrally identical,” Phys. Rev. Lett. 118, 083201 (2017). https://doi.org/10.1103/PhysRevLett.118.083201
G. McCaul, C. Orthodoxou, K. Jacobs, G. H. Booth, and D. I. Bondar, “Driven imposters: Controlling expectations in many-body systems,” Phys. Rev. Lett. 124, 183201 (2020). https://doi.org/10.1103/PhysRevLett.124.183201
G. McCaul, A. F. King, and D. I. Bondar, “Optical indistinguishability via twinning fields,” Phys. Rev. Lett. 127, 113201 (2021). https://doi.org/10.1103/PhysRevLett.127.113201
A. B. Magann, T.-S. Ho, C. Arenz, and H. A. Rabitz, “Quantum tracking control of the orientation of symmetric-top molecules,” Phys. Rev. A 108, 033106 (2023). https://doi.org/10.1103/PhysRevA.108.033106
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour, Policy gradient methods for reinforcement learning with function approximation, in Advances in Neural Information Processing Systems, edited by S. A. Solla, T. K. Leen, and K.-R. Müller (MIT Press, Cambridge, MA, 2000), Vol. 12, p. 1057.
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization, in: Proceedings of the 32nd international conference on machine learning (ICML 2015),” Proc. Mach. Learn. Res. 37, 1889 (2015).
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, Proximal policy optimization algorithms, arXiv:1707.06347 (2017).
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, Continuous control with deep reinforcement learning, arXiv:1509.02971 (2015).
S. Fujimoto, H. van Hoof, and D. Meger, Addressing function approximation error in actor-critic methods, arXiv:1802.09477 (2018).
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, Soft actor-critic algorithms and applications, arXiv:1812.05905 (2018).
K. Doya, “Reinforcement learning: Computational theory and biological mechanisms,” HFSP Journal 1, 30 (2007). https://doi.org/10.2976/1.2732246/10.2976/1
S. Y.-C. Chen, An introduction to quantum reinforcement learning (QRL), arXiv.2409.05846 (2024).
T. H. Su, S. Shresthamali, and M. Kondo, “Quantum framework for reinforcement learning: Integrating the markov decision process, quantum arithmetic, and trajectory search,” Phys. Rev. A 111, 062421 (2025). https://doi.org/10.1103/5lfr-xb8m
S. hua Wang and H. Lin, Quantum reinforcement learning (2023).
M. Ismail, M. Shaban, and S. Y.-C. Chen, “Hands-on introduction to quantum machine learning,” The International FLAIRS Conference Proceedings 37 (2024). https://doi.org/10.32473/flairs.37.1.135478
M. Schuld and F. Petruccione, Machine Learning with Quantum Computers (Springer International Publishing, 2021). https://doi.org/10.1007/978-3-030-83098-4