Reinforcement learning -based multi-stage decision optimisation in career path planning
Author Names:
Mengdie Guo
Author Affiliation:
School of Marxism, Shanghai Normal University Tianhua College, Shanghai 201815, China
Author Email:
15000367686@163.com
Publication Date:
June 5, 2026
Page numbers:
DOI Number:
https://doi.org/10.1177/14727978251361399
Abstract:
In a career development environment that is very dynamic and complicated, standard career path planning approaches can’t really deal with the uncertainty of the environment in multi-stage decision-making because they rely on static models and expert expertise. This research suggests a reinforcement learning (RL)-based career path planning model that uses a Markov decision process (MDP) to simulate multi-stage decision optimisation and deep Q network (DQN) to perform adaptive iterative optimisation of strategies. A cumulative reward normalised score comparison experiment and a learning rate sensitivity analysis experiment show that the model works best at a medium learning rate (0.01) and that it reaches its highest cumulative reward score during the training period of 25,000 to 35,000 steps. In conclusion, the RL algorithm suggested in this research provides both theoretical and technical support for making an effective career planning system.
Keywords:
career path planning, RL, multi-stage decision-making, MDP, DQN
You need to register before accessing this content.