
Learning performance
Did not solve the environment. Best 100-episode average reward was 121.10 ± 5.61. (CartPole-v0 is considered "solved" when the agent obtains an average reward of at least 195.0 over 100 consecutive episodes.)
Did not solve the environment. Best 100-episode average reward was 121.10 ± 5.61. (CartPole-v0 is considered "solved" when the agent obtains an average reward of at least 195.0 over 100 consecutive episodes.)