From DQN to continuous action spaces: the deterministic policy gradient, TD3's 3 fixes, SAC's maximum-entropy formulation, the squashed Gaussian policy, and where these ideas stand in modern robot learning.

9/9/2026 tech

给定一个整数数组 A,长度为 n

  1. 数组相邻元素之间的差的绝对值为 1: abs(A[i] - A[i+1]) = 1
  2. 数组仅有一个波峰或波谷

问题:请快速反回这个波峰或波谷的位置

2/28/2026 tech

MIT 6.S184, Flow Matching and Diffusion Models.

2/24/2026 tech
2/23/2026 tech
1/20/2026 tech

CS 285 Deep Reinforcement Learning, Sergey Levine, Lecture 6, Actor-Critic Algorithms notes.

12/12/2025 tech

CS 285 Deep Reinforcement Learning, Sergey Levine, Lecture 5, Policy Gradients notes.

11/9/2025 tech

CS 285 Deep Reinforcement Learning, Sergey Levine, Lecture 4 notes.

11/9/2025 tech

CS 285 Deep Reinforcement Learning, Sergey Levine, Lecture 7, Value Function Methods notes.

4/2/2025 tech