Abstract
Traditional gait design for snake robots has long relied on geometric models inspired by biological snakes. However, due to the inherent morphological discrepancies between biological snakes and robotic systems, these conventional models are not always optimal for snake robots. In this study, we conducted locomotion task training for a snake robot with ten redundant degrees of freedom (DoF) across random step fields. Experimental results demonstrated that the acquired policy autonomously emerged dynamic gaits capable of traversing unstructured terrains where conventional locomotion based on simple sine waves typically fails. The robot effectively utilized environmental protrusions as pivot points for propulsion, three-dimensionally transforming its body to overcome obstacles. Although these gaits deviate from mathematically predefined periodic motions, they exhibit rational characteristics approximating biological sidewinding as a direct consequence of physical constraints. Comparative experiments with geometric models validated that the learned policy possesses enhanced robustness and environmental adaptability in rugged terrains. Our findings suggest that locomotion strategies leveraging a robot’s unique embodiment can further enhance performance, extending the capabilities of snake robots beyond the conventional framework of pure biomimetics.
The policy takes a 56-dimensional state as input. Per-link contact forces (highlighted) let the robot sense terrain through each link.
Key Results
Emergent Sidewinding
We provided no gait template. A sidewinding-like motion emerged from training, driven purely by the physical constraints of the robot and terrain.
Obstacle-Aided Locomotion
The robot learned to push off terrain protrusions as pivot points, twisting its body in 3D to get past obstacles.
Comparison with Geometric Model
The conventional sine-wave model tends to pile up on rough terrain and stall. The trained policy does not show this behavior.
Snake robot with 11 links and 10 alternating yaw/pitch joints (total length: 1.6 m).
Results
Navigation trajectories on the Random Grid terrain. The learned policy (top) reaches the goal in 103 s; the geometric sidewinding model (bottom) takes 165 s and covers less ground.
Simulation Videos
Learned policy on rough terrain (Random Grid). The robot uses terrain protrusions as pivot points to move forward.
Conventional geometric sidewinding model (A = 1.2 rad, ω = 0.50 Hz) on rough terrain. Baseline comparison.
Learned policy on flat ground. A sidewinding-like gait emerged without any template.
Geometric sidewinding model on flat ground. Maintains a constant periodic waveform across all joints.