Emergent Locomotion Patterns of a Snake Robot through Reinforcement Learning

Yuya Shimizu1, Yongdong Wang1,2, So Shimooka1, Tetsushi Kamegawa1
1Graduate School of Environmental, Life, Natural Science and Technology, Okayama University, 2Department of Precision Engineering, The University of Tokyo
*Corresponding author: shimizu0y0mif@s.okayama-u.ac.jp
This work was supported by OU-SPRING, JSPS KAKENHI Grant Number 26K07423, and JSPS KAKENHI Grant Number 26K21350..
arXiv icon Paper Coming Soon Hugging Face icon Model Video icon Video

Abstract

Traditional gait design for snake robots has long relied on geometric models inspired by biological snakes. However, due to the inherent morphological discrepancies between biological snakes and robotic systems, these conventional models are not always optimal for snake robots. In this study, we conducted locomotion task training for a snake robot with ten redundant degrees of freedom (DoF) across random step fields. Experimental results demonstrated that the acquired policy autonomously emerged dynamic gaits capable of traversing unstructured terrains where conventional locomotion based on simple sine waves typically fails. The robot effectively utilized environmental protrusions as pivot points for propulsion, three-dimensionally transforming its body to overcome obstacles. Although these gaits deviate from mathematically predefined periodic motions, they exhibit rational characteristics approximating biological sidewinding as a direct consequence of physical constraints. Comparative experiments with geometric models validated that the learned policy possesses enhanced robustness and environmental adaptability in rugged terrains. Our findings suggest that locomotion strategies leveraging a robot’s unique embodiment can further enhance performance, extending the capabilities of snake robots beyond the conventional framework of pure biomimetics.

Overview of the proposed framework

The policy takes a 56-dimensional state as input. Per-link contact forces (highlighted) let the robot sense terrain through each link.

Key Results

Emergent Sidewinding

We provided no gait template. A sidewinding-like motion emerged from training, driven purely by the physical constraints of the robot and terrain.

Obstacle-Aided Locomotion

The robot learned to push off terrain protrusions as pivot points, twisting its body in 3D to get past obstacles.

Comparison with Geometric Model

The conventional sine-wave model tends to pile up on rough terrain and stall. The trained policy does not show this behavior.

Snake robot structure

Snake robot with 11 links and 10 alternating yaw/pitch joints (total length: 1.6 m).

Results

Navigation trajectories comparison

Navigation trajectories on the Random Grid terrain. The learned policy (top) reaches the goal in 103 s; the geometric sidewinding model (bottom) takes 165 s and covers less ground.

Simulation Videos

Learned policy on rough terrain (Random Grid). The robot uses terrain protrusions as pivot points to move forward.

Conventional geometric sidewinding model (A = 1.2 rad, ω = 0.50 Hz) on rough terrain. Baseline comparison.

Learned policy on flat ground. A sidewinding-like gait emerged without any template.

Geometric sidewinding model on flat ground. Maintains a constant periodic waveform across all joints.