The invention relates to the technical field of ocean platform pipeline design, in particular to an ocean platform pipeline laying method based on
reinforcement learning, and the method comprises the following steps: S1, dispersing a pipeline laying region into a three-dimensional grid matrix according to the physical size of an actual cabin of an ocean platform, and marking a pipeline starting point, a pipeline ending point and an impassable region; s2, performing Q-Learning
algorithm parameter initialization configuration, defining an action space adaptive to the linear movement characteristics of the ocean platform pipeline, constructing a Q value matrix adaptive to three-dimensional space coordinates and actions, and performing initialization; s3, entering a training round, and continuously updating the value evaluation matrix by dynamically adjusting a greedy criterion, a multi-dimensional reward mechanism and a
time sequence difference
learning rule; and S4, after the training is completed, starting from the starting point based on the converged Q value matrix, selecting an optimal action through a greedy to generate a final pipeline path, and improving the quality and search efficiency of pipeline laying on the ocean platform.