| This paper considers a satellite network with free space optical links, where satellites are able to form intra and inter satellite links. Briefly, intra-satellite links connect satellites in the same orbit plane, and inter-satellite links (ISLs) connect satellites on different orbital planes. In this respect, a key problem is to jointly determine when to establish inter-satellite links and the routing between a source-destination pair of satellites or ground users. Critically, unlike prior works, the resulting solution must ensure there is no congestion, which leads to packet loss. Henceforth, this paper proposes the first {\em safe} reinforcement learning (RL) approach that allows agents to learn a policy to optimize inter-ISL activations and routing of traffic whilst ensuring zero packet loss. In particular, it incorporates a shield mechanism that ensures agents do not take actions that would lead to packet loss. The results show that, as compared to conventional approaches, the proposed RL approach achieves a 33\% improvement in terms of average throughput, and reduces average delay and queue lengths by 45\%. |