For AI agents: the complete documentation index is available at https://docs.ovhcloud.com/it/llms.txt, the full documentation bundle is available at https://docs.ovhcloud.com/it/llms-full.txt, and this page is available as Markdown at https://docs.ovhcloud.com/it/guides/hosted-private-cloud/opcp/how-to-reset-opcp.md.

How to reset OPCP

Vedi come Markdown

Find out how to reset OPCP, from uninstalling the control plane and resetting the storage to deploying a fresh installation

Objective

This procedure describes how to do a full redeployment of OPCP without reinstalling the host Operating system of the Controller server.

Danger

This is a destructive operation. You will lose all your data including OPCP configuration as well as all data stored on OPCP servers when performing a complete reset of OPCP.

Requirements

You need SSH access to the controller node and sudo permissions. To check, run:

ssh <your-node-ip>
sudo -i

You will need those permissions on all controller nodes.

You also need SSH access to the switches. To check, connect to them:

ssh admin@<switch-ip>

You will find your switch IPs in the OPCP hardware config. Run opcp-cli config edit --hardware and look under the netbox_devices for devices with the role tor or tor_ipmi. Get the value from primary_ipv4.

To get the passwords for the switches, use opcp-cli secrets passwords --edit and look for the keys network_devices_admin_arista_secret, network_devices_admin_cisco_secret and network_devices_enable_secret.

Instructions

The OPCP control plane runs on K3s, with Ceph as its storage backend.

Get Ceph information

The first step is to get the information about the Ceph cluster:

ceph device ls

This will output something similar to this:

DEVICE                                     HOST:DEV                  DAEMONS  WEAR  LIFE EXPECTANCY
SAMSUNG_MZQL21T9HCJR-00A07_S64GNL0Y414610  opcp-crxdemo-6-0:nvme3n1  osd.0      0%
SAMSUNG_MZQL21T9HCJR-00A07_S64GNL0Y414613  opcp-crxdemo-6-0:nvme1n1  osd.3      0%
SAMSUNG_MZQL21T9HCJR-00A07_S64GNL0Y414614  opcp-crxdemo-6-0:nvme2n1  osd.2      0%
SAMSUNG_MZQL21T9HCJR-00A07_S64GNL0Y414621  opcp-crxdemo-6-0:nvme4n1  osd.1      0%

This lists every device running as a Ceph OSD. Keep this list: you will need it later to wipe the disks.

Stop K3s cluster

After that, stop K3s on every controller node:

systemctl stop k3s.service

Wipe Ceph cluster

The next step is to erase all Ceph data:

rm -rf /var/lib/rook

Then, on every controller node, wipe all signatures (partition table, filesystem, LVM, RAID, etc.) from the disks listed earlier:

wipefs -a -f /dev/nvmeXnX # run this for each disk listed in ceph device ls

You can now securely erase all data (crypto erase is faster and preferred):

nvme format /dev/nvmeXnX --ses=1 # crypto erase
# or
nvme format /dev/nvmeXnX --ses=2 # user data erase (overwrite)

Reset K3s

Reset the iptables rules, then uninstall K3s, to avoid switch partitioning after the reset:

iptables -P INPUT ACCEPT
iptables -P OUTPUT ACCEPT
iptables -P FORWARD ACCEPT
iptables -F
iptables -X
iptables -t nat -F
iptables -t nat -X
iptables -t mangle -F
iptables -t mangle -X
/usr/local/bin/k3s-uninstall.sh

Reset network equipment

To reset the network equipment, connect to the switches via SSH. There are two types of switches: IPMI and TOR.

Warning

Using SSH to access the switches means you have to have access to the passwords for the switches. Public key authentication is not configured.

Look in the Requirements section on how to get the passwords for the network equipment.

Reset IPMI switches (Cisco)

SSH into the IPMI switch and run the following commands:

en
write erase
reload

Reset the TOR switches (Arista)

Reset the startup config on every TOR switch. Start with TOR-B: otherwise you will no longer be able to reach it.

en
bash
sudo -i
rm -rf /mnt/flash/.ragnarok
rm -rf /mnt/flash/EOS*.swi.sha512sum
rm -rf /mnt/flash/NogTaskMgr.py
rm -rf /mnt/flash/ca-bundle.crt
rm -rf /mnt/flash/network-ovh-pyclient-.tar.gz

rm -rf /mnt/flash/nog.*.cfg
rm -rf /mnt/flash/nog_diff_startup.py
rm -rf /mnt/flash/nog_task_exec.py
rm -rf /mnt/flash/ragnarok
rm -rf /mnt/flash/ragnctl
rm -rf /mnt/flash/rc.eos
rm -rf /mnt/flash/startup-config
exit
exit
reload force
Warning

Do not run rm -rf /mnt/flash/*! This will erase the firmware: the switch can no longer boot to the ZTP config, and you will then have to upload a firmware via a serial cable.

Redeploy controller

On the first controller node, run the following command:

opcp-cli deploy --force-reconcile

Go further

How to install a controller

Getting started with your OPCP

For training or technical assistance implementing our solutions, contact your sales representative or visit our Professional Services page to request a quote and have your project analysed by our experts.

Join our community of users.

Questa pagina ti è stata utile?