↓ Skip to main content

System reproducibility (a primitive design)

·856 words·5 mins

Defining The Problem
#

The importance of data in our century is stated all the time but people have not yet realised it. Our data is in danger of being erased. If your work was deleted, the only way to get it back would be to do it all over again. The same applies for your software configuration and the software installed in the machine. You would have to install the software and write the configs again from scratch.

What can cause such disaster you ask? I can state some of the causes (not being 100% exhaustive). (disaster being an event that leads to a catastrophic failure and even data loss)

  • Theft: someone steals your computer just to sell the hardware. You may get another machine but you still loose your data.
  • Natural disasters: fires, floods, earthquakes etc. Your computer may be destroyed in the event of one.
  • Hardware faults: your disk/SSD breaks. Hard disks are mechanical devices which means they have moving parts and that leads to wear and tear given enough time (the hard disk is the most failure-prone part of your computer). SSDs on the other hand may not be mechanical but they still wear out by the nature of their construction which is the reason they are not used for long-term, reliable data storage.
  • Software faults: some idiot programmer or vibe coder pushed a bug (I won’t even go into the details of what kind of bug), they did not test the code and it lead to catastrophic failure. This is not as frequent in professionally developed software because QA systems and CI/CD pipelines detect such cases but it’s still possible. (Kevin Fang is a channel that has examples of outages/failures etc)
  • Your fault: its your fault! You deleted it! That case might be the most common. The user may delete stuff by accident. Culprits include: running a command (or an option/flag of a command) that is not understood, miss-typing an option, target file etc, accidental overwrites, removing an external drive or USB stick… I could go on.

In conclusion, the need to make your system “disaster-tolerant” is mandatory, it’s not a “nice to have” feature for hackers and techies.

We will achieve this “disaster-tolerant” system by leveraging redundancy and basic, “primitive” CLI tools.

Note

Redundancy means having extra copies or spare components beyond the minimum needed, so that if one fails, another can take over and the system keeps working.

(A nice byproduct of this is easier system migration, having your system in a new machine in under 10 minutes is a great convenience)

3+1 System Layers
#

We will divide the system into some layers, one sitting above the other.

Layer NumberLayer Name
3User Data
2Configurations
1User Software
0Operating System

Layer 0 (Operating system)
#

As expected a new machine requires an operating system, so we have to install one. Unfortunately to reproduce this layer you would need to go through the installation process.

Layer 1 (User Software / User Space)
#

In this layer lives every user space program that you use (browser, text editor, file manager, CLI tools, etc). To reproduce this layer you would need a post-install script containing all the names of the wanted packages. By running this script all of those packages will be installed on the system. A bash script will do but people use a lot of other tools like chezmoi, Ansible or even NixOS which is a little overkill for the scope of this article.

Layer 2 (Configurations)
#

In this layer belongs every configuration file of the aforementioned installed software. In Linux those configs are hidden in the home directory, they are also called dotfiles because the system treats files that start with a dot as hidden. I won’t go into much detail in regards to dotfiles management, you can find countless tutorials online. What I recommend and use is just a GitHub (GitLab, Bitbucket, Codeberg) repo and GNU stow. What stow does is to generate all the links from the repo directory to home and/or .config so the programs can read and use them. To reproduce this layer you would just need to clone the dotfiles repo and then run stow to generate the links.

Layer 3 (User Data)
#

In this layer belongs all the data you have. They may include documents, pictures, password databases, ssh keys, anything that you consider sensitive information. You need to copy, package-archive and store them in another machine, ideally in a remote location. This is called a backup and for your convenience it should be automated. People usually use one of the countless graphical backup tools but what I recommend is rsync+cron. The destination of the backup may be an external drive, a home server or cloud storage. To reproduce this layer you would just get the latest backup from the storage medium. (backups probably deserve their own article)

Acknowledgements
#

The scope of this article was to build a disaster-tolerant system using just primitive CLI tools (bash, git, stow, rsync, cron). I am aware that other components of the system (for example systemd services) are not reproduced.

Angelos T. Dimoglis
Author
Angelos T. Dimoglis
Software And DevOps Engineer