Skip to content

Containers Architecture

Fundamental

Operating System (Linux Kernel)

Kernel Primitives

Kernel primitives are the fundamental & low-level operations provided by the operating system kernel which is utilized to contruct a sandboxed container environment.

Container Environment

A container environment constructred using the kernel primitives with OverlayFS.

An example instance could be as follows,


Container Instance

Diagram 1: Rough Architecture of container system

To create a container, a container image is used. The contents of this image file are duplicated into the sandboxed environment as the root filesystem using OverlayFS and chroot. There are many strategies for mounting the root filesystem in the container, but OverlayFS is quite the common one.

In Diagram 1 above, the sandboxed container environment can be seen as the last third box. The middle box holds the kernel primitives that play the key role in container technologies and they are as follows:

Container Primitives Definitions
cgroups cgroup allows putting limits on a process and its children. Commonly used for limiting CPU and RAM usage. cgroups are technically optional for containers. However, you may need it in production.
pid_ns The PID namespace (pid_ns) allows a process and its children to run in a new process tree that maps back to the host process tree.
pic_ns The Inter-Process Communication Namespace (ipc_ns) limits the processes ability to share memory.
net_ns The Network Namespace (net_ns) allows a new network stack to exist in the sandbox. This means our sandboxed environment can have its own network interfaces, routing tables, DNS lookup servers, IP addresses, and etc…​ you name!
uts_ns Ironic as it is, The Unix Time Sharing Namespace (uts_ns) exists purely to isolate the system identity strings. This allows a container to assign its own hostname without conflicting with the host.
mnt_ns The Mount Namespace (mnt_ns) is the part of the kernel that stores the mount table. When the sandboxed environment runs in a new Mount Namespace, it can mount filesystems not present on the host. This is very important as you’ll see.
user_ns The User Namespace (user_ns) the sandboxed environments to have its own set of user and group IDs that will map to unique user and group IDs back on the host system.
seccomp seccomp is a utility acts as a filter for kernel calls. This allows us to drop Kernel capabilities in the sandboxed environment. Utilizing seccomp is also not strictly vital to containers.

Unshare

Container technology relies on namespaces to create the sandbox. New Linux Namespaces are typically spawned by using either the clone or unshare system calls. These exist as C functions that have wrapper in many languages. For using in a shell, unshare is more straight forward.

unshare --help

The mnt_ns is what responsible for storing the mount table in Linux kernel. When a sandboxed environment like a container runs in a new Mount Namespace, it can mount filesystems not present on the host.

sudo unshare -m /bin/bash
[sudo] password for localhost:
root@localhost:/var/home/localhost/# whoami
root

depending on where you are using, you may need sudo.

  • Unshare wraps the unshare kernel sys call.
  • -m requests a mount new Namespace.
  • /bin/bash tells what program to run after the mount. We ran bash thus, it gave a bash terminal
mount -t tmpfs tmpfs /mnt
mount | grep mnt

It creates a tem fs in ram and covers up the existing up one.

Now lets add something to the new /mnt

date > /mnt/date
cat /mnt/date
Fri Sep 11 10:49:31 AM +0530 2026

The shell that spawned the namespace can be closed by using the exit command.

exit
exit

Now open a shell and read from /mnt/date

cat /mnt/date
cat: /mnt/date: No such file or directory

As it was created in a temporary mount namespace, it is gone after exiting the shell that spawned it.

chroot

Unshare carves out a new namespace for the container. Since a container needs to be as if a whole operating system on its own, a root for it is needed.

A fakeroot for the container that the container finds as real can be constructed by chroot.

A chroot is an operation that changes the apparent root directory for the current running process and their children.

A program that is run in such a modified environment cannot access files and commands outside the specified directory tree. This modified environment is called a chroot jail.

chroot shouldn’t be used as an isolating. it must be paired with other tools and practices for it.

A new root means it should have a filesystem just small enough and the filesystem can be created inside of a directory as follows.

mkdir -p chroot/{bin,lib64}
cd chroot
cp -a /bin/{bash,cat,mkdir,pwd,cd} bin

the chroot dir is going to be the root for itself and its recursives.

bin, lib, and lib64 inside chroot and copied the binaries of cat, echo, cd, and mkdir into bin but, Binary alone can’t run because of missing dependencies. Missing dependencies is found using ldd and must be added for chroot to work.

ldd /bin/{bash,cat,mkdir,pwd,cd}
/bin/bash:
	linux-vdso.so.1 (0x00007f585e7fc000)
	libtinfo.so.6 => /lib64/libtinfo.so.6 (0x00007f585e634000)
	libc.so.6 => /lib64/libc.so.6 (0x00007f585e43b000)
	/lib64/ld-linux-x86-64.so.2 (0x00007f585e7fe000)
/bin/cat:
	linux-vdso.so.1 (0x00007f8057629000)
	libc.so.6 => /lib64/libc.so.6 (0x00007f805740b000)
	/lib64/ld-linux-x86-64.so.2 (0x00007f805762b000)
/bin/mkdir:
	linux-vdso.so.1 (0x00007f8302c10000)
	libselinux.so.1 => /lib64/libselinux.so.1 (0x00007f8302bb3000)
	libc.so.6 => /lib64/libc.so.6 (0x00007f83029ba000)
	libpcre2-8.so.0 => /lib64/libpcre2-8.so.0 (0x00007f8302909000)
	libgcc_s.so.1 => /lib64/libgcc_s.so.1 (0x00007f83028dc000)
	/lib64/ld-linux-x86-64.so.2 (0x00007f8302c12000)
/bin/pwd:
	linux-vdso.so.1 (0x00007f3445946000)
	libc.so.6 => /lib64/libc.so.6 (0x00007f3445729000)
	/lib64/ld-linux-x86-64.so.2 (0x00007f3445948000)
/bin/cd:
	not a dynamic executable

A binary that has all necessary libraries within itself and does not rely on external shared libraries is considered not a dynamic executable. Some of the dep here might be a symlink and it will fail if just the link name added without its corresponding target

VDSO, shorthand for virtual dynamic shared object, injected by the kernel at execve() time. it’s mapped directly from kernel memory into the process address space. There’s no file on disk, so there’s nothing to copy. It works automatically in any chroot as long as the kernel supports it

for lib in            \
    libtinfo.so.6      \
    libc.so.6           \
    ld-linux-x86-64.so.2 \
    libselinux.so.1       \
    libpcre2-8.so.0        \
    libgcc_s.so.1; do
ls -la /lib64/$lib
done
lrwxrwxrwx. 1 root root 15 Jan  1  1970 /lib64/libtinfo.so.6 -> libtinfo.so.6.6
-rwxr-xr-x. 1 root root 2482096 Jan  1  1970 /lib64/libc.so.6
-rwxr-xr-x. 1 root root 1002568 Jan  1  1970 /lib64/ld-linux-x86-64.so.2
-rwxr-xr-x. 1 root root 209040 Jan  1  1970 /lib64/libselinux.so.1
lrwxrwxrwx. 1 root root 20 Jan  1  1970 /lib64/libpcre2-8.so.0 -> libpcre2-8.so.0.15.0
lrwxrwxrwx. 1 root root 25 Jan  1  1970 /lib64/libgcc_s.so.1 -> libgcc_s-16-20260819.so.1

Symlinks are denoted by ->. It seems three files are symlinked. lets add all files excluding vdso.

for lib in            \
    libtinfo.so.6*     \
    libc.so.6           \
    ld-linux-x86-64.so.2 \
    libselinux.so.1       \
    libpcre2-8.so.0*       \
    libgcc_s*; do
cp -a /lib64/$lib lib64/
done

All dep and target if the dep is a symlink are added, now a chroot can be initialised as follows,

sudo chroot . bin/bash

As you might now, bash or sh is also a program. It spawns a shell that can run commands.

mkdir hi
cd hi
pwd
/hi
Navigation

Type to search…

↑↓ navigate↵ selectEsc close