Lesson 2
You deleted the file and the image stayed the same size
A Docker image is a stack of read-only layers. A later layer can hide a file from an earlier one,
but it cannot modify that earlier layer — it is already packed and already has a digest. So an
rm in a later RUN only hides the file: the bytes stay in the image and
still travel over the wire on every docker pull.
Three images, one payload
The three Dockerfiles below all start from alpine:3.22 and all write the same 20 MB of
random data. They differ only in when the deletion happens, or whether it happens at all:
# keep — write it and leave it there
RUN dd if=/dev/urandom of=/big bs=1M count=20
# two — delete it in a separate RUN
RUN dd if=/dev/urandom of=/big bs=1M count=20
RUN rm -f /big
# one — write and delete inside a SINGLE RUN
RUN dd if=/dev/urandom of=/big bs=1M count=20 && rm -f /big
| Image | inspect .Size | Layers | Is /big in the container? |
|---|---|---|---|
alpine:3.22 (base) | 4,133,954 | 1 | — |
keep | 25,103,572 | 2 | yes, 20,971,520 bytes |
two | 25,104,035 | 3 | no |
one | 4,125,528 | 2 | no |
two measures 25,104,035 bytes; keep — which still has the whole 20 MB
file in it — measures 25,103,572. The difference is 463 bytes, and it goes the wrong way. The
rm recovers nothing; it only adds a layer recording that the file is gone.
Folding both operations into one RUN is a completely different outcome:
one measures 4,125,528 bytes, which is 6.09 times smaller than
two, and neither of them has /big at runtime.
docker history shows where the bytes went
docker history prints each layer's size, so the missing 21 MB is easy to find:
$ docker history dsize:two --format '{{.Size}}\t{{.CreatedBy}}'
4.1kB RUN /bin/sh -c rm -f /big
21MB RUN /bin/sh -c dd if=/dev/urandom of=/big bs=1M count=20
0B CMD ["/bin/sh"]
9.2MB ADD alpine-minirootfs-3.22.6-aarch64.tar.gz /
$ docker history dsize:one --format '{{.Size}}\t{{.CreatedBy}}'
4.1kB RUN /bin/sh -c dd if=/dev/urandom of=/big bs=1M count=20 && rm -f /big
0B CMD ["/bin/sh"]
9.2MB ADD alpine-minirootfs-3.22.6-aarch64.tar.gz /
In two the 21 MB layer is still intact and the rm layer weighs 4.1 kB. In
one, the very instruction that wrote 20 MB produces a 4.1 kB layer — because by the
time the layer was packed, the file had already been removed.
docker history gives 9.2 + 21 = 30.2 MB,
docker image inspect .Size says 25,104,035 bytes, and docker image ls
says 55.3 MB — all three are right by their own accounting, because the containerd image store
keeps both the compressed blob and the unpacked copy. Whenever you quote an image size, name the
command that produced it. This lesson uses
docker image inspect --format '{{.Size}}'.
Why a later layer cannot erase an earlier one
Each layer is a tarball of changes relative to the layer beneath it, with a SHA-256 digest computed over that content. That digest is what lets layers be shared between images and verified on pull. If an instruction in a later layer could modify an earlier one, the digest would change and every other image sharing that layer would break.
So a deletion is recorded differently: the union filesystem writes a whiteout — under
overlayfs, an empty character device with the name of the deleted file. Processes inside the
container see the file disappear; the bytes in the lower layer stay exactly where they were. That
whiteout is the 4.1 kB that docker history attributes to the rm.
The lab
Compose your RUN steps, set how many MB each one writes and deletes, and mark which
steps share a single RUN with the one above. The model differs from
docker image inspect by at most 12,552 bytes across the three images in the table.
Your RUN commands
The images that were actually built:
The image you get
Three ways to avoid it
-
Fold it into one
RUN. Download, unpack, build, clean up — all chained with&&in a single instruction. That is what produced the 4,125,528-byteoneimage above. The price is that the whole block is a single cache link: change one character in it and all of it runs again. -
Use multi-stage builds. The first stage has the compilers and the headers; the
final stage
COPY --froms only what it needs. None of the first stage's layers end up in the final image, so anything you do not copy across simply does not exist in it. -
Do not write what you are going to delete. The common case is package-manager
caches:
apk add --no-cache,apt-get … && rm -rf /var/lib/apt/lists/*in the sameRUN, or BuildKit's--mount=type=cacheso the cache lives outside the image entirely.
Check yourself
Your image is 1.2 GB. You add RUN rm -rf /root/.cache at the end of the Dockerfile and rebuild. How big is it now?
Still about 1.2 GB, plus a few kB for the layer recording the deletion. That is exactly the two case in the measured table: deleting in a later layer recovers nothing. You have to delete inside the same RUN that created the cache, or use a multi-stage build.
Why is two larger than keep rather than equal to it?
Because two has a third layer that exists only to record the deletion. docker history puts that layer at 4.1 kB; the difference in inspect .Size is 463 bytes. Either way, the deletion only makes the image bigger.
You want a small image and a good build cache. Are those goals in conflict?
They are, if folding everything into oneRUN is your only tool: the bigger the block, the coarser the cache. Multi-stage builds resolve both — the build stage keeps many small steps for a fine-grained cache, and the final stage only COPY --froms the result, so image size no longer depends on how many steps the build stage had.