Class 2: LIBBPF

Date: 03.03.2026

Scenario

Introduction

On the previous labs, we learned how to use bpftrace tool to write very short BPF programs. However, bpftrace is not the only way to write BPF programs. Today, we will learn how to write and load BPF programs in C: first using linux/bpf.h header provided by kernel, then using the libbpf library.

Important

Before starting, make sure the following packages are installed:

sudo apt install clang libbpf-dev bpftool
  • clang — the BPF back-end compiler (--target=bpf); GCC does not support this target.

  • libbpf-dev — libbpf headers (bpf/bpf_helpers.h, etc.) and the shared library for -lbpf.

  • bpftool — utility for inspecting and managing BPF objects (programs, maps, BTF). On older kernels it may be packaged as linux-tools-$(uname -r).

linux/bpf.h

The bpf() system call (declared in <linux/bpf.h>) is the fundamental kernel interface for working with extended BPF (eBPF) — see the bpf(2) man page [1] for the complete specification. It accepts a command (cmd), an attribute union (union bpf_attr *attr), and a size argument.

For our classes, the two most important commands are BPF_PROG_LOAD, which verifies and loads a BPF program into the kernel, returning a file descriptor, and BPF_MAP_CREATE, which creates a key/value data structure (a map) that BPF programs and user-space can share. Before every program is executed, an in-kernel verifier statically proves that it terminates and is safe—so a buggy or malicious program is rejected at load time, not at run time.

Because the syscall operates on raw struct bpf_insn arrays, even a minimal program requires you to express BPF bytecode directly. The header <linux/bpf.h> provides macros such as BPF_MOV64_IMM, BPF_EXIT_INSN, and others that make this slightly more readable, but the overall workflow remains low-level: fill in a union bpf_attr, call bpf(), and check the returned file descriptor. Understanding this plumbing is important because every higher-level framework (libbpf, BCC, bpftrace) eventually calls this same syscall under the hood.

Hands-on

Work through Examples 1–3 below to get familiar with the raw bpf() syscall: defining BPF bytecode, observing verifier errors, and creating a map.

Example 1: loading and running a minimal BPF program

This example loads a BPF socket-filter program into the kernel using the raw bpf() syscall, attaches it to a socket, and triggers it by sending a packet. The BPF program uses the bpf_trace_printk helper to write "Hello from BPF!" to the kernel trace log.

/* hello_bpf.c — load a BPF program that prints via bpf_trace_printk */

#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <errno.h>
#include <unistd.h>
#include <sys/socket.h>
#include <sys/syscall.h>
#include <linux/bpf.h>

/* BPF_REG_10 is the frame-pointer register; define a readable alias */
#ifndef BPF_REG_FP
#define BPF_REG_FP BPF_REG_10
#endif

static inline int sys_bpf(enum bpf_cmd cmd, union bpf_attr *attr,
                           unsigned int size)
{
    return syscall(__NR_bpf, cmd, attr, size);
}

int main(void)
{
    /*
     * BPF program that calls bpf_trace_printk("Hello from BPF!\n").
     *
     * The string is pushed onto the stack in two 8-byte halves
     * (little-endian), then trace_printk is called with r1=stack ptr,
     * r2=length, and finally r0 is set to 0 and we exit.
     *
     * Equivalent pseudo-code:
     *   char fmt[] = "Hello from BPF!";
     *   bpf_trace_printk(fmt, sizeof(fmt));
     *   return 0;
     */
    struct bpf_insn prog[] = {
        /* store "Hello fr" (little-endian) at fp-16 */
        { .code = BPF_ST | BPF_DW | BPF_MEM,
          .dst_reg = BPF_REG_FP, .src_reg = 0,
          .off = -16, .imm = 0x6f727620 },       /* "o fr" (upper) — see note */
        /* We need two 32-bit stores for each 8-byte chunk because
         * BPF_ST | BPF_DW stores a 32-bit imm sign-extended to 64 bits.
         * Instead, we use 32-bit (BPF_W) stores for full control. */

        /* fp-16: "Hell" */
        { .code = BPF_ST | BPF_W | BPF_MEM,
          .dst_reg = BPF_REG_FP, .src_reg = 0,
          .off = -16, .imm = 0x6c6c6548 },       /* "lleH" LE = "Hell" */
        /* fp-12: "o fr" */
        { .code = BPF_ST | BPF_W | BPF_MEM,
          .dst_reg = BPF_REG_FP, .src_reg = 0,
          .off = -12, .imm = 0x7266206f },       /* "rf o" LE = "o fr" */
        /* fp-8: "om B" */
        { .code = BPF_ST | BPF_W | BPF_MEM,
          .dst_reg = BPF_REG_FP, .src_reg = 0,
          .off = -8, .imm = 0x42206d6f },        /* "B mo" LE = "om B" */
        /* fp-4: "PF!\0" */
        { .code = BPF_ST | BPF_W | BPF_MEM,
          .dst_reg = BPF_REG_FP, .src_reg = 0,
          .off = -4, .imm = 0x00214650 },        /* "\0!FP" LE = "PF!\0" */

        /* r1 = fp - 16  (pointer to format string) */
        { .code = BPF_ALU64 | BPF_MOV | BPF_X,
          .dst_reg = BPF_REG_1, .src_reg = BPF_REG_FP,
          .off = 0, .imm = 0 },
        { .code = BPF_ALU64 | BPF_ADD | BPF_K,
          .dst_reg = BPF_REG_1, .src_reg = 0,
          .off = 0, .imm = -16 },

        /* r2 = 17  (length including '\n' + '\0'... we use 16 for safety) */
        { .code = BPF_ALU64 | BPF_MOV | BPF_K,
          .dst_reg = BPF_REG_2, .src_reg = 0,
          .off = 0, .imm = 16 },

        /* call bpf_trace_printk (helper #6) */
        { .code = BPF_JMP | BPF_CALL,
          .dst_reg = 0, .src_reg = 0,
          .off = 0, .imm = 6 },

        /* r0 = 0 */
        { .code = BPF_ALU64 | BPF_MOV | BPF_K,
          .dst_reg = BPF_REG_0, .src_reg = 0,
          .off = 0, .imm = 0 },

        /* exit */
        { .code = BPF_JMP | BPF_EXIT,
          .dst_reg = 0, .src_reg = 0,
          .off = 0, .imm = 0 },
    };

    char log_buf[4096];
    memset(log_buf, 0, sizeof(log_buf));

    union bpf_attr attr;
    memset(&attr, 0, sizeof(attr));
    attr.prog_type = BPF_PROG_TYPE_SOCKET_FILTER;
    attr.insns     = (unsigned long)prog;
    attr.insn_cnt  = sizeof(prog) / sizeof(prog[0]);
    attr.license   = (unsigned long)"GPL";
    attr.log_level = 1;
    attr.log_size  = sizeof(log_buf);
    attr.log_buf   = (unsigned long)log_buf;

    int prog_fd = sys_bpf(BPF_PROG_LOAD, &attr, sizeof(attr));
    if (prog_fd < 0) {
        fprintf(stderr, "BPF_PROG_LOAD failed: %s\nVerifier log:\n%s\n",
                strerror(errno), log_buf);
        return EXIT_FAILURE;
    }
    printf("BPF program loaded, fd = %d\n", prog_fd);
    printf("Verifier log:\n%s\n", log_buf);

    /* Attach to a socket and send a packet to trigger the program */
    int sv[2];
    if (socketpair(AF_UNIX, SOCK_DGRAM, 0, sv) < 0) {
        perror("socketpair");
        return EXIT_FAILURE;
    }

    if (setsockopt(sv[0], SOL_SOCKET, SO_ATTACH_BPF,
                    &prog_fd, sizeof(prog_fd)) < 0) {
        perror("SO_ATTACH_BPF");
        return EXIT_FAILURE;
    }

    /* Sending data through the socket triggers the BPF filter */
    write(sv[1], "x", 1);
    printf("Packet sent — check the trace log:\n");
    printf("  sudo cat /sys/kernel/debug/tracing/trace_pipe\n");

    close(sv[0]);
    close(sv[1]);
    close(prog_fd);
    return EXIT_SUCCESS;
}

Compile and run:

gcc -o hello_bpf hello_bpf.c -Wall -Wextra
sudo ./hello_bpf

Then, in another terminal (or before running), read the trace log:

sudo cat /sys/kernel/debug/tracing/trace_pipe

You should see a line containing Hello from BPF! — this was printed by the BPF program running inside the kernel.

The key takeaways:

  • A BPF program is an array of struct bpf_insn instructions. Even a simple printk requires manually building a format string on the stack.

  • bpf_trace_printk is BPF helper #6 — called via BPF_JMP | BPF_CALL with imm = 6.

  • The program must be loaded with bpf(BPF_PROG_LOAD, …) and then attached to something (here, a socketpair via SO_ATTACH_BPF) before it can execute.

  • This is why higher-level tooling (libbpf, bpftrace) exists — writing bytecode by hand is difficult.

Example 2: three ways the verifier rejects a program

The BPF verifier runs before a program is executed. Below is a single source file that attempts to load three intentionally broken programs so you can see the verifier error messages.

/* verifier_errors.c — three programs the verifier will reject */
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <errno.h>
#include <unistd.h>
#include <sys/syscall.h>
#include <linux/bpf.h>

static inline int sys_bpf(enum bpf_cmd cmd, union bpf_attr *attr,
                           unsigned int size)
{
    return syscall(__NR_bpf, cmd, attr, size);
}

static void try_load(const char *label, struct bpf_insn *insns, int cnt)
{
    char log_buf[4096];
    memset(log_buf, 0, sizeof(log_buf));

    union bpf_attr attr;
    memset(&attr, 0, sizeof(attr));
    attr.prog_type = BPF_PROG_TYPE_SOCKET_FILTER;
    attr.insns     = (unsigned long)insns;
    attr.insn_cnt  = cnt;
    attr.license   = (unsigned long)"GPL";
    attr.log_level = 1;
    attr.log_size  = sizeof(log_buf);
    attr.log_buf   = (unsigned long)log_buf;

    int fd = sys_bpf(BPF_PROG_LOAD, &attr, sizeof(attr));
    printf("=== %s ===\n", label);
    if (fd < 0)
        printf("Rejected: %s\nVerifier log:\n%s\n\n", strerror(errno), log_buf);
    else {
        printf("Unexpectedly accepted!\n\n");
        close(fd);
    }
}

int main(void)
{
    /* --- Case 1: empty program (zero instructions) --- */
    {
        printf("--- Case 1: empty program (0 instructions) ---\n");
        struct bpf_insn prog[] = {
            /* nothing */
            { 0 }  /* placeholder so the array is non-empty in C */
        };
        try_load("Empty program", prog, 0);
    }

    /* --- Case 2: no exit instruction (falls off the end) --- */
    {
        printf("--- Case 2: no exit instruction ---\n");
        struct bpf_insn prog[] = {
            /* r0 = 0, but no BPF_EXIT follows */
            { .code = BPF_ALU64 | BPF_MOV | BPF_K,
              .dst_reg = BPF_REG_0, .src_reg = 0,
              .off = 0, .imm = 0 },
        };
        try_load("No exit", prog, 1);
    }

    /* --- Case 3: unreachable code after exit --- */
    {
        printf("--- Case 3: unreachable code after exit ---\n");
        struct bpf_insn prog[] = {
            { .code = BPF_ALU64 | BPF_MOV | BPF_K,
              .dst_reg = BPF_REG_0, .src_reg = 0,
              .off = 0, .imm = 0 },           /* r0 = 0 */
            { .code = BPF_JMP | BPF_EXIT,
              .dst_reg = 0, .src_reg = 0,
              .off = 0, .imm = 0 },           /* exit */
            /* dead code — verifier rejects this */
            { .code = BPF_ALU64 | BPF_MOV | BPF_K,
              .dst_reg = BPF_REG_0, .src_reg = 0,
              .off = 0, .imm = 1 },           /* r0 = 1 (unreachable) */
        };
        try_load("Unreachable code", prog, 3);
    }

    return 0;
}

Compile and run:

gcc -o verifier_errors verifier_errors.c -Wall -Wextra
sudo ./verifier_errors

Each case prints the verifier log explaining why the program was rejected:

  • Case 1 — empty program: the kernel rejects the bpf() syscall with E2BIG (Argument list too long) before the verifier even runs, because insn_cnt = 0 fails a basic sanity check.

  • Case 2 — no exit: the last instruction is not BPF_EXIT, so the program could "fall off" the end. The verifier reports last insn is not an exit or jmp.

  • Case 3 — unreachable code: the exit at instruction 1 makes instruction 2 dead code. The verifier reports unreachable insn.

These examples show that the verifier is strict: every instruction must be reachable, the program must end with an explicit exit, and it must contain at least one instruction.

Example 3: creating a BPF map and inspecting it with bpftool

This program creates a BPF hash map, writes value 12345 at key 42, and pauses so you can inspect the map from another terminal.

/* map_demo.c — create a map, write a value, pause for inspection */
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <errno.h>
#include <unistd.h>
#include <sys/syscall.h>
#include <linux/bpf.h>

static inline int sys_bpf(enum bpf_cmd cmd, union bpf_attr *attr,
                           unsigned int size)
{
    return syscall(__NR_bpf, cmd, attr, size);
}

int main(void)
{
    /* 1. Create a BPF_MAP_TYPE_HASH: up to 64 entries, key=u32, value=u32 */
    union bpf_attr map_attr;
    memset(&map_attr, 0, sizeof(map_attr));
    map_attr.map_type    = BPF_MAP_TYPE_HASH;
    map_attr.key_size    = sizeof(__u32);
    map_attr.value_size  = sizeof(__u32);
    map_attr.max_entries = 64;

    int map_fd = sys_bpf(BPF_MAP_CREATE, &map_attr, sizeof(map_attr));
    if (map_fd < 0) {
        perror("BPF_MAP_CREATE");
        return EXIT_FAILURE;
    }
    printf("Map created, fd = %d\n", map_fd);

    /* 2. Write key=42 → value=12345 */
    __u32 key   = 42;
    __u32 value = 12345;

    union bpf_attr upd;
    memset(&upd, 0, sizeof(upd));
    upd.map_fd = map_fd;
    upd.key    = (unsigned long)&key;
    upd.value  = (unsigned long)&value;
    upd.flags  = BPF_ANY;

    if (sys_bpf(BPF_MAP_UPDATE_ELEM, &upd, sizeof(upd)) < 0) {
        perror("BPF_MAP_UPDATE_ELEM");
        return EXIT_FAILURE;
    }
    printf("Set map[%u] = %u\n", key, value);

    /* 3. Pause so bpftool can inspect the map */
    printf("\nProcess paused — use bpftool in another terminal.\n");
    printf("Press Enter to exit and clean up...\n");
    getchar();

    close(map_fd);
    return EXIT_SUCCESS;
}

Compile and run:

gcc -o map_demo map_demo.c -Wall -Wextra
sudo ./map_demo

While the program is paused, open a second terminal as root and list all BPF maps:

sudo bpftool map show

You will see your array map, for example:

7: hash  flags 0x0
    key 4B  value 4B  max_entries 64  memlock 4096B

Now dump its contents in a human-readable format (replace 7 with the ID shown above):

sudo bpftool map dump id 7

The output will show only the entries you wrote, confirming your write:

key: 2a 00 00 00  value: 39 30 00 00

(0x2a = 42, 0x00003039 = 12345 in little-endian.

libbpf

As we saw in the examples above, working directly with the bpf() syscall is tedious: you have to hand-assemble bytecode, manage file descriptors for every map and program, and manually attach programs to kernel hooks. libbpf is a C library (maintained in the kernel tree under tools/lib/bpf/) that automates all of this. It is the recommended way to write production-quality BPF applications.

The typical libbpf workflow

  1. Write the BPF program in a .bpf.c file. You use normal C with special macros:

    • SEC("kprobe/…"), SEC("tracepoint/…"), SEC("xdp"), etc. — tells libbpf where to attach the program.

    • bpf_helpers.h — provides wrappers for in-kernel helper functions (bpf_printk, bpf_get_current_pid_tgid, bpf_map_lookup_elem, …).

    • bpf_tracing.h — gives access to function arguments via PT_REGS_PARM1 and similar macros.

    • Maps are declared as global structs with the SEC(".maps") attribute (see the stats map in the assignment for an example).

  2. Compile to a BPF ELF object with clang:

    clang --target=bpf -g -O2 -c example.bpf.c -o example.bpf.o
    

    The -g flag embeds BTF (BPF Type Format) information, which libbpf uses for CO-RE relocations.

  3. Generate a skeleton header (optional but very convenient):

    bpftool gen skeleton example.bpf.o name example > example.skel.h
    

    This header provides type-safe example__open(), example__load(), example__attach(), and example__destroy() functions, plus direct pointers to your maps and global variables.

  4. Write a user-space loader that #includes the skeleton, opens the object, loads it into the kernel, attaches the programs, and then enters an event loop or sleeps.

CO-RE: Compile Once — Run Everywhere

A key advantage of libbpf over older approaches (like BCC, which compiles BPF source at run time) is CO-RE. Instead of compiling against the exact kernel headers of the target machine, you compile once with a vmlinux.h header generated from BTF:

bpftool btf dump file /sys/kernel/btf/vmlinux format c > vmlinux.h

This single header contains every kernel type. At load time libbpf uses BTF information embedded in the ELF object to relocate field accesses so they match the running kernel's struct layout — even if fields were added, removed, or reordered between kernel versions.

Helper macros from bpf_core_read.h (e.g., BPF_CORE_READ, bpf_core_field_exists) let you write portable field accesses that the CO-RE relocator can adjust.

Hands-on

Work through Examples 4–6 below to practice writing and loading BPF programs with libbpf: attaching a kprobe, using maps from user-space, and comparing kprobes with tracepoints.

Example 4: probing do_nanosleep with a kprobe

Write a program that probes do_nanosleep and prints the process name by reading task_struct->comm via BPF_CORE_READ. This requires vmlinux.h because we are accessing a kernel-internal structure.

First, generate vmlinux.h (once, from the running kernel's BTF):

bpftool btf dump file /sys/kernel/btf/vmlinux format c > vmlinux.h
/* example.bpf.c */
#include "vmlinux.h"
#include <bpf/bpf_helpers.h>
#include <bpf/bpf_tracing.h>
#include <bpf/bpf_core_read.h>

SEC("kprobe/do_nanosleep")
int handle(void *ctx)
{
    struct task_struct *task = (void *)bpf_get_current_task();
    int pid = BPF_CORE_READ(task, tgid);
    char comm[16];
    BPF_CORE_READ_STR_INTO(&comm, task, comm);

    bpf_printk("PID %d (%s) is sleeping", pid, comm);
    return 0;
}

char LICENSE[] SEC("license") = "GPL";

Key differences from the earlier examples:

  • #include "vmlinux.h" replaces #include <linux/bpf.h> — it provides the definition of struct task_struct (and every other kernel type).

  • bpf_get_current_task() returns a pointer to the current task_struct.

  • BPF_CORE_READ(task, tgid) safely reads the tgid field using CO-RE relocations — if the field offset changes in a future kernel, libbpf adjusts it automatically at load time.

  • BPF_CORE_READ_STR_INTO copies the comm string from kernel memory.

And compile it with:

clang --target=bpf -g -Og -c  example.bpf.c -o example.bpf.o

The simplest way is to generate a skeleton file like:

bpftool gen skeleton  example.bpf.o name example > example.skel.h

And write a loader file like:

#include <unistd.h>
#include "example.skel.h"

int main()
{
      struct example *skel;
      int err = 0;

      skel = example__open();
      if (!skel)
              goto cleanup;

      err = example__load(skel);
      if (err)
              goto cleanup;

      err = example__attach(skel);
      if (err)
              goto cleanup;

      pause();

cleanup:
      example__destroy(skel);
      return err;
}

Which may be compiled and executed with:

gcc example.user.c -o example.user -lbpf
./example.user

In either way, open the trace printk log with:

bpftool prog tracelog

And execute a sleep program in another terminal.

Example 5: using BPF maps from libbpf

Maps are the primary mechanism for communication between a BPF program running in the kernel and user-space. With libbpf, you declare maps as global structs in the .bpf.c file; the skeleton gives you typed access from user-space.

The following example declares an array map that is written from the BPF program every time a process calls nanosleep. The BPF program stores value 12345 at key 42 in the map. The user-space loader then reads the map to confirm the write, showing that both sides can access the same map.

BPF side (mapex.bpf.c):

#include <linux/bpf.h>
#include <bpf/bpf_helpers.h>

/* Array map: key (u32 index) → value (u32) */
struct {
    __uint(type, BPF_MAP_TYPE_ARRAY);
    __uint(max_entries, 64);
    __type(key, __u32);
    __type(value, __u32);
} my_map SEC(".maps");

SEC("kprobe/do_nanosleep")
int set_map_entry(void *ctx)
{
    __u32 key   = 42;
    __u32 value = 12345;
    bpf_map_update_elem(&my_map, &key, &value, BPF_ANY);
    return 0;
}

char LICENSE[] SEC("license") = "GPL";

Key points:

  • The my_map map is declared with the SEC(".maps") attribute. libbpf parses this section from the ELF object and calls BPF_MAP_CREATE automatically at load time.

  • Inside the BPF program, bpf_map_update_elem is a kernel helper function — it operates on the map by pointer (&my_map), not by file descriptor.

  • The map write happens in kernel context every time do_nanosleep is called (i.e., every time any process sleeps).

Build:

clang --target=bpf -g -O2 -c mapex.bpf.c -o mapex.bpf.o
bpftool gen skeleton mapex.bpf.o name mapex > mapex.skel.h

User-space loader (mapex.user.c):

#include <stdio.h>
#include <stdlib.h>
#include <unistd.h>
#include <bpf/libbpf.h>
#include <bpf/bpf.h>
#include "mapex.skel.h"

int main(void)
{
    struct mapex *skel = mapex__open_and_load();
    if (!skel) {
        fprintf(stderr, "Failed to open/load BPF object\n");
        return 1;
    }
    if (mapex__attach(skel)) {
        fprintf(stderr, "Failed to attach BPF program\n");
        goto cleanup;
    }

    printf("BPF program attached — trigger it by running 'sleep 1' "
           "in another terminal.\n");
    printf("Press Enter after triggering...\n");
    getchar();

    /* Read back what the BPF program wrote */
    int map_fd = bpf_map__fd(skel->maps.my_map);
    __u32 key = 42;
    __u32 got = 0;
    if (bpf_map_lookup_elem(map_fd, &key, &got)) {
        perror("bpf_map_lookup_elem");
        goto cleanup;
    }
    printf("map[%u] = %u  (written by the BPF program)\n", key, got);

cleanup:
    mapex__destroy(skel);
    return 0;
}

Compile and run:

gcc mapex.user.c -o mapex -lbpf -lelf -lz
sudo ./mapex

Then, in another terminal, run sleep 1 to trigger the kprobe. After pressing Enter, the loader reads key 42 from the map and prints 12345 — the value that was written by the BPF program running inside the kernel.

You can also inspect the map while the program is running:

sudo bpftool map show
sudo bpftool map dump id <ID>
Example 6: kprobes vs. tracepoints

BPF programs can hook into the kernel at two main attachment points: kprobes and tracepoints. Understanding the difference is essential for choosing the right tool.

Kprobes (SEC("kprobe/…"))

  • Attach to any kernel function by name (e.g., do_nanosleep, __x64_sys_write).

  • The BPF program fires every time the function is entered (kprobe) or is about to return (kretprobe).

  • Pro: you can hook virtually any function in the kernel.

  • Con: function names and signatures are internal to the kernel and may change between versions. Your program might break on a different kernel.

  • The context pointer (void *ctx) is a struct pt_regs * — you read function arguments with PT_REGS_PARM1(ctx), PT_REGS_PARM2(ctx), etc. from <bpf/bpf_tracing.h>.

Tracepoints (SEC("tracepoint/…"), SEC("tp_btf/…"))

  • Attach to stable, explicitly defined instrumentation points in the kernel (e.g., tracepoint/syscalls/sys_enter_write, tracepoint/sched/sched_process_exit).

  • The kernel guarantees the tracepoint format across versions (within reason), so programs are more portable.

  • Pro: stable ABI, well-documented arguments.

  • Con: you can only hook where a tracepoint exists — not arbitrary functions.

  • The context is a struct matching the tracepoint's declared fields, which you can find under /sys/kernel/debug/tracing/events/.

You can list all available tracepoints with:

sudo ls /sys/kernel/debug/tracing/events/

And see the fields of a specific tracepoint:

sudo cat /sys/kernel/debug/tracing/events/syscalls/sys_enter_write/format

The following example attaches to the tracepoint syscalls/sys_enter_write instead of kprobing the syscall function. Compare this with the kprobe-based wrcnt.bpf.c above.

BPF side (tp_write.bpf.c):

#include <linux/bpf.h>
#include <bpf/bpf_helpers.h>

/* Tracepoint context for syscalls/sys_enter_write.
 * Fields come from the tracepoint format file.
 * The first two fields (common_type, etc.) are always present;
 * we skip them by declaring the struct starting at the args. */
struct sys_enter_write_args {
    unsigned long long unused;       /* padding / common fields */
    long               syscall_nr;
    unsigned long      fd;
    const char        *buf;
    unsigned long      count;
};

struct {
    __uint(type, BPF_MAP_TYPE_HASH);
    __uint(max_entries, 10240);
    __type(key, __u32);
    __type(value, __u64);
} write_bytes SEC(".maps");

SEC("tracepoint/syscalls/sys_enter_write")
int tp_count_write(struct sys_enter_write_args *ctx)
{
    __u32 pid = bpf_get_current_pid_tgid() >> 32;
    __u64 bytes = ctx->count;  /* direct field access — no PT_REGS needed */

    __u64 *total = bpf_map_lookup_elem(&write_bytes, &pid);
    if (total) {
        __sync_fetch_and_add(total, bytes);
    } else {
        bpf_map_update_elem(&write_bytes, &pid, &bytes, BPF_ANY);
    }
    return 0;
}

char LICENSE[] SEC("license") = "GPL";

Notice the differences from the kprobe version:

Kprobe

Tracepoint

SEC

SEC("kprobe/__x64_sys_write")

SEC("tracepoint/syscalls/sys_enter_write")

Context

void *ctx (struct pt_regs *)

Typed struct matching tracepoint fields

Argument access

PT_REGS_PARM1(ctx)

ctx->count (direct)

Portability

Kernel-version-specific

Stable ABI across versions

Scope

Any kernel function

Only where tracepoints are defined

Build:

clang --target=bpf -g -O2 -c tp_write.bpf.c -o tp_write.bpf.o
bpftool gen skeleton tp_write.bpf.o name tp_write > tp_write.skel.h

User-space loader (tp_write.user.c):

#include <stdio.h>
#include <stdlib.h>
#include <unistd.h>
#include <bpf/libbpf.h>
#include <bpf/bpf.h>
#include "tp_write.skel.h"

int main(void)
{
    struct tp_write *skel = tp_write__open_and_load();
    if (!skel) {
        fprintf(stderr, "Failed to open/load BPF object\n");
        return 1;
    }
    if (tp_write__attach(skel)) {
        fprintf(stderr, "Failed to attach BPF program\n");
        goto cleanup;
    }

    int map_fd = bpf_map__fd(skel->maps.write_bytes);

    printf("Tracing write() calls — Ctrl-C to stop.\n");
    while (1) {
        sleep(1);

        __u32 key = 0, next_key;
        printf("\n--- write bytes ---\n");
        while (bpf_map_get_next_key(map_fd, &key, &next_key) == 0) {
            __u64 value = 0;
            bpf_map_lookup_elem(map_fd, &next_key, &value);
            printf("  PID %u: %llu bytes written\n", next_key, value);
            key = next_key;
        }
    }

cleanup:
    tp_write__destroy(skel);
    return 0;
}

Compile and run:

gcc tp_write.user.c -o tp_write -lbpf -lelf -lz
sudo ./tp_write

Tip

Rule of thumb: prefer tracepoints when one exists for the event you care about (especially for syscalls and scheduler events). Fall back to kprobes when you need to instrument an internal function that has no tracepoint.

References