Class 6: Kernel interfaces part II and schedext¶
Date 31.03.2026
Data exchange between user and kernel address space¶
To read/write something from/to the memory space of user programs, use the following functions (actually macros):
put_user(kptr, ptr).write a byte/word/long word into user program memory space (from under the address
ptr); the macro definition works automagically - the size is determined by the type to whichkptrpoints.get_user(kptr, ptr).as above, but reading
Use the following functions to copy larger areas of memory:
unsigned long copy_from_user(void *to, const void __user *from, unsigned long n);
unsigned long copy_to_user(void __user *to, const void *from, unsigned long n);
The former allows copying data from the user address space to the kernel address space, the latter the opposite.
In general, they behave like memcpy,
but it is important to note that in case of an address page error of the user space,
they can cause the process to sleep until the page is downloaded from the swap.
Before copying, the correctness of the address in user space is checked.
If the beginning of the area is correct, but the rest is not, the longest possible fragment is copied.
The value returned by both functions is the number of NOT copied bytes - a non-zero value indicates a copy error.
The functions and their corresponding macro definitions are defined in the asm/uaccess.h file.
Note that functions for blocks of powers of two are optimized.
In case of an error when copying from/to user space, syscalls should return -EFAULT.
Scheduling in the kernel¶
Build in schedulers¶
If you would want to implement your own scheduler, you would have to implement sched_class.
There are few schedulers implemented that are present in the kernel.
sched/stop_task.cstop-task scheduling class
sched/deadline.cDeadline Scheduling Class (SCHED_DEADLINE)
sched/rtc.cReal-Time Scheduling Class
sched/fair.c.Completely Fair Scheduling (CFS) Class (SCHED_NORMAL/SCHED_BATCH).
sched/idle.cGeneric entry points for the idle threads and implementation of the idle task scheduling class.
In most cases, processes will be scheduled with the fair policy. The scheduler class used by the given
process can be changed through sched_setscheduler() syscall
Hands-on
Play around with the scheduler. Add debug messages to see how the task is scheduled. I would highly recommended printing this messages only to selected process. You can achieve it with this helper function:
Helper function:
helper.h
Do the following:
For functions that take the task struct as the argument, print the message in the function implementation.
For functions that return task struct, print the message in the place it's called (
sched/core.c) [recommended],
or just before the return value.
Sched_ext - BPF scheduler¶
Hands-on
To compile a kernel with support for schedext, you need to have few required options enabled in config file. You can try to manually enable them, according to kernel documentation or use this already prepared version:
In the November of 2022 a group of engineers from Google and Meta proposed a new scheduler class.
The motivation for the new class was following:
Ease of experimentation and exploration: Enabling rapid iteration of new scheduling policies.
Customization: Building application-specific schedulers which implement policies that are not applicable to general-purpose schedulers.
Rapid scheduler deployments: Non-disruptive swap outs of scheduling policies in production environments.
After almost 2 years of discussion, and 6 revisions, the sched_ext was merged and was included in the 6.11 kernel release.
Similarly to in kernel scheduler classes, BPF scheduler can be controlled
by implementing a set of operations, here defined in the struct sched_ext_ops.
BPF scheduler also introduces a new concept of local and global DSQ.
Documentation
When a CPU is ready to schedule, it first looks at its local DSQ. If empty, it then looks at the global DSQ. If there still isn't a task to run, ops.dispatch() is invoked which can use the following two functions to populate the local DSQ.
// This is a reduced version of the struct.
// See EBPF documentation, or kernel source code for full definition:
// https://docs.ebpf.io/linux/program-type/BPF_PROG_TYPE_STRUCT_OPS/sched_ext_ops/
struct sched_ext_ops {
void (*enable)(struct task_struct *p);
s32 (*init_task)(struct task_struct *p, struct scx_init_task_args *args);
s32 (*select_cpu)(struct task_struct *p, s32 prev_cpu, u64 wake_flags);
void (*runnable)(struct task_struct *p, u64 enq_flags);
void (*enqueue)(struct task_struct *p, u64 enq_flags);
void (*dispatch)(s32 cpu, struct task_struct *prev);
void (*running)(struct task_struct *p);
void (*tick)(struct task_struct *p);
void (*stopping)(struct task_struct *p, bool runnable);
void (*quiescent)(struct task_struct *p, u64 deq_flags);
void (*disable)(struct task_struct *p);
void (*exit_task)(struct task_struct *p, struct scx_exit_task_args *args);
};
The following pseudo code can help understand when each callback is invoked:
// taken from kernel documentation (scheduler/sched-ext.rst)
ops.init_task(); /* A new task is created */
ops.enable(); /* Enable BPF scheduling for the task */
while (task in SCHED_EXT) {
if (task can migrate)
ops.select_cpu(); /* Called on wakeup (optimization) */
ops.runnable(); /* Task becomes ready to run */
while (task is runnable) {
if (task is not in a DSQ && task->scx.slice == 0) {
ops.enqueue(); /* Task can be added to a DSQ */
/* Any usable CPU becomes available */
ops.dispatch(); /* Task is moved to a local DSQ */
}
ops.running(); /* Task starts running on its assigned CPU */
while (task->scx.slice > 0 && task is runnable)
ops.tick(); /* Called every 1/HZ seconds */
ops.stopping(); /* Task stops running (time slice expires or wait) */
/* Task's CPU becomes available */
ops.dispatch(); /* task->scx.slice can be refilled */
}
ops.quiescent(); /* Task releases its assigned CPU (wait) */
}
ops.disable(); /* Disable BPF scheduling for the task */
ops.exit_task(); /* Task is destroyed */
Readings¶
Some documentation about sched_ext, including a repository containing implementation of few production-ready schedulers