{"id":53,"date":"2019-12-10T21:36:00","date_gmt":"2019-12-11T02:36:00","guid":{"rendered":"https:\/\/www.vociferousvoid.org\/?p=53"},"modified":"2023-07-08T21:57:37","modified_gmt":"2023-07-09T01:57:37","slug":"risc-v-bare-metal-programming-chapter-5-its-a-trap","status":"publish","type":"post","link":"https:\/\/www.vociferousvoid.org\/index.php\/2019\/12\/10\/risc-v-bare-metal-programming-chapter-5-its-a-trap\/","title":{"rendered":"RISC-V Bare Metal Programming &#8211; Chapter 5: It&#8217;s a Trap!"},"content":{"rendered":"\n<p>Up to this point, the RISC-V tutorial has focused on single applications running on a single hardware thread. The application&#8217;s environment was composed of the processor&#8217;s state and memory map. The processor state was controlled via assembly instructions, and the memory map was defined at build time via a linker script. However, modern systems are almost always multiprogrammed and will execute many applications concurrently (by interleaving their instructions in a single stream). Moreover, multiple hardware threads are common, allowing application instructions to execute simultaneously. This workflow requires a lot more care to ensure correctness and proper separation of memory. This idea was touched upon briefly in <a href=\"https:\/\/www.vociferousvoid.org\/index.php\/2019\/11\/30\/risc-v-bare-metal-programming-chapter-4-another-brick-in-the-wall\/\">chapter 4<\/a> while discussing the <strong>A<\/strong> extension which provides atomic memory operations that were used to define synchronization primitives. In addition to memory synchronization, the application must control the execution environment of all of the active hardware threads, this chapter will explore the mechanisms available for this purpose.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"orgbf395c9\">The ABI<\/h2>\n\n\n\n<p>The examples presented thus far can all be logically separated into three layers: The application execution environment (AEE), the application binary interface (ABI), and the application code. This organisation allows for a single application to execute in a single AEE.<\/p>\n\n\n\n<p>The AEE is defined by the processor state and the static memory map which is defined by the linker script. The ABI describes the conventions that programs must follow and manages dynamic memory; in this case the stack. The upshot of defining an ABI layer is that application code in the third layer can interact with an abstract view of the machine implementation. This simplifies the development of applications by hiding some of the hardware details. For example, the preamble and postamble code for function calls, which are part of this ABI, were used by the application code to ensure that dynamic memory was managed correctly. These code templates can be provided by developer tools, such as high-level language compilers.<\/p>\n\n\n\n<p>This becomes even more important as application environments become more complex. As complexity increases in the application environment, more can be done by the ABI to hide some of the tedious details of the layer beneath it. This is where the operating system (OS) comes into play. The OS is a supervisor process which is sandwiched between the ABI and the supervisor binary interface (SBI). In this configuration, the SBI hides more of the hardware details from the OS, which provides an even more abstract view of the system to the applications. The definition of an SBI also improves the portability of the OS layer.<\/p>\n\n\n\n<p>Another advantage of this layered approach is that it is easier to enforce separation between applications. One application should not interfere with other applications, or with the supervisor itself. The RISC-V ISA defines three different privilege modes to enforce this. From most to least privileged, the three levels are: machine, supervisor, and user.<\/p>\n\n\n\n<p>ISA implementations may provide 1-3 of the defined privilege modes. All hardware implementations must provide machine mode (or M-mode). A secure embedded system may provide machine and user modes. A virtualized multiprogramming system should provide all three modes.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"org4951020\">What&#8217;s my Pay Grade?<\/h2>\n\n\n\n<p>The capabilities of the ISA implementation can be queried via the <code>misa<\/code> control and status register. The function illustrated in the following listing will determine which integer ISA and the number of privilege modes that are supported by the current hardware thread:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code> 1:         .text\n 2:         .align 2\n 3:         .global __system_check\n 4: __system_check:\n 5:         # Input\n 6:         # None\n 7:         #\n 8:         # Returns\n 9:         # - a0: The number of supported privilege modes (1, 2 or 3).\n10:         # - a1: The register width used by the ISA (in bytes).\n<span class=\"coderef-off\" id=\"coderef-load-misa\">11:         csrr    t0, misa   # Load misa into t0\n12:         li      a1, 4      # Load minimum register width\n13:         li      a0, 1      # M-mode is always supported\n14:         # Probe for user-mode\n15:         lui     t1, 0x100  # Set the u-mode mask in t1\n16:         and     t1, t0, t1\n17:         beqz    t1, 1f     # Determine if the U-mode bit is set in misa\n18:         addi    a0, a0, 1  # U-mode is available, increment a0\n19: 1:      # Probe for supervisor-mode\n20:         lui     t1, 0x40   # Set the S-mode mask in t1\n21:         and     t1, t0, t1\n22:         beqz    t1, 2f     # Determine if the S-mode bit is set in misa\n23:         addi    a0, a0, 1  # S-mode is available, increment a0\n24: 2:      # Determine register width\n25:         bgez    t0, 3f     # Determine if a1 holds the register width\n26:         slli    a1, a1, 1  # Multiply register width by 2\n27:         slli    t0, t0, 1\n28:         j       2b\n29: 3:      \n30:         ret\n31: \n<\/span><\/code><\/pre>\n\n\n\n<p>This function loads the <code>misa<\/code> Control and Status Register (CSR) into the temporary register <code>t0<\/code> on line <a href=\"#coderef-load-misa\">11<\/a> using the <code>csrr<\/code> pseudo instruction. The minimum register width is 4-bytes, therefore the immediate 4 is loaded into register <code>a1<\/code> as an initial value (line <a href=\"#coderef-load-xlen\">12<\/a>). Moreover, machine-mode is required by hardware implementations, thus the initial value of <code>a1<\/code> is 1.<\/p>\n\n\n\n<p>The least-significant 26-bits of the <code>misa<\/code> CSR are flags which indicate supported extensions; one for each letter of the alphabet. The <strong>S<\/strong> and <strong>U<\/strong> extensions are for supervisor and user mode respectively. Therefore, user mode will be available if bit 20 is set to 1. A bit mask is created on line <a href=\"https:\/\/www.vociferousvoid.org\/main\/riscv_bare_metal_chapter5#coderef-umode-mask\">15<\/a>.<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p>Since the mask is an immediate value that is wider than what is allowed by RISC-V I-type instructions, the <code>lui<\/code> instruction can be used to load the 20-bit immediate value then shift it left by 12-bits with a single instruction. Thus the immediate value 0x100 becomes 0x100000 when loaded following the load operation.<\/p>\n<\/blockquote>\n\n\n\n<p>A bit-wise <strong>AND<\/strong> is performed using the bit-mask and the <code>misa<\/code> CSR value to determine if user-mode is available. If the result of the <code>and<\/code> instruction is zero on line <a href=\"#coderef-check-umode\">17<\/a>, user-mode is not supported. Otherwise the value of <code>a1<\/code> is incremented on line <a href=\"#coderef-set-umode\">18<\/a>.<\/p>\n\n\n\n<p>Similarly supervisor mode is supported if bit 18 is set to 1. A bit mask is created on line <a href=\"#coderef-smode-mask\">20<\/a>, then a bit-wise <strong>AND<\/strong> is performed with the value if the <code>misa<\/code> CSR on the next line. If the result is zero, then supervisor mode is not supported. Otherwise the number of privilege modes is incremented by 1.<\/p>\n\n\n\n<p>The final step of the <code>__system_check<\/code> function is to determine the width of the hart&#8217;s registers. The <code>misa<\/code> CSR&#8217;s most-significant 2-bits encodes the width of the registers used by the ISA: 1) 32-bits, 2) 64-bits, 3) 128-bits. Determining this is complicated by the fact that the CSR&#8217;s width is not known prior to checking this field. To overcome this, a check is performed to determine if the register&#8217;s value is negative, in which case the most-significant bit must be set to 1. Therefore the number of bytes in <code>a1<\/code> will be multiplied by 2 (by shifting it to the left by 1 bit on line <a href=\"#coderef-scale-xlen\">26<\/a>), the value of <code>t0<\/code> is shifted by 1 to the left, and the check is performed again. If the value of <code>t0<\/code> is determined to be positive on line <a href=\"#coderef-have-xlen\">25<\/a> (i.e. the MSB is 0), then the <code>__system_check<\/code> function has completed its probe. The following listing illustrates the main program which performs the system check:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>1:         .section \".text.init\"\n2:         .align 2\n3:         .global _start\n4:         .global _stack_end\n5: _start:\n6:         la      sp, _stack_end\n7:         call    __system_check\n8: stop:   j       stop\n9: \n<\/code><\/pre>\n\n\n\n<p>The <code>__system_check<\/code> function is called on line <a href=\"#coderef-call-systemcheck\">7<\/a>. When this function returns the register <code>a0<\/code> should hold the number of supported privilege modes, and register <code>a1<\/code> should hold the number of bytes in a register. If this program is run in QEMU, the <code>info registers<\/code> command can be used at the console to determine the hardware details:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>make chapter5\nqemu-system-riscv64 -M virt -serial \/dev\/null -nographic -kernel chapter5.elf\nQEMU 3.1.0 monitor - type 'help' for more information\n(qemu) info registers\n pc       000000008000000c\n mhartid  0000000000000000\n mstatus  0000000000000000\n mip      0000000000000000\n mie      0000000000000000\n mideleg  0000000000000000\n medeleg  0000000000000000\n mtvec    0000000000000000\n mepc     0000000000000000\n mcause   0000000000000000\n zero 0000000000000000 ra   000000008000000c sp   0000000080009000 gp   0000000000000000\n tp   0000000000000000 t0   000000000028225a t1   0000000000040000 t2   0000000000000000\n s0   0000000000000000 s1   0000000000000000 a0   0000000000000003 a1   0000000000000008\n...\n<\/code><\/pre>\n\n\n\n<p>The value of <code>a0<\/code> is 3, therefore the QEMU VirtIO platform supports all three privilege modes. The value of <code>a1<\/code> is 8, which means that the register width is 64-bits (8 bytes). This value could be used as a stack offset allowing for the definition of the function call preamble and postamble as part of a library. This library could then be used to define an operating system&#8217;s ABI.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"org59973e7\">It&#8217;s a Trap!<\/h2>\n\n\n\n<p>Typically, systems should run in the most restricted environment possible in order to minimize catastrophes in the event of a system fault. However, when an exceptional event occurs, the system may want to raise its privilege level in order to deal with it. These types of events are often associated with an interrupt.<\/p>\n\n\n\n<p>The machine-mode status CSR, <code>mstatus<\/code>, allows some control over a hart&#8217;s operating state, including enabling or disabling global interrupts. Three fields are defined in the register&#8217;s least significant four bits for this purpose; one for each of the privilege modes.<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p>specific interrupt types for each privilege mode must be enabled individually via the <code>mie<\/code> register which will be described later<\/p>\n<\/blockquote>\n\n\n\n<p>The fields defined in the <code>mstatus<\/code> register are illustrated in the following table:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table><thead><tr><th>Bit<\/th><th>0<\/th><th>1<\/th><th>2<\/th><th>3<\/th><th>4<\/th><th>5<\/th><th>6<\/th><th>7<\/th><\/tr><\/thead><tbody><tr><td>0<\/td><td>UIE<\/td><td>SIE<\/td><td>&nbsp;<\/td><td>MIE<\/td><td>UPIE<\/td><td>SPIE<\/td><td>&nbsp;<\/td><td>MPIE<\/td><\/tr><tr><td>+8<\/td><td>MPP<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><td>MPP[0]<\/td><td>MPP[1]<\/td><td>FS[0]<\/td><td>FS[1]<\/td><td>XS[0]<\/td><\/tr><tr><td>+16<\/td><td>XS[1]<\/td><td>MPRV<\/td><td>SUM<\/td><td>MXR<\/td><td>TVM<\/td><td>TW<\/td><td>TSR<\/td><td>&nbsp;<\/td><\/tr><tr><td>+24<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><\/tr><tr><td>+32<\/td><td>UXL[0]<\/td><td>UXL[1]<\/td><td>SXL[0]<\/td><td>SXL[1]<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><\/tr><tr><td>+40<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><\/tr><tr><td>+48<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><\/tr><tr><td>+56<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><td>&nbsp;<\/td><td>SD<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p>The <strong>UIE<\/strong>, <strong>SIE<\/strong>, and <strong>MIE<\/strong> fields will enable interrupts globally for the user, supervisor, and machine modes respectively. If the <em>x<\/em>IE fields value is set to 1, interrupts will be enabled globally for privilege mode <em>x<\/em> and any privilege mode <em>y<\/em>&lt;<em>x<\/em> provided <em>y<\/em>IE is also set to 1. If the hart is operating at privilege level <em>x<\/em>, interrupts at privilege levels inferior to <em>x<\/em> will be disabled regardless of the state of the associated interrupt bit in <code>mstatus<\/code>.<\/p>\n\n\n\n<p>To be of any use, there must be a mechanism to handle interrupts. The <strong>BASE<\/strong> field of the <code>mtvec<\/code> CSR can be set to the base address of a trap-vector to handle interrupts. The <strong>MODE<\/strong> field, that occupies the least-significant two bits of <code>mtvec<\/code>, specify the trap mode. When the <strong>MODE<\/strong> field is set to 0, the trap mode will be set to call the handler directly. Otherwise the base address expresses the base of a vector of trap handlers; one for each interrupt type indexed by the interrupt code. For the time being, a single interrupt handler will be used to dispatch to the appropriate handler for interrupts or exceptions.<\/p>\n\n\n\n<p>The code that follows illustrates a skeleton for a trap handler implementation in direct mode:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code> 1:         .text\n 2:         .align 2\n 3:         .global trap_handler\n 4: trap_handler:\n 5:         # Trap handler preamble (save registers).\n 6:         csrrw   a0, mscratch, a0\n 7:         sd      a1, 0(a0)\n 8:         sd      a2, 8(a0)\n 9:         sd      a3, 16(a0)\n10:         sd      a4, 24(a0)\n11:         # Decode the cause of the interrupt.\n12:         csrr    a1, mcause\n13:         bgez    a1, exception\n14: interrupt:\n15:         andi    a1, a1, 0x3f # Isolate the cause field\n16:         # TODO: Dispatch to specific interrupt handler\n17:         j       trap_handler_restore_state\n18: exception:\n19:         addi    a1, a1, 0x3f # Isolate the cause field\n20:         # TODO: Dispatch to specific exception handler\n21: trap_handler_restore_state:\n22:         ld      a4, 24(a0)\n23:         ld      a3, 16(a0)\n24:         ld      a2, 8(a0)\n25:         ld      a1, 0(a0)\n26:         csrrw   a0, mscratch, a0\n27:         mret\n<\/code><\/pre>\n\n\n\n<p>The first instruction on line <a href=\"#coderef-trap-load scratch\">6<\/a> will atomically swap the values of the <code>mscratch<\/code> CSR and <code>a0<\/code>. This register is defined to provide additional data to trap handlers. Typically, this should be set to an memory address of a buffer where register data can be saved while the handler is active.<\/p>\n\n\n\n<p>The next four lines save the contents of registers <code>a1-a4<\/code> in the memory buffer located at the address in <code>mscratch<\/code>, then the value of the <code>mcause<\/code> CSR is copied into <code>a1<\/code> on line <a href=\"#coderef-trap-get cause\">12<\/a>. The <code>mcause<\/code> register is used to indicate the cause of synchronous and asynchronous exceptions. If the most significant bit in this register is zero, then the trap was caused by a synchronous exception. This can be determined by testing the value of <code>a1<\/code> to see if it is greater than or equal to zero (on line <a href=\"#coderef-trap-check exception\">13<\/a>; if the MSB is 1, the register value is negative and the execution will fall through to the interrupt handler code.<\/p>\n\n\n\n<p>To use the trap handler, it&#8217;s address must be set in the base field of the <code>mtvec<\/code> CSR. This is illustrated in the following program:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code> 1:         .section \".text.init\"\n 2:         .align 2\n 3:         .global _start\n 4:         .global _stack_end\n 5: _start:\n 6:         la      sp, _stack_end\n 7:         la      a0, scratch\n 8:         csrrw   a0, mscratch, a0\n 9:         la      a0, trap_handler\n10:         csrrw   a0, mtvec, a0\n11:         call    __system_check\n12: stop:   j       stop\n13:         .bss\n14: scratch:        .dword 0, 0, 0, 0\n<\/code><\/pre>\n\n\n\n<p>The scratch area where the register state can be saved is defined on line <a href=\"#coderef-define scratch\">14<\/a>. This allocates 32 bytes of space to save the contents of up to 4 registers which is enough for the current <code>trap_handler<\/code> implementation. The address of the scratch is loaded into register <code>a0<\/code> on line <a href=\"#coderef-load scratch\">7<\/a>. This register&#8217;s value is then swapped with the value of the <code>mscratch<\/code> CSR on line <a href=\"#coderef-set scratch\">8<\/a>.<\/p>\n\n\n\n<p>The address of the trap handler function is loaded into register <code>a0<\/code> on line <a href=\"#coderef-load trap handler\">9<\/a>, then it is swapped with the contents of the <code>mtvec<\/code> CSR on line <a href=\"#coderef-set trap handler\">10<\/a>. The <strong>MODE<\/strong> field is left as zero to set the trap mode to direct; which will cause all synchronous and asynchronous interrupts to branch to the base address.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"orga01113b\">What&#8217;s the Time?<\/h2>\n\n\n\n<p>A platform&#8217;s real-time counter is one of the possible sources of asynchronous exceptions that can cause an interrupt. The timer is typically external to the processing core. The VirtIO machine of the QEMU emulator includes the core-level interrupt module (CLINT). This module defines the <code>mtime<\/code> register that exposes the current value of the real-time counter. This value expresses the number of clock cycles that have elapsed since the processor was reset. This does not represent the real time, but a count of real-time intervals (determined by the oscillator frequency).<\/p>\n\n\n\n<p>The <strong>mtime<\/strong> register is mapped to a particular address in physical memory. The actual address is specified in the memory map of the VirtIO machine in the QEMU system (see <a href=\"https:\/\/git.qemu.org\/?p=qemu.git;a=blob_plain;f=hw\/riscv\/virt.c;hb=refs\/heads\/stable-3.1\">hw\/riscv\/virt.c<\/a>). The listing that follows illustrates a function to retrieve the current real-time counter value:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code> 1:         .equ    CLINT_BASE, 0x2000000 # The base address of the CLINT module\n 2:         .equ    CLINT_MTIME, 0xbff8   # The offset of the MTIME register\n 3:         .macro  ld_mtime rd # Macro to access the MTIME memory mapped register\n 4:         li      t0, CLINT_BASE\n 5:         li      t1, CLINT_MTIME\n 6:         add     t0, t0, t1 # Determine the absolute address of MTIME\n 7:         ld      \\rd, 0(t0) # Read the counter value\n 8:         .endm\n 9: \n10:         .text\n11:         .align 2\n12:         .global __clock_cycle\n13: __clock_cycle:\n14:         # Retrieve the current value of the real-time counter.\n15:         #\n16:         # Inputs: None\n17:         #\n18:         # Returns:\n19:         # - a0: The current value of the mtime register.\n20:         ld      ld_mtime a0\n21:         ret\n<\/code><\/pre>\n\n\n\n<p>The <code>CLINT_BASE<\/code> symbol is defined on line <a href=\"#coderef-clint base\">1<\/a>, its value is the absolute base address of the memory mapped registers of the CLINT module (see <a href=\"https:\/\/git.qemu.org\/?p=qemu.git;a=blob_plain;f=include\/hw\/riscv\/sifive_clint.h;hb=refs\/heads\/stable-3.1\">include\/hw\/riscv\/sifive_clint.h<\/a> of the QEMU source). The <code>CLINT_MTIME<\/code> symbol, defined on line <a href=\"#coderef-mtime offset\">2<\/a>, specifies the offset of the <code>mtime<\/code> register relative to the CLINT&#8217;s base memory address. The offset is added to the base address and its result stored in register <code>t0<\/code> on line <a href=\"#coderef-mtime addr\">6<\/a>. Finally the value of the the real-time counter is retrieved on line <a href=\"#coderef-read counter\">7<\/a>. All of these operations are defined in a macro starting on line <a href=\"#coderef-ld-mtime macro\">3<\/a>.<\/p>\n\n\n\n<p>The <code>mtime<\/code> register is useful to determine the time since the board was reset. However, it cannot generate interrupts by itself. The <code>mtimecmp<\/code> memory mapped register will cause a timer interrupt to be posted when its value is less than the value contained in <code>mtime<\/code>. In other words, if when the periodically increasing value of <code>mtime<\/code> exceeds that contained in <code>mtimecmp<\/code>, a timer interrupt will be posted (provided that timer interrupts are enabled). Therefore to receive an interrupt after some fixed interval of time, the current value of <code>mtime<\/code> must be retrieved, then the timeout value must be added thereto and the result saved in the <code>mtimecmp<\/code> register. This process is illustrated in the following listing:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code> 1:         .equ    CLINT_MTIMECMP, 0x4000 # The offset of the MTIMECMP register\n 2:         .macro st_mtimecmp rs\n 3:         li      t0, CLINT_BASE\n 4:         li      t1, CLINT_MTIMECMP\n 5:         add     t1, t0, t1\n 6:         sd      \\rs, 0(t1)\n 7:         .endm\n 8:         .global __timer_create\n 9: __timer_create:\n10:         # Set a timer to trigger an interrupt when a given number of\n11:         # clock cycles have elapsed.\n12:         #\n13:         # Inputs:\n14:         # - a0: The timeout value in clock cycles.\n15:         ld_mtime        t1\n16:         ld      t1, 0(t1)\n17:         add     a0, a0, t1             # Add the timeout to the clock cycle\n18:         st_mtimecmp     a0\n19:         ret\n<\/code><\/pre>\n\n\n\n<p>The <code>__timer_create<\/code> function will set a timer to trigger an interrupt after the given number of clock cycles have elapsed. This code re-uses the macro defined previously to read the current value of the real-time counter, then adds the desired number of cycles thereto, and writes the result to the <code>mtimecmp<\/code> register. When <code>mtime<\/code>&#8216;s value is greater than the value in <code>mtimecmp<\/code>, the <strong>MTIP<\/strong> field of the <code>mip<\/code> CSR register will be asserted to indicate that a timer interrupt is pending, and the trap handler will be called. The following code will set up this process:<\/p>\n\n\n\n<pre id=\"org0bbc010\" class=\"wp-block-code\"><code> 1:         .section \".text.init\"\n 2:         .align 2\n 3:         .global _start\n 4:         .global _stack_end\n 5:         .global CLOCK_MONOTONIC\n 6: _start:\n 7:         la      sp, _stack_end\n 8:         la      t0, __mtrap_handler # Load trap vector address\n 9:         csrrw   zero, mtvec, t0\n10:         li      t0, 0b1&lt;&lt;3\n11:         csrrs   t0, mstatus, t0 # Enable interrupts globally (ref:set-mstatus.MIE)\n12:         li      t0, 0b1&lt;&lt;7\n13:         csrrs   t0, mie, t0     # Enable timer interrupts (ref:set-mie.MTIE)\n14: 1:      li      a0, 0x10000     # Set the timeout value\n15:         call    __timer_create  # Set the timer\n16:         wfi                     # Wait for interrupts\n17:         j       1b\n18: \n19:         .align 2\n20: __mtrap_handler: # Machine interrupt handler\n21:         csrrc   t0, mcause, zero # Get the cause of the interrupt\n22:         bgez    t0, 2f # Exit on an exception\n23:         slli    t0, t0, 1\n24:         srli    t0, t0, 1\n25:         li      t1, 7 # The timer interrupt has code 7.\n26:         bne     t0, t1, 2f # Check for timer interrupts\n27:         addi    s0, s0, 1 # Increment the interrupt count.\n28: 2:      mret    # Machine trap return\n<\/code><\/pre>\n\n\n\n<p>After setting up the stack, this program loads the address for the trap handler on line <a href=\"#coderef-load-mtrap\">8<\/a>, and stores this address in the <code>mtvec<\/code> CSR. Machine-mode interrupts are then enabled globally by setting bit 3 of the <code>mstatus<\/code> CSR to 1 (on line <a href=\"#coderef-set-mstatus.MIE\">11<\/a>), and M-mode timer interrupts are enabled (on line <a href=\"#coderef-set-mie.MTIE\">13<\/a>) by setting bit 7 of the <code>mie<\/code> CSR to 1. On line <a href=\"#coderef-load-timeout\">14<\/a>, an immediate value is loaded into register <code>a0<\/code>. This value is used to create the timer on line <a href=\"#coderef-trap-settimer\">15<\/a>. Once the timer is armed, the <code>wfi<\/code> instruction is invoked on line <a href=\"#coderef-wfi\">16<\/a> to wait for an exception to occur at which point control jumps to the interrupt handler. When the handler returns, control jumps back to line <a href=\"#coderef-load-timeout\">14<\/a>, and the timer is armed again. This should cause periodic calls to the trap vector.<\/p>\n\n\n\n<p>A machine-level trap handler is defined on line <a href=\"#coderef-mtrap-handler\">20<\/a>. This handler will increment the value in register <code>s0<\/code> every time a timer exception is triggered. The cause of the interrupt is determined on line <a href=\"#coderef-trap-get-cause\">21<\/a> which atomically reads the value of the <code>mcause<\/code> CSR and sets its value to zero. If the value of this register was greater-than, or equal to, zero (line <a href=\"#coderef-trap-exception\">22<\/a>), the interrupt was caused by a synchronous exception (whereby the most-significant bit of the register will be set to 1). In the case of a synchronous exception, the handler simply exists by calling the <code>mret<\/code> instruction which returns control to the instruction that was executing when the exception occurred. Otherwise the next two lines will shift off the most significant bit (i.e. set it to zero). The interrupt code is checked on line <a href=\"#coderef-trap-check-mtip\">26<\/a>, if it corresponds with the timer interrupt, the value of <strong>s0<\/strong> is incremented by 1.<\/p>\n\n\n\n<p>The state of the hart can be inspected via the <code>info registers<\/code> command in QEMU. If the timing of the snapshot is such that the trap handler is executing the machine CSRs will have the following state:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>(qemu) info registers\n pc       000000008000004c\n mhartid  0000000000000000\n mstatus  0000000000001880\n mip      0000000000000080\n mie      0000000000000080\n mideleg  0000000000000000\n medeleg  0000000000000000\n mtvec    0000000080000048\n mepc     0000000080000040\n mcause   8000000000000007\n<\/code><\/pre>\n\n\n\n<p>As usual, the <code>pc<\/code> register shows the current address of the active instruction, however, in this case control has jumped into the trap handler. When an exception is triggered, the following operations are executed atomically:<\/p>\n\n\n\n<ol class=\"wp-block-list\">\n<li>Interrupts are disabled globally (bit 3 of <code>mstatus<\/code> is set to 0).<\/li>\n\n\n\n<li>The <code>mstatus<\/code><strong>.MPIE<\/strong> field (which represents the previous global interrupt mode) is set the value of <code>mstatus<\/code><strong>.MIE<\/strong><\/li>\n\n\n\n<li>Interrupts are disabled globally by setting the <code>mstatus<\/code><strong>.MIE<\/strong> field is set to zero.<\/li>\n\n\n\n<li>The <code>mstatus<\/code><strong>.MPP<\/strong> field (bits 12:11) is set to <code>0b11<\/code> which indicates the privilege level prior to the exception being raised.<\/li>\n\n\n\n<li>The address of the instruction that follows the last one to execute before the exception was raised is saved in the <code>mepc<\/code> register.<\/li>\n\n\n\n<li>The cause of the exception is written in register <code>mcause<\/code>.<\/li>\n\n\n\n<li>The <strong>pc<\/strong> register is set to the address in <code>mtvec<\/code>.<\/li>\n<\/ol>\n\n\n\n<p>When the interrupt handler completes its task, the <code>mret<\/code> instruction is executed which will reverse this process. Control will be set to the address in <code>mepc<\/code>, and the value of <code>mstatus<\/code><strong>.MPIE<\/strong> will be written to <code>mstatus<\/code><strong>.MIE<\/strong> and cleared. The <code>mip<\/code> register will be cleared to indicate that there are no longer interrupts pending. Control flow will continue from where it was interrupted, and the core will be ready to handle new interrupts. More importantly, The hart will be returned to the privilege mode specified in the <code>mstatus<\/code><strong>.MPP<\/strong> field. This behaviour can be exploited to set the current privilege mode of the processor.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"orgd8dddb4\">Enter the User<\/h2>\n\n\n\n<p>The default privilege level when the processor is reset is machine mode. Programs running at this level have full control over the processor via the control and status registers. However, this opens up the system to abuse. To limit any damage that is possible by a wayward program, it is better to run in user mode. Getting to user mode is a simple matter of setting up the <code>mstatus<\/code> register and returning from M-mode via the <code>mret<\/code> instruction. The following macro will set the priviledge mode to the specified level:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code> 1:         # Set the privilege mode to that specified by the immediate\n 2:         # value.\n 3:         .macro  setmode imm\n 4:         li      t0, 0x1800 #(ref:clear mstatus.MPP)\n 5:         li      t0, \\imm\n 6:         slli    t0, t0, 11\n 7:         csrrs   t0, mstatus, t0 #(ref:set mstatus.MPP)\n 8:         la      t0, 1f\n 9:         csrrw   zero, mepc, t0\n10:         mret\n11: 1:\n12:         .endm\n<\/code><\/pre>\n\n\n\n<p>This macro will set the privilege mode to the value specified as an immediate by first clearing the <code>mstatus<\/code><strong>.MPP<\/strong> field on line <a href=\"#coderef-clear mstatus.MPP\">4<\/a>, then replacing it with the encoding of the desired privilege mode. The following table lists the privilege modes and their encodings:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table><thead><tr><th>Level<\/th><th>Name<\/th><th>Encoding<\/th><th>Immediate<\/th><\/tr><\/thead><tbody><tr><td>U<\/td><td>User\/Application<\/td><td><code>0b00<\/code><\/td><td>0<\/td><\/tr><tr><td>S<\/td><td>Supervisor<\/td><td><code>0b01<\/code><\/td><td>1<\/td><\/tr><tr><td>M<\/td><td>Machine<\/td><td><code>0b11<\/code><\/td><td>3<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p>The fourth column of this table lists the immediate value that should be supplied to the macro to set the associated privilege mode. The encoded privilege value is loaded into register <code>t0<\/code> on line <a href=\"#coderef-load priv-mode\">5<\/a>, then shifted to the right position on the next line. This field is set via the <code>csrrs<\/code> instruction on line <a href=\"#coderef-set mstatus.MPP\">7<\/a> which effectively sets <code>mstatus<\/code> to the bit-wise <strong>OR<\/strong> of its previous value with the value in <code>t0<\/code>. However, before returning from machine-mode, the <code>mepc<\/code> register must be updated with the address of the instruction to which control will return.<\/p>\n\n\n\n<p>The address immediately following the <code>mret<\/code> instruction is loaded into register <code>t0<\/code> on line <a href=\"#coderef-load return address\">8<\/a>, then written to <code>mepc<\/code> on line <a href=\"#coderef-set mepc\">9<\/a>. When the <code>mret<\/code> instruction is executed, <code>pc<\/code> will be set to this address. If the main program is updated to invoke this macro with the argument 0 just before the <code>wfi<\/code>, the program should be in U-mode by the time the timer expires:<\/p>\n\n\n\n<pre id=\"org6f4fee5\" class=\"wp-block-code\"><code>26: 1:      li      a0, 0x10000    # Set the timeout value\n27:         call    __timer_create # Set the timer\n28:         setmode 0\n29:         wfi\n30:         j       1b\n31: \n<\/code><\/pre>\n\n\n\n<p>Anything that runs following the call to the <code>setmode<\/code> macro will be executing in U-mode. However, the trap handler will execute in machine mode. Therefore the trap handler is useful for performing tasks that require machine mode privilege. Fortunately, traps can occur for asynchronous interrupts as well as synchronous exceptions. Therefore the trap handler can be used to implement system calls.<\/p>\n\n\n\n<p>The <code>ecall<\/code> instruction is an environment call which raises a synchronous exception, and sets <code>mcause<\/code> to the code indicating the active privilege mode when it was executed. The trap handler can be updated to jump to the appropriate function, which will execute in machine mode, then returning to the original pivilege mode when it is complete. The trap handler must be updated to handle the system call:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code> 1:         .align 2\n 2: __mtrap_handler:\n 3:         csrr    t0, mcause\n 4:         bgez    t0, 1f\n 5:         slli    t0, t0, 1\n 6:         srli    t0, t0, 1\n 7:         li      t1, 7\n 8:         bne     t0, t1, 2f\n 9:         addi    s0, s0, 1\n10:         j       2f\n11: 1:\n12:         li      t1, 8\n13:         bne     t0, t1, 2f\n14:         push_stack\n15:         call    __syscall\n16:         pop_stack\n17:         csrr    t0, mepc\n18:         addi    t0, t0, 4\n19:         csrrw   zero, mepc, t0\n20: 2:      csrrw   t0, mcause, zero\n21:         mret\n22: \n<\/code><\/pre>\n\n\n\n<p>This handler is updated by adding some code for synchronous exceptions (starting at line <a href=\"#coderef-U-mode ecall code\">12<\/a>). The value of <code>mcause<\/code> is compared with the user-environment call exception code (8) to see if a system call was requested. If so, the handler will save the registers on the stack via the <code>push_stack<\/code> macro on line <a href=\"#coderef-save registers\">14<\/a>, call the <code>__syscall<\/code> function, then resture the registers to their values prior to the call via the <code>pop_stack<\/code> macro on line <a href=\"#coderef-restore registers\">16<\/a>.<\/p>\n\n\n\n<p>After the system call has finished processing, the handler loads the value of the <code>mepc<\/code> CSR which should contain the address of the <code>ecall<\/code> instruction that caused the trap. This address is incremented by 4 to skip to the instruction that follows <code>ecall<\/code> on line <a href=\"#coderef-Set mret address\">18<\/a>, then stores the updated address in <code>mepc<\/code> before executing <code>mret<\/code>. This will set the program counter to the instruction immediately following the one that triggered the system call, and restore the privilege mode to U-mode.<\/p>\n\n\n\n<p>Typically system calls are identified by a number. If its arguments are stored in the registers <code>a0<\/code> to <code>a7<\/code>, and its return value in <code>a0<\/code>, system calls will follow the same convention as regular function calls (albeit with greater privilege). The following code snippet will invoke the system call associated with the identifier <code>0x100<\/code>:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>1: li      a0, 0x100\n2: ecall\n<\/code><\/pre>\n\n\n\n<p>The timer can now be set from user mode via a system call. The following implementation of <code>__syscall<\/code> will invoke <code>__timer_create<\/code> with the timeout value specified in register <code>a1<\/code> when system call <code>0x100<\/code> is requested:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code> 1: __syscall:\n 2:         mv      s1, a0\n 3:         li      t0, 0x100\n 4:         bne     t0, a0, 1f\n 5:         mv      a0, a1\n 6:         push_stack\n 7:         call    __timer_create\n 8:         pop_stack\n 9: 1:      ret\n10: \n<\/code><\/pre>\n\n\n\n<p>This function will check the register <code>a0<\/code> to determine the system call id that is requested. If it is <code>0x100<\/code>, then it will move the timeout value from register <code>a1<\/code> to <code>a0<\/code> on line <a href=\"#coderef-set timeout arg\">5<\/a>, push the stack, then invoke the <code>__timer_create<\/code> function with M-mode privilege. The main program can be updated to set the privilege level to U-mode, then make a system call <code>0x100<\/code> to set the timer. The updated main program is illustrated in the following listing:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code> 1:         .section \".text.init\"\n 2:         .align 2\n 3:         .global _stack_end\n 4:         .global CLOCK_MONOTONIC\n 5:         .global _start\n 6: _start:\n 7:         la      sp, _stack_end\n 8:         call    __system_check\n 9:         la      t0, __mtrap_handler\n10:         csrrw   zero, mtvec, t0\n11:         li      t0, 0b1&lt;&lt;3\n12:         csrrs   zero, mstatus, t0\n13:         li      t0, 0b1&lt;&lt;7\n14:         csrrs   zero, mie, t0\n15:         setmode 0\n16: 1:      li      a0, 0x100\n17:         li      a1, 0x10000\n18:         ecall\n19:         wfi\n20:         j       1b\n<\/code><\/pre>\n\n\n\n<p>The major difference from the previous main program is that the timer is not set directly, but via a system call. The syscall id is loaded into register <code>a0<\/code> on line <a href=\"#coderef-set timer syscall\">16<\/a>, and the timeout value into register <code>a1<\/code> on line <a href=\"#coderef-set system timeout\">17<\/a>. This set of instructions will be repeated each time a timeout is triggered.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"orgdb17f69\">Conclusion<\/h2>\n\n\n\n<p>This chapter has delved into how RISC-V handles synchronous and asynchronous exceptions as well as the privilege mode instructions available in the ISA. This may be one of the key components in the development of an operating system by allowing privileged functions to run separately from user code. System calls allow U-mode applications to request services which require M-mode (or S-mode) privilege to execute via the <code>ecall<\/code> instruction.<\/p>\n\n\n\n<p>Moreover, the asynchronous exception handling provides a good introduction into how RISC-V processors deal with events originating from external peripherals. This will become important when creating programs intended to interact with the system user.<\/p>\n\n\n\n<p>Although the examples in this chapter were restricted to M-mode and U-mode privilege levels. The supervisor mode was briefly discussed. Having three privilege levels is useful when creating hypervisors: guest operating systems can operate in S-mode while the virtualization environment uses M-mode.<\/p>\n\n\n\n<p>The next chapter will leverage many of the capabilities discussed here to interface with more of the external components in a RISC-V system. In particular the UART module will allow users to interact with applications via a serial console.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Up to this point, the RISC-V tutorial has focused on single applications running on a single hardware thread. The application&#8217;s environment was composed of the processor&#8217;s state and memory map. The processor state was controlled via assembly instructions, and the memory map was defined at build time via a linker script. However, modern systems are [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_jetpack_memberships_contains_paid_content":false,"footnotes":""},"categories":[4],"tags":[],"class_list":["post-53","post","type-post","status-publish","format-standard","hentry","category-risc-v"],"jetpack_featured_media_url":"","jetpack_sharing_enabled":true,"_links":{"self":[{"href":"https:\/\/www.vociferousvoid.org\/index.php\/wp-json\/wp\/v2\/posts\/53","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.vociferousvoid.org\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.vociferousvoid.org\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.vociferousvoid.org\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.vociferousvoid.org\/index.php\/wp-json\/wp\/v2\/comments?post=53"}],"version-history":[{"count":2,"href":"https:\/\/www.vociferousvoid.org\/index.php\/wp-json\/wp\/v2\/posts\/53\/revisions"}],"predecessor-version":[{"id":55,"href":"https:\/\/www.vociferousvoid.org\/index.php\/wp-json\/wp\/v2\/posts\/53\/revisions\/55"}],"wp:attachment":[{"href":"https:\/\/www.vociferousvoid.org\/index.php\/wp-json\/wp\/v2\/media?parent=53"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.vociferousvoid.org\/index.php\/wp-json\/wp\/v2\/categories?post=53"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.vociferousvoid.org\/index.php\/wp-json\/wp\/v2\/tags?post=53"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}