{"id":33,"date":"2019-11-30T01:41:00","date_gmt":"2019-11-30T06:41:00","guid":{"rendered":"https:\/\/www.vociferousvoid.org\/?p=33"},"modified":"2023-07-21T22:41:13","modified_gmt":"2023-07-22T02:41:13","slug":"risc-v-bare-metal-programming-chapter-4-another-brick-in-the-wall","status":"publish","type":"post","link":"https:\/\/www.vociferousvoid.org\/index.php\/2019\/11\/30\/risc-v-bare-metal-programming-chapter-4-another-brick-in-the-wall\/","title":{"rendered":"RISC-V Bare Metal Programming &#8211; Chapter 4: Another Brick in the Wall"},"content":{"rendered":"\n<p><a href=\"https:\/\/www.vociferousvoid.org\/index.php\/2019\/11\/12\/risc-v-bare-metal-programming-chapter-3-a-link-to-the-past\/\">Chapter 3<\/a> of this RISC-V bare metal tutorial studied the linking process and how a developer can control where code and data are placed in memory. Constants, initialized variables and uninitialized variables were defined and explicitly positioned in RAM as prescribed by a linker script. The running example program was updated to read operands from RAM to perform its task, and subsequently store the result in a different location in RAM. However, up to this point only the base RV64I instruction set has been used. This chapter will explore some of the standard extensions available in the RISC-V ISA.<\/p>\n\n\n\n<p>One of the objectives in the design of the RISC-V ISA is to support many different deployment environments which may have varying constraints for efficiency, performance, and cost. For this reason, the base instruction set was restricted to the minimum required to build a useful program. This reduces the processor complexity potentially yielding performance and efficiency gains. However, these gains may be lost when performing more complex computations. To address potential limitations in the base instruction set, optional standard extensions have been defined to expand the available set of instructions. The standard extensions available for 32 and 64-bit instruction sets include: <strong>M<\/strong> Support for multiply and divide (RV32M and RV64M). <strong>A<\/strong> Atomic operations (RV32A and RV64A). <strong>F<\/strong> Floating point support (RV32F and RV64F). <strong>D<\/strong> Double precision floating point support (RV32D and RV64D).<\/p>\n\n\n\n<p>This set of standard extensions are typically included in most implementations of RISC-V cores. The base set plus these extensions is often referred to as the <strong>G<\/strong> instruction set (RV32G or RV64G). Each of these standard extensions will be explored in this chapter.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"org657fc9a\">Multiply<\/h2>\n\n\n\n<p>The <strong>M<\/strong> extension provides instructions for multiplying and dividing integers using both word and double-word length operands. When using word length operands, the result will not require more than 64-bits of memory which fits in an RV64I registers. The following listing of the <code>product.s<\/code> source file shows the assembly code of a function to multiply word sized integer operands:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code> 1:         .text\n 2:         .align 2\n 3:         .global __imul32\n 4: __imul32:\n 5:         # Input:\n 6:         # a0: 32-bit multiplicand\n 7:         # a1: 32-bit multiplier\n 8:         # Result:\n 9:         # a0: 64-bit product\n10:         addi    sp, sp, -32\n11:         sd      ra, 24(sp)\n12:         mulw    a0, a0, a1\n13:         ld      ra, 24(sp)\n14:         addi    sp, sp, 32\n15:         ret<\/code><\/pre>\n\n\n\n<p>Due to the fact that the arguments of this function are expected to be word-length data, the calculation of the product can be performed using a single instruction (<code>mulw<\/code> on line <a href=\"#coderef-multiplication\">12<\/a>). The main program can be updated as follows to invoke the <code>__imul32<\/code> function:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code> 1:         .section \".text.init\"\n 2:         .align 2\n 3:         .global _start\n 4:         .global _stack_end\n 5: _start:\n 6:         lw      a0, operand1\n 7:         lw      a1, operand2\n 8:         la      sp,_stack_end\n 9:         call    sum\n10:         la      t1, result1\n11:         sw      a0, 0(t1)\n12:         call    __imul32\n13:         la      t1, result2\n14:         sd      a0, 0(t1)\n15: stop:   j       stop\n16:         .section \".rodata\"\n17: operand1:       .word   4\n18:         .data\n19: operand2:       .word   5\n20:         .bss\n21: result1:        .word   0\n22: result2:        .dword  0<\/code><\/pre>\n\n\n\n<p>The <code>.bss<\/code> section of the ELF file was updated to declare two result variables: <code>result1<\/code> on line <a href=\"#coderef-sum_result\">21<\/a> which will hold the sum of the operands in a word, and <code>result2<\/code> on line <a href=\"#coderef-__imul32_result\">22<\/a> which will hold their product in a double-word.<\/p>\n\n\n\n<p>After the sum of the operands is calculated, and the result is saved in memory, it is kept in register <code>a0<\/code> to be used as the multiplicand. The value of <code>operand2<\/code> will be used as the multiplier; its value should still be in the <code>a1<\/code> register since its content is not modified by the <code>sum<\/code> function. The <code>__imul32<\/code> function is then called on line <a href=\"#coderef-call_imul32\">12<\/a> and the result is saved in memory at line <a href=\"#coderef-save_imul32\">14<\/a>.<\/p>\n\n\n\n<p>The program can be compiled and executed in <kbd>QEMU<\/kbd> using the following sequence of commands:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>riscv64-unknown-elf-as -o add.o add.s\nriscv64-unknown-elf-as -o main.o main.s\nriscv64-unknown-elf-as -o product.o product.s\nriscv64-unknown-elf-ld -T chapter3.lds -o main.elf add.o main.o product.o\nqemu-system-riscv64 -M virt -serial \/dev\/null -nographic -kernel main.elf\nQEMU 3.1.0 monitor - type 'help' for more information\n(qemu) <\/code><\/pre>\n\n\n\n<p>The <code>chapter3.lds<\/code> linker script is the same one that was used in chapter 3. The result values can be inspected from the QEMU console using the <kbd>xp<\/kbd> command:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>(qemu) xp \/1wd 0x80001004\n0000000080001004:          9\n(qemu) xp \/1gd 0x80001008\n0000000080001008:         45\n(qemu) <\/code><\/pre>\n\n\n\n<p>The location of <code>result1<\/code> in memory is the same as <code>result<\/code> from the previous chapter. The memory location of <code>result2<\/code> will be 4-bytes beyond <code>result1<\/code> since this value is 32-bits wide. Therefore the product result can be found at memory offset <code>0x80001008<\/code>. This can easily be verified using the <kbd>objdump<\/kbd> utility:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>$ riscv64-unknown-elf-objdump -D -j.bss main.elf \n\nsum.elf:     file format elf64-littleriscv\n\n\nDisassembly of section .bss:\n\n0000000080001004 &lt;result1&gt;:\n    80001004:\t0000                \tunimp\n\t...\n\n0000000080001008 &lt;result2&gt;:\n\t...\n<\/code><\/pre>\n\n\n\n<p>As expected the multiplication of 9 and 5 is 45.<\/p>\n\n\n\n<p>Multiplication using registers is a little more complicated when dealing with 64-bit values. This is due to the fact that the product will be wider (in bits) than either the multiplier or multiplicand. The <code>__imul32<\/code> function assumes that the operands are word-length values, therefore the result will fit in a single double-word register. However, the calculated product will be truncated if double-word length operands are provided. The product of two 64-bit values may have as many as 128 bits which is wider than any available register in the RV64I instruction set. To mitigate this problem, the RISC-V ISA requires two instructions to perform a multiplication: one to calculate the most significant double-word (<code>mulh<\/code>), and a second to calculate the least significant double-word (<code>mul<\/code>). The following listing illustrates the <code>__imul64<\/code> function that can handle 64-bit operands:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code> 1:         .global __imul64\n 2: __imul64:\n 3:         # Input:\n 4:         # a0: 64-bit multiplicand\n 5:         # a1: 64-bit multiplier\n 6:         # Result:\n 7:         # a0: low 64-bits of the product\n 8:         # a1: high 64-bits of the product\n 9:         addi    sp, sp, -32\n10:         sd      ra, 24(sp)\n11:         sd      t1, 16(sp)\n12:         sd      t0, 8(sp)\n13:         mv      t0, a0\n14:         mv      t1, a1\n15:         mul     a0, t0, t1\n16:         mulh    a1, t0, t1\n17:         ld      t0, 8(sp)\n18:         ld      t1, 16(sp)\n19:         ld      ra, 24(sp)\n20:         addi    sp, sp, 32\n21:         ret<\/code><\/pre>\n\n\n\n<p>This code can be added to the <code>product.s<\/code> source file to provide a multiplication operation that uses 64-bit integers. The first thing this function does is save the contents of registers <code>t1<\/code> (line <a href=\"#coderef-save_t1\">11<\/a>) and <code>t0<\/code> (line <a href=\"#coderef-save_t0\">12<\/a>) which will be used by this function.<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"has-small-font-size\"><strong>Note<\/strong>: these are supposed to be caller saved registers, presumably the caller of the product function would have saved them. However, we are saving them here anyway<\/p>\n<\/blockquote>\n\n\n\n<p>The values of the function arguments are then moved into the temporary registers (lines <a href=\"#coderef-mv_arg0\">13<\/a> and <a href=\"#coderef-mv_arg1\">14<\/a>). This is required because, unlike the first version of this function, the arguments need to be reused and the value of <code>a0<\/code> will be overwritten by the first mutiplication on line <a href=\"#coderef-__imul64_low\">15<\/a> which calculates the product of the low 32-bits of the operands. The second multiplication (line <a href=\"https:\/\/www.vociferousvoid.org\/main\/riscv_bare_metal_chapter4#coderef-__imul64_high\">1<\/a><a href=\"#coderef-__imul64_high\">6<\/a>) will calculate the product of the high 32-bits of the operands and store the result in <code>a1<\/code>.<\/p>\n\n\n\n<p>The main program must be updated to handle a potential 128-bit result from the <code>__imul64<\/code> function:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code> 1:         .section \".text.init\"\n 2:         .align 2\n 3:         .global _start\n 4:         .global _stack_end\n 5: _start:\n 6:         lw      a0, operand2\n 7:         lw      a1, operand1\n 8:         la      sp,_stack_end\n 9:         call    sum\n10:         la      t1, result1\n11:         sw      a0, 0(t1)\n12:         call    __imul64\n13:         la      t1, result2\n14:         sd      a0, 8(t1)\n15:         sd      a1, 0(t1)\n16: stop:   j       stop\n17:         .section \".rodata\"\n18: operand1:       .word   4\n19:         .data\n20: operand2:       .word   5\n21:         .bss\n22: result1:        .word   0\n23: result2:        .dword  0, 0<\/code><\/pre>\n\n\n\n<p>The most significant change is that the result must be stored to memory using two instructions: one to store the product of the low 32-bits (line <a href=\"#coderef-save_product_low\">14<\/a>), and one to store the product of the high 32-bits (line <a href=\"#coderef-save_product_high\">15<\/a>). The <code>result2<\/code> variable on line <a href=\"#coderef-product128_result\">23<\/a> must also be updated to reserve 128-bits for the product. The arguments of the <code>__imul64<\/code> function are the same as those of the <code>__imul32<\/code> function. Therefore the new function can be invoked by simply changing the call label on line <a href=\"#coderef-call__imul64\">12<\/a>.<\/p>\n\n\n\n<p>After recompiling and linking the modified source files, the result can be inspecetd in the QEMU console by printing out 2 double-word values at offset <code>0x80001008<\/code>:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>$ qemu-system-riscv64 -M virt -serial \/dev\/null -nographic -kernel main.elf\nQEMU 3.1.0 monitor - type 'help' for more information\n(qemu) xp \/2gd 0x80001008\n0000000080001008:                   45                    0\n(qemu) quit\n<\/code><\/pre>\n\n\n\n<p>Note that the <code>__imul64<\/code> function can also be used with 32-bit operands. The value of the high double-word will be zero in this case since no overflow occurred.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"org9ca4bd4\">Divide<\/h2>\n\n\n\n<p>The <strong>RVM<\/strong> extension also provides instructions to calculate the quotient and remainder a division of an integer by another integer. This is slightly less complicated than multiplication because the result cannot be wider than the operands. However, this also limits divisions to dividends and divisors with a maximum of 64-bits. Therefore this is not a true reciprocal of the multiplication which can have a 128-bit result.<\/p>\n\n\n\n<p>The following listing illustrates the contents of the <code>divide.s<\/code> source file which defines the function to divide an unsigned 64-bit integer divisor by an unsigned 64-bit integer dividend.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code> 1:         .text\n 2:         .align 2\n 3:         .global __idiv64u\n 4: __idiv64u:\n 5:         addi    sp, sp, -32\n 6:         sd      ra, 24(sp)\n 7:         beqz    a1, __idiv64u_exit\n 8:         div     a0, a0, a1\n 9: __idiv64u_exit: \n10:         ld      ra, 24(sp)\n11:         addi    sp, sp, 32\n12:         ret<\/code><\/pre>\n\n\n\n<p>This function is fairly straight forward, after ensuring that the dividend is not zero, it simply calls the <code>div<\/code> instruction to calculate the quotient. The check to ensure that the dividend is not zero on line <a href=\"#coderef-__idiv64u_check_divzero\">7<\/a> is necessary because <strong>R64M<\/strong> does not trap on a divide by zero error. If the dividend is zero, the <code>div<\/code> instruction will be skipped.<\/p>\n\n\n\n<p>Since the result of the <code>__imul64<\/code> function is a 64-bit value due to its small operands, the <code>__idiv64<\/code> function can be invoked on the result to verify its accuracy. The <code>main.s<\/code> program can be updated as follows to divide the result of <code>__imul64<\/code> by <code>operand2<\/code>, and save the result in a variable in the <code>.data<\/code> section named <code>result3<\/code>.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code> 1:         .section \".text.init\"\n 2:         .align 2\n 3:         .global _start\n 4:         .global _stack_end\n 5: _start:\n 6:         lw      a0, operand1\n 7:         lw      a1, operand2\n 8:         la      sp,_stack_end\n 9:         call    sum\n10:         la      t1, result1\n11:         sw      a0, 0(t1)\n12:         call    __imul64\n13:         la      t1, result2\n14:         sd      a0, 0(t1)\n15:         sd      a1, 8(t1)\n16:         bnez    a1, stop\n17:         lw      a1, operand2\n18:         call    __idiv64u\n19:         la      t0, result3\n20:         sd      a0, 0(t0)\n21: stop:   j       stop\n22:         .section \".rodata\"\n23: operand1:       .word   4\n24:         .data\n25: operand2:       .word   5\n26:         .bss\n27: result1:        .word   0\n28: result2:        .dword  0, 0\n29: result3:        .dword  0<\/code><\/pre>\n\n\n\n<p>After <code>__imul64<\/code> returns, the value is checked for overflow (line <a href=\"#coderef-check_overflow\">16<\/a>) by asserting that the value returned in <code>a1<\/code> is zero. This will ensure that the result of the multiplication fits in a single 64-bit register. If the result is greater than 64-bits wide, the division will be skipped. Otherwise <code>operand2<\/code> is loaded into register <code>a1<\/code>. This check is not strictly necessary unless different operand values are used which may result in an overflow.<\/p>\n\n\n\n<p>The divide function will determine the quotient of the <code>__imul64<\/code> result by the value of <code>operand2<\/code>. The quotient will be stored in the <code>result3<\/code> variable. This should be the same as the result of the <code>sum<\/code> function (in <code>result1<\/code>). This can be verified by assembling and linking this program and running the binary in QEMU. The value of <code>result3<\/code><\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>riscv64-unknown-elf-as -o add.o add.s\nriscv64-unknown-elf-as -o divide.o divide.s\nriscv64-unknown-elf-as -o main.o main.s\nriscv64-unknown-elf-as -o product.o product.s\nriscv64-unknown-elf-ld -T chapter3.lds -o main.elf add.o divide.o main.o product.o\nqemu-system-riscv64 -M virt -serial \/dev\/null -nographic -kernel main.elf\nQEMU 3.1.0 monitor - type 'help' for more information\n(qemu) xp \/1wd 0x80001004\n0000000080001004:          9\n(qemu) xp \/1gd 0x80001018\n0000000080001018:                    9\n(qemu) <\/code><\/pre>\n\n\n\n<p>The offset of the <code>result3<\/code> variable will be <code>0x80001018<\/code>; it is 16-bytes beyond the <code>result2<\/code> variable which is located at <code>0x80001008<\/code> (therefore +<code>0x10<\/code>). This can be verified using <kbd>objdump<\/kbd> as in the previous example.<\/p>\n\n\n\n<p>As expected, <code>result3<\/code> contains the integer 9 which is the result of the <code>sum<\/code> function in variable <code>result1<\/code> at offset <code>0x80001004<\/code>.<\/p>\n\n\n\n<p>This value is convenient because 5 divides 45 exactly. If we divided the result of <code>__imul64<\/code> by <code>operand1<\/code> instead, the result would be 11 and there would be a remainder of 1. In the current implementation, this value is lost. However, the divide function can be updated to calculate the quotient and the remainder of a division. The updated <code>__idiv64u<\/code> function is illustrated in the following listing.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code> 1: __idiv64u:\n 2:         # Input:\n 3:         # a0: 64-bit divisor\n 4:         # a1: 64-bit dividend\n 5:         # Returns:\n 6:         # a0 =&gt; 64-bit quotient\n 7:         # a1 =&gt; 64-bit remainder\n 8:         addi    sp, sp, -32\n 9:         sd      ra, 24(sp)\n10:         sd      t1, 16(sp)\n11:         sd      t0, 8(sp)\n12:         beqz    a1, __idiv64u_exit\n13:         mv      t0, a0\n14:         mv      t1, a1\n15:         div     a0, t0, t1\n16:         rem     a1, t0, t1\n17: __idiv64u_exit: \n18:         ld      t0, 8(sp)\n19:         ld      t1, 16(sp)\n20:         ld      ra, 24(sp)\n21:         addi    sp, sp, 32\n22:         ret<\/code><\/pre>\n\n\n\n<p>This new implementation will save the argument values in temporary registers because this is a two-step function and the first argument would be overridden in the first step. The <code>divide<\/code> function then calculates the quotient on line <a href=\"#coderef-calculate_quotient\">15<\/a>, and the remainder on line <a href=\"#coderef-calculate_remainder\">16<\/a>. The <kbd>main.s<\/kbd> program must also be updated to save the result of the new <code>divide<\/code> function in two double words.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code> 1:         .section \".text.init\"\n 2:         .align 2\n 3:         .global _start\n 4:         .global _stack_end\n 5: _start:\n 6:         lw      a0, operand1\n 7:         lw      a1, operand2\n 8:         la      sp,_stack_end\n 9:         call    sum\n10:         la      t1, result1\n11:         sw      a0, 0(t1)\n12:         call    __imul64\n13:         la      t1, result2\n14:         sd      a0, 0(t1)\n15:         sd      a1, 8(t1)\n16:         bnez    a1, stop\n17:         lw      a1, operand1\n18:         beqz    a1, stop\n19:         call    __idiv64u\n20:         la      t0, result3\n21:         sd      a0, 0(t0)\n22:         sd      a1, 8(t0)\n23: stop:   j       stop\n24:         .section \".rodata\"\n25: operand1:       .word   4\n26:         .data\n27: operand2:       .word   5\n28:         .bss\n29: result1:        .word   0\n30: result2:        .dword  0, 0\n31: result3:        .dword  0, 0<\/code><\/pre>\n\n\n\n<p>The only changes are that <code>operand1<\/code> is used as the dividend on line <a href=\"#coderef-load_operand1_dividend\">17<\/a> and an instruction was added on line <a href=\"#coderef-save_remainder\">22<\/a> to store the remainder in RAM. The <code>result3<\/code> variable was also updated to allocate two double-words of memory on line <a href=\"#coderef-result3_2dwords\">31<\/a>. If this program is assembled and linked, then executed in QEMU (as in the previous example), the contents of <code>operand3<\/code> can be inspected to see that both the quotient and remainder have been calculated:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>(qemu) xp \/2gd 0x80001018\n0000000080001018:                   11                    1<\/code><\/pre>\n\n\n\n<p>This provides a more flexible implementation of <code>__idiv64u<\/code>, but if a true reciprocal of the <code>__imul64<\/code> function is desired, the function must allow for a 128-bit divisor argument. The <strong>RV64M<\/strong> extension does not define an instruction to calculate this, therefore the calculation must be performed in parts.<\/p>\n\n\n\n<p>If the 128-bit divisor is broken up into four words, the division can be carried out on each part individually and the result combined. This is possible because of the following:<\/p>\n\n\n\n<p>$$ x=2^{32}w_h+w_l $$<\/p>\n\n\n\n<p>The quotient of <em>x<\/em> by some integer <em>d<\/em> can be calculated as:<\/p>\n\n\n\n<p>$$ \\frac{x}{d}=\\frac{2^{32}w_h}{d} + \\frac{2^{32}w_h \\mod({d+w_l)}}{d} $$<\/p>\n\n\n\n<p>This calculation can be implemented with the following RISC-V assembly code:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code> 1:         .global __idiv128u\n 2: __idiv128u:\n 3:         # Input:\n 4:         # a0: Address where the 128-bit quotient will be stored (high\n 5:         #     dword, low dword).\n 6:         # a1: 64-bit dividend\n 7:         # a2: Address of the 128-bit divisor (high dword, low dword)\n 8:         # Returns:\n 9:         # a0: Address of the 128-bit quotient\n10:         # a1: 64-bit remainder\n11:         addi    sp, sp, -32\n12:         sd      ra, 24(sp)\n13:         # Check for divide by zero\n14:         beqz    a1, __idiv128u_exit\n15:         addi    t2, a2, 16\n16:         li      t3, 0           # t3 = remainder\n17: __idiv128u_next_dword:\n18:         lwu     t1, (a2)        # t1 = low word\n19:         ld      t0, (a2)\n20:         srli    t0, t0, 32      # t0 = high word\n21: __idiv128u_high_word:\n22:         slli    t3, t3, 32      \n23:         add     t0, t0, t3\n24:         divu    t4, t0, a1      # t4 = t0\/a1\n25:         slli    t5, t4, 32      # t5 = t4 * 2^32\n26:         remu    t3, t0, a1      # t3 = t0 mod a1\n27: __idiv128u_low_word:\n28:         slli    t3, t3, 32      # t3 = t3 * 2^32\n29:         add     t0, t1, t3\n30:         divu    t4, t0, a1\n31:         add     t5, t5, t4\n32:         remu    t3, t0, a1\n33:         sd      t5, (a0)\n34:         addi    a2, a2, 8\n35:         addi    a0, a0, 8\n36:         bne     t2, a2, __idiv128u_next_dword\n37:         mv      a0, t3\n38: __idiv128u_exit:        \n39:         ld      ra, 24(sp)\n40:         addi    sp, sp, 32\n41:         ret<\/code><\/pre>\n\n\n\n<p>This function iteratively performs a 64-bit division on 32-bit words of the divisor. The remainder is scaled (<a href=\"#coderef-__idiv128u_scale_remainder\">28<\/a>), then added to the next word of the divisor (line <a href=\"#coderef-__idiv128u_add_remainder\">29<\/a>) and the process is repeated for the next 64-bit double word.<\/p>\n\n\n\n<p>The following listing illustrates an updated <code>main.s<\/code> file:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code> 1:         .section \".text.init\"\n 2:         .align 2\n 3:         .global _start\n 4:         .global _stack_end\n 5: _start:\n 6:         lw      a0, operand1\n 7:         lw      a1, operand2\n 8:         la      sp, _stack_end\n 9:         call    sum\n10:         la      t1, result1\n11:         sw      a0, 0(t1)\n12:         call    __imul64\n13:         la      t1, divisor\n14:         sd      a0, 8(t1)\n15:         sd      a1, 0(t1)\n16:         la      a0, quotient\n17:         lw      a1, operand1\n18:         la      a2, divisor\n19:         call    __idiv128u\n20:         la      t0, remainder\n21:         sd      a0, (t0)\n22: stop:   j       stop\n23:         .section \".rodata\"\n24: operand1:       .word   4\n25:         .data\n26: operand2:       .word   5\n27:         .bss\n28: result1:        .word   0\n29: result2:        .dword  0, 0\n30: result3:        .dword  0, 0\n31: divisor:        .dword  0, 0\n32: quotient:       .dword  0, 0\n33: remainder:      .dword  0<\/code><\/pre>\n\n\n\n<p>This updated main program does not perform an overflow check since the <code>__idiv128u<\/code> function can handle a 128-bit divisor. This function also reads its operands directly from memory rather than from registers due to the fact that the divisor may not fit in a single register. The memory at label <code>quotient<\/code> will be updated with the result of the division. The remainder will be returned by the function, which is then saved to the memory at label <code>remainder<\/code> on line <a href=\"#coderef-save__idiv128u_remainder\">20<\/a>.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"org2adc904\">Atomic Instructions<\/h2>\n\n\n\n<p>Synchronization is an important feature in multiprocessing systems. Thus far, the examples have used a single hardware thread, or hart, therefore there has not been any need to synchronize memory access. RISC-V defines the <strong>A<\/strong> extension which provides instructions to atomically read-modify-write data in memory. These instructions can be used to support synchronization between multiple hardware threads running in the same memory space.<\/p>\n\n\n\n<p>The most basic synchronization primitive is the atomic compare and swap operation. This will compare a value in a register with a value in memory. If the two values are equal, the value in another register will be swapped with the value in memory. The pseudo code for this is as follows:<\/p>\n\n\n\n<ol class=\"wp-block-list\">\n<li>Load value in register <strong>R1<\/strong><\/li>\n\n\n\n<li>Load address of the second value in <strong>R2<\/strong><\/li>\n\n\n\n<li>Load the value at address <strong>R2<\/strong> into a temporary register <strong>T1<\/strong><\/li>\n\n\n\n<li>Load swap value in register <strong>R3<\/strong><\/li>\n\n\n\n<li>If <strong>R1<\/strong> == <strong>T1<\/strong>:\n<ol class=\"wp-block-list\">\n<li>Store <strong>R3<\/strong> at memory location <strong>R2<\/strong><\/li>\n\n\n\n<li><strong>R3<\/strong> := <strong>T1<\/strong><\/li>\n<\/ol>\n<\/li>\n<\/ol>\n\n\n\n<p>This entire sequence is expected to be performed atomically (i.e. there can be no interrupt between the time the value <strong>T1<\/strong> is read from memory, and the end of the procedure. This can be implemented using Load Reserved\/Store Conditional instructions provided by the <strong>RVA<\/strong> extension. The following listing illustrates the implementation of a compare-and-swap function:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code> 1:         .text\n 2:         .align 2\n 3:         .global compare_and_swap\n 4:         # a0: Address of value operand\n 5:         # a1: Value to compare\n 6:         # a2: Value to swap if (a0) == a1\n 7:         # return: a0 == 0 =&gt; CAS successful\n 8:         # return: a0 == 1 =&gt; CAS failed\n 9: compare_and_swap:\n10:         lr.d    t0, (a0)\n11:         bne     t0, a1, nomatch\n12:         sc.d    a0, a2, (a0)\n13:         bnez    a0, compare_and_swap\n14:         j       exit\n15: nomatch:\n16:         li      a0, 1\n17: exit:\n18:         ret<\/code><\/pre>\n\n\n\n<p>This function will atomically compare the value in memory located at the address in <code>a0<\/code> with the value in register <code>a1<\/code>, and store the value of <code>a2<\/code> at the location in <code>a0<\/code> if they match.<\/p>\n\n\n\n<p>The load-reserved instruction on line <a href=\"#coderef-load-reserved\">10<\/a> loads the value at memory location <code>a0<\/code> into register <code>t0<\/code>, and registers a reservation on the address in memory. The nature of the memory reservation is specific to the implementation of the RISC-V core and is transparent to the program. The memory range that is reserved can be arbitrarily sized, however, it must be at least large enough to enclose the value that was loaded.<\/p>\n\n\n\n<p>The value of <code>t0<\/code> is then compared with <code>a1<\/code>. If the values match, the store-conditional instruction on line <a href=\"#coderef-match-store-conditional\">12<\/a> will save the value in <code>a2<\/code> to the memory location of <code>a0<\/code>. This will also release the reservation on the memory address. If the values do not match, the memory is not updated (this instruction is skipped).<\/p>\n\n\n\n<p>If another hardware thread writes data to the memory for which there is a reservation, then the store-conditional instruction will fail and a non-zero error code will be written to the destination register which is <code>a0<\/code> in this function (line <a href=\"#coderef-match-store-conditional\">12<\/a>. In this case, the compare-and-swap operation is restarted (line <a href=\"#coderef-cas-failed\">13<\/a>.<\/p>\n\n\n\n<p>The main program shown in the following listing will invoke the compare-and-swap function:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code> 1:         .section \".text.init\"\n 2:         .align 2\n 3:         .global _start\n 4:         .global _stack_end\n 5: _start:\n 6:         la      sp, _stack_end\n 7:         la      a0, n\n 8:         li      a1, 5\n 9:         li      a2, 6\n10:         call    compare_and_swap\n11:         la      a0, n\n12:         li      a1, 5\n13:         li      a2, 7\n14:         call    compare_and_swap\n15: stop:   j       stop\n16:         .balign 8\n17: n:              .dword  5\n<\/code><\/pre>\n\n\n\n<p>Starting at line <a href=\"https:\/\/www.vociferousvoid.org\/main\/riscv_bare_metal_chapter4#coderef-setup-cas-arguments\">7<\/a>, the function arguments are setup by first loading the address of the variable <code>n<\/code> into <code>a0<\/code>. Note that the alignment of the data loaded by the <code>lr.d<\/code> instruction must be aligned on an 8-byte boundary (similarly the <code>lr.w<\/code> instruction expects the data to be aligned to a 4-byte boundary). The <code>.balign<\/code> (byte align) assembler directive on line <a href=\"#coderef-load-alignment\">16<\/a> ensures that this is the case.<\/p>\n\n\n\n<p>The first invocation of the function on line <a href=\"#coderef-successful-cas\">10<\/a> will succeed, thus the value of <code>n<\/code> will be updated to 6. the second invocation will fail, this the value of <code>n<\/code> will not be changed. This can be verified by assembling the program and inspecting the memory from the QEMU monitor:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>riscv64-unknown-elf-as  -o chapter4_cas_main.o chapter4_cas_main.s\nriscv64-unknown-elf-as  -o cas.o cas.s\nriscv64-unknown-elf-ld -T chapter3.lds -o chapter4-cas.elf chapter4_cas_main.o cas.o\nqemu-system-riscv64 -M virt -serial \/dev\/null -nographic -kernel chapter4-cas.elf\nQEMU 3.1.0 monitor - type 'help' for more information\n(qemu) xp \/1gd 0x80001008\n0000000080001008:                    6\n(qemu) \n<\/code><\/pre>\n\n\n\n<p>In addition to the load-reserved\/store-conditional instructions, the <strong>RVA<\/strong> extension also provides atomic memory operations. These atomically perform an operation on a value in memory, and swap the previous content of the memory location into the targetted register. The supported operations include: <code>add<\/code>, <code>and<\/code>, <code>or<\/code>, <code>xor<\/code>, <code>max<\/code>, <code>min<\/code>, and <code>swap<\/code>. Moreover, the <code>min<\/code> and <code>max<\/code> instructions have signed and unsigned variants. These instructions are convenient for defining another useful synchronization primitive: the test-and-set spinlock.<\/p>\n\n\n\n<p>Spinlocks can be acquired by setting a sentinel value in a specific memory location, but only if that value is not already set therein. If the target memory location already contains the sentinel value, the spinlock will loop until it is released. The lock is released by clearing the memory location (i.e. setting it to zero). The implementation of a spinlock acquire\/release pair is illustrated in the listing that follows:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code> 1:         .text\n 2:         .align 2\n 3:         .global spinlock_acquire\n 4: spinlock_acquire:\n 5:         # a0 = memory address of the spinlock\n 6:         li      t1, 1 #\n 7:         amoswap.d.aq    t0, t1, (a0) #\n 8:         bnez    t1, spinlock_acquire #\n 9:         ret\n10: \n11:         .global spinlock_release\n12: spinlock_release:\n13:         # a0 = memory address of the spinlock\n14:         amoswap.d.rl zero, zero, (a0) #\n15:         ret\n16: \n<\/code><\/pre>\n\n\n\n<p>This listing defines two sub-routines: one to acquire a spinlock, and one to release it. The <code>spinlock_acquire<\/code> function loads the value 1 to use as the sentinel on line <a href=\"#coderef-load_sentinel\">6<\/a>. Then the atomic memory operation <code>amoswap<\/code> is used on line <a href=\"#coderef-set_sentinel\">7<\/a> to swap the value of the sentinel with the contents of the memory location specified in <code>a0<\/code>. The value contained in the lock location will be saved in register <code>t0<\/code>. If this value is not zero, the lock was already acquired by another thread, therefore the function will try again (line <a href=\"#coderef-test_if_locked\">8<\/a>), otherwise the function returns.<\/p>\n\n\n\n<p>the <code>spinlock<\/code> release function will simply write zero into the memory location specified in <code>a0<\/code>. This will allow another thread that is spinning on the lock to acquire it.<\/p>\n\n\n\n<p>The <code>amoswap<\/code> instruction has two variants: one for double-words (<code>amoswap.d<\/code>) and one for word values (<code>amoswap.w<\/code>). Moreover, there are flags which define define the release consistency semantics of the memory operation (the <code>.aq<\/code> and <code>.rl<\/code> suffixes). Basically by setting the <code>.aq<\/code> suffix on the operation, then the effect of memory operations that occur after this one in the current hardware thread will not be observed by another thread before the effect of the current instruction. Conversely, when the <code>.rl<\/code> suffix is specified, the effects of memory operations preceding that of the current instruction will not be observed by other threads after its own effect.<\/p>\n\n\n\n<p>The following program illustrates the use of the spinlock functions to define a critical section:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code> 1:         .section \".text.init\"\n 2:         .align 2\n 3:         .global _start\n 4:         .global _stack_end\n 5: _start:\n 6:         la      sp, _stack_end\n 7:         la      a0, lock #\n 8:         call    spinlock_acquire #\n 9:         la      t0, n #\n10:         ld      a0, (t0)\n11:         li      a1, 1\n12:         call    sum\n13:         la      t0, n\n14:         sd      a0, (t0) #\n15:         la      a0, lock\n16:         call    spinlock_release #\n17: stop:   j       stop\n18:         .data\n19:         .balign 8\n20: lock:   .dword  0\n21: n:      .dword 0\n22: \n<\/code><\/pre>\n\n\n\n<p>This program will attempt to acquire the spinlock on line <a href=\"#coderef-spinlock_acquire\">8<\/a> (the address of the lock variable is loaded on line <a href=\"#coderef-load_lockaddr\">7<\/a>). This function call will block until the lock is acquired. Since there is only a single hardware thread, the lock should be acquired immediately. The critical section starts on line <a href=\"#coderef-critical-section-start\">9<\/a>. The variable <code>n<\/code> is loaded and incremented by calling the <code>sum<\/code> function (defined in a previous chapter). The critical section ends on line <a href=\"#coderef-critical-section-end\">14<\/a>, at which point the program releases the spinlock (<a href=\"#coderef-spinlock_release\">16<\/a>. Following the execution of this program, the contents of the variable <code>n<\/code> should be 1:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>riscv64-unknown-elf-as  -o chapter4_spinlock_main.o chapter4_spinlock_main.s\nriscv64-unknown-elf-as  -o spinlock.o spinlock.s\nriscv64-unknown-elf-as  -o add.o add.s\nriscv64-unknown-elf-ld -T chapter3.lds -o chapter4-spinlock.elf chapter4_spinlock_main.o spinlock.o add.o\nqemu-system-riscv64 -M virt -serial \/dev\/null -nographic -kernel chapter4-spinlock.elf\nQEMU 3.1.0 monitor - type 'help' for more information\n(qemu) xp \/1gd 0x80001008\n0000000080001008:                    1\n\n<\/code><\/pre>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"org86a22e8\">Floating Point<\/h2>\n\n\n\n<p>In <a href=\"https:\/\/www.vociferousvoid.org\/index.php\/2019\/11\/06\/risc-v-bare-metal-programming-chapter-2-opcodes-assemble\/\">chapter 2<\/a>, the base set of the base <strong>I<\/strong> (integer) registers were enumerated. However, when inspecting the VirtIO machine in QEMU, using the <code>info registers<\/code> command, certain registers were listed that are not described in the table. These registers exist to support the <strong>F<\/strong> or <strong>D<\/strong> extensions which provide floating point arithmetic instructions that work with operands which conform to the IEEE 754-2008 standard. The <strong>F<\/strong> extension provides support for single-precission values and operands, and the <strong>D<\/strong> extension provides the same instructions for double-precision values.<\/p>\n\n\n\n<p>The 32 additional registers, <code>f0<\/code>&#8211;<code>f31<\/code>, are used exclusively by the instructions provided by the <strong>RVF<\/strong> and <strong>RVD<\/strong> extensions. This doubles the number of registers available to the processor without increasing the space required for the register specifier in the instruction op-code since only enough bits to enumerate 32 registers are required (5 bits).<\/p>\n\n\n\n<p>If only the <strong>RVF<\/strong> extension is supported, the <code>f<\/code> registers will be 32-bits wide. If the <strong>RVD<\/strong> extension is supported, the <code>f<\/code> registers will be 64-bits wide. If both <strong>RVF<\/strong> and <strong>RVD<\/strong> are supported, the <strong>RVF<\/strong> instructions will use only the lower 32-bits of the 64-bit registers.<\/p>\n\n\n\n<p>The <code>f<\/code> registers are enumerated in the following table with their ABI name and a description:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table><thead><tr><th>Register(s)<\/th><th>ABI Name(s)<\/th><th>Description<\/th><\/tr><\/thead><tbody><tr><td>f0-f7<\/td><td>ft0-ft7<\/td><td>Temporary<\/td><\/tr><tr><td>f8-f9<\/td><td>fs0-fs1<\/td><td>Saved register<\/td><\/tr><tr><td>f10-f11<\/td><td>fa0-fa1<\/td><td>Function argument\/Return value<\/td><\/tr><tr><td>f12-f17<\/td><td>fa2-fa7<\/td><td>Function argument<\/td><\/tr><tr><td>f18-f27<\/td><td>fs2-fs11<\/td><td>Saved register<\/td><\/tr><tr><td>f28-f31<\/td><td>ft8-ft9<\/td><td>Temporary<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p>These registers roughly mirror the base integer registers with two notable exception: unlike <code>x0<\/code>, <code>f0<\/code> is not hardwired to <code>0<\/code>, it can be used just like every other register. Moreover there are no registers to manage return addresses, stacks, globals, or threads. The equivalent <code>f<\/code> registers are used as temporaries.<\/p>\n\n\n\n<p>The convention for who is responsible for saving the contents of the registers is essentially the same as the equivalent base integer registers: Saved registers and temporary registers are to be saved by the callee. All other registers must be saved by the caller.<\/p>\n\n\n\n<p>In addition to the 32 <code>f<\/code> registers, the <strong>RVF<\/strong> and <strong>RVD<\/strong> extensions define a status and control register: <code>fcsr<\/code>. The <strong>RVF<\/strong> and <strong>RVD<\/strong> extensions provide the <code>frcsr<\/code> instruction to read this register, storing its value into the targetted integer register. Similarly, the <code>fscsr<\/code> instruction will copy the original value of <code>fcsr<\/code> into the destination integer register, and the write the value in the source integer register thereto.<\/p>\n\n\n\n<p>The <code>fcsr<\/code> prescribes the rounding mode used by floating point operations. The rounding mode field occupies bits 5-7 of the register. The <strong>RVF<\/strong> and <strong>RVD<\/strong> extensions also define the <code>frrm<\/code> instruction to retrieve the rounding mode.<\/p>\n\n\n\n<p>The <code>fcsr<\/code> register also contains flags to indicate exception conditions that may have occured while executing floating-point arithmetic since it was last reset. These errors include:NV Invalid operation (<code>fcsr<\/code>[4]) <code>DZ<\/code> Divide by zero (<code>fcsr<\/code>[3]) <code>OF<\/code> Overflow (<code>fcsr<\/code>[2]) <code>UF<\/code> Underflow (<code>fcsr<\/code>[1]) <code>NX<\/code> Inexact (<code>fcsr<\/code>[0])<\/p>\n\n\n\n<p>The floating-point exception flags can also be retrieved using the <code>frflags<\/code> instruction which saves their state in the specified integer registers.<\/p>\n\n\n\n<p>The <strong>RVF<\/strong> and <strong>RVD<\/strong> extensions define two load instructions and two store instructions. These are essentially mirrors of the base load and store instructions that use the <code>f<\/code> registers rather than the <code>x<\/code> integer registers. Therefore their addressing mode and format are the same as the <code>lw<\/code>, <code>ld<\/code>, <code>sw<\/code> and <code>sd<\/code> instructions.<\/p>\n\n\n\n<p>The <strong>RVF<\/strong> and <strong>RVD<\/strong> extensions also provide a set of arithmetic instructions including:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><code>fadd<\/code><\/li>\n\n\n\n<li><code>fsub<\/code><\/li>\n\n\n\n<li><code>fmul<\/code><\/li>\n\n\n\n<li><code>fdiv<\/code><\/li>\n\n\n\n<li><code>fsqrt<\/code><\/li>\n<\/ul>\n\n\n\n<p>Each instruction has a single- and double-precision variant which can be specified by adding a <code>.s<\/code> or <code>.d<\/code> suffix to the instruction respectively.<\/p>\n\n\n\n<p>The floating-point arithmetic instructions will operate using only the <code>f<\/code> registers, therefore the extensions provide instructions to move data from integer to floating point registers.<\/p>\n\n\n\n<p>The following function implementation will demonstrate some of these instructions. The function in <code>fvector.s<\/code> will multiply each element from an array of floating-point values by a floating-point scalar:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code> 1:         .text\n 2:         .align 2\n 3:         .global __vec_scalef\n 4:         # a0: number of elements, 'n', in the array\n 5:         # fa0: A double-precision floating-point scalar 'a'\n 6:         # a1: Address of array of x&#91;n] double-precision floating-point values.\n 7: __vec_scalef:\n 8:         addi    sp, sp, -32\n 9:         sd      ra, 24(sp)\n10:         beqz    a0, __vec_scalef_exit\n11: __vec_scalef_loop:\n12:         fld     fa5,0(a1)\n13:         fmul.d  fa5, fa5, fa0\n14:         fsd     fa5,0(a1)\n15:         addi    a1, a1,8\n16:         addi    a0, a0,-1\n17:         bnez    a0, __vec_scalef_loop\n18: __vec_scalef_exit:\n19:         ld      ra, 24(sp)\n20:         addi    sp, sp, 32\n21:         ret\n<\/code><\/pre>\n\n\n\n<p>In this function each value of the double array is loaded on line <a href=\"#coderef-load_double\">12<\/a> at each iteration (up to a maximum set by the integer value in <code>a0<\/code>). The loaded value is multipled by the double-precision floating-point value in <code>fa0<\/code> on line <a href=\"#coderef-mul_double\">13<\/a>, then stored to the same memory location on line <a href=\"#coderef-store_double\">14<\/a>.<\/p>\n\n\n\n<p>The source data for the function can be defined using the <code>.double<\/code> assembler directive. This directive will store double-precision floating-point values in successive memory double-words. The <code>.float<\/code> directive will do the same for single-precision floating-point values.<\/p>\n\n\n\n<p>There are many more instructions defined in the <strong>RVF<\/strong> and <strong>RVD<\/strong> extensions. Enough to dedicate an entire chapter to this topic. Moreover, the QEMU support for the <strong>RVF<\/strong> and <strong>RVD<\/strong> does not seem to be fully implemented for the version available in the Debian 10 packages. A more thorough investigation of these extensions will be reserved for a future chapter.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"orgfdc44a9\">Conclusion<\/h2>\n\n\n\n<p>The RISC-V architecture is designed to be a simple as possible but no simpler. Therefore a building block philosophy is followed to allow chip designers to include as many or as few instructions as needed. This provides some flexibility to system designers to satisfy cost, efficiency, and performance constraints specific to the application domain.<\/p>\n\n\n\n<p>Breaking out instructions into optional extensions is like having Lego bricks representing sub-sets of the total RISC-V ISA. In this chapter the <strong>M<\/strong>, <strong>A<\/strong>, <strong>F<\/strong>, and <strong>D<\/strong> extensions were used to create a small library of functions that can be re-used in the future to perform more complex calculations, and to synchronize memory access across hardware threads.<\/p>\n\n\n\n<p>In addition to these there are two other optional standard extensions that were not covered in this chapter:C Compressed instructions. <strong>V<\/strong> Vector instructions for SIMD processing.<\/p>\n\n\n\n<p>Discussion of these extensions will be reserved for future chapter.<\/p>\n\n\n\n<p>In the next chapter the privileged instruction set will be described. This allows for varying levels of support for the base instructions. In this chapter, the utility functions defined so far will be used to create more complex programs. The synchronization utilities will be particularly useful when dealing with interrupts.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Chapter 3 of this RISC-V bare metal tutorial studied the linking process and how a developer can control where code and data are placed in memory. Constants, initialized variables and uninitialized variables were defined and explicitly positioned in RAM as prescribed by a linker script. The running example program was updated to read operands from [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_jetpack_memberships_contains_paid_content":false,"footnotes":""},"categories":[4],"tags":[],"class_list":["post-33","post","type-post","status-publish","format-standard","hentry","category-risc-v"],"jetpack_featured_media_url":"","jetpack_sharing_enabled":true,"_links":{"self":[{"href":"https:\/\/www.vociferousvoid.org\/index.php\/wp-json\/wp\/v2\/posts\/33","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.vociferousvoid.org\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.vociferousvoid.org\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.vociferousvoid.org\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.vociferousvoid.org\/index.php\/wp-json\/wp\/v2\/comments?post=33"}],"version-history":[{"count":6,"href":"https:\/\/www.vociferousvoid.org\/index.php\/wp-json\/wp\/v2\/posts\/33\/revisions"}],"predecessor-version":[{"id":72,"href":"https:\/\/www.vociferousvoid.org\/index.php\/wp-json\/wp\/v2\/posts\/33\/revisions\/72"}],"wp:attachment":[{"href":"https:\/\/www.vociferousvoid.org\/index.php\/wp-json\/wp\/v2\/media?parent=33"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.vociferousvoid.org\/index.php\/wp-json\/wp\/v2\/categories?post=33"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.vociferousvoid.org\/index.php\/wp-json\/wp\/v2\/tags?post=33"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}