Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Skyarch Instruction Set

Gen Design

Word Size: 32 bit

Memory Load/Store Order: Little Endian

Instruction Size/Alignment: 4 bytes

Variable Sized Instructions: No

Common Flags/Condition Code Driven

Instruction format

8-bit opcode in lowest byte, 24-bit payload in upper three bytes.

Bit order is notated as MSB-first, e.g. “31 down to 0”. In memory, data and instructions should be stored in little endian, so the instruction will be in the first/lowest address, followed by the three payload bytes.

Registers

Register Maps

There are 8 maps of registers:

  • Map 0: General Purpose
  • Map 1: System Configuration
  • Map 2: I/O Transfer Registers
  • Map 3: Information
  • Map 4: Coprocessor Control
  • Maps 8-15: Co-processor Registers.

There are 32 registers of each type. Except for Map 0, not all registers may be defined.

Map 0: General Purpose

Assembly syntax: rn.

All Registers of the Map are defined. Certain Registers have special meaning:

  • r0 is the zero-register. It reads as zero, and writes are ignored.

Map 1: Interrupt Support

Assembly syntax: intn or alias.

Refer to the following table of defined registers. Some registers define a specific format

Regno.AliasesDescription
0intctlInterrupt Status Register
1intret1Return for Priority 1 Interrupts
2intret2Return for Priority 2 Interrupts
3intret3Return for Priority 3 Interrupts
4ints0Scratch Register
5ints1Scratch Register
6ints2Scratch Register
7ints3Scratch Register
8intd0Misc Config Register
9intd1Misc Config Register
10intd2Misc Config Register
11intd3Misc Config Register
31inttabInterrupt Table Pointer

Reading or Writing an undefined register causes EX[2]. Writing an invalid value to a defined register causes EX[4]

Interrupt Control (Map 1, Register 0)

+31-----------------------------0+
|r00000000000000000000000000000mm|
+--------------------------------+

(All bits indicated as 0 must be written with 0)

BitsNameDescription
rAbort TriggeredSet to 1 when an Abort (Ex[0]) occurs.
mPriority MaskInterrupts with priority value > m are blocked

Both fields are set to 0 on startup.

Interrupt Priority

Interrupt Priority is used to ensure that overlapping Interrupts do not interfere. There are 4 Priority levels, numbered in descending order of priority (0 is the highest priority, 3 is the lowest priority)

  • Priority 0: Abort (Ex[0])
  • Priority 1: Synchronous Exceptions (Ex[1], Ex[2], Ex[3], Ex[4])
  • Priority 2: Asynchronous High Priority Event (Ex[7], EX[8-15])
  • Priority 3: IRQs

An interrupt/trap is blocked when the priority level is greater than m. The behaviour depends on the kind of exception:

  • Synchronous Exceptions (other than Abort) Reset the processor if r = 1, else they set r = 1 and raise Ex[0]
  • Asynchronous Events are discarded
  • IRQs are buffered (up to an implementation-specific capacity until an intret occurs that sets m to be 3) or are discarded.

Interrupt Return Registers

Each priority of interrupt (other than priority 0) has a distinct return register, labeled intretn where n is the priority value, which corresponds to Register n in map 1. Aborts are not recoverable, so no return register is provided.

+31-----------------------------0+
|aaaaaaaaaaaaaaaaaaaaaaaaaaaaaamm|
+--------------------------------+
BitsNameDescription
aAddressContains the high 30 bits of the return address
mPriority MaskStores the priority mask before the interrupt

Interrupt Scratch/Config Registers

Registers 4 through 11 in Map 1 are unused, freely writable registers, labeled intsn for registers 4+n and intdn for registers 8+n. The register intsn is intended for use as a scratch register for interrupts with priority n (used during the interrupt procedure) and intdn i intended for use as a data/configuration register for such interrupts (written by the program and read during each interrupt invocation).

Interrupt/Exception Table (Map 1, Register 31)

+31-----------------------------0+
|aaaaaaaaaaaaaaaaaaaaaaaaaaaaa000|
+--------------------------------+

(All bits indicated as 0 must be written with 0)

Bits a contain the 29 most significant bits of an 8-byte aligned address which points to the interrupt table. 512 bytes starting from this address refer to 64 8-byte entries of the interrupt table, which use the following format, in LSB-first order using little-endian byte encoding:

+31-----------------------------0+
|tttttttttttttttttttttttttttttt0p|
+63----------------------------32+
|00000000000000000000000000000000|
+--------------------------------+

The t bits are the 30 most significant bits of the address to transfer control to when the specified interrupt occurs.

The p bit must be set for all interrupt vectors that are present and valid to execute. If the CPU tries to execute a not-present interrupt vector, EX[4] is raised.

All bits indicated as 0 must be 0 when the vector is read from memory, or EX[4] is raised.

Interrupts

The first 32 interrupt entries are reserved for hardware exceptions, these interrupts are allocated as follows (and the nth entry in this list is designated elsewise as EX[n]):

  • Entry 0: Exception Handling Fault - an exception is raised when the t flag is set.
  • Entry 1: Bus Fault - accessing memory in a particular manner causes an error, or attempts to access memory that doesn’t exist.
  • Entry 2: Invalid Instruction - An instruction that is executed is an unknown opcode, reserved, malformed, or invalid
  • Entry 3: Unaligned Branch Target - an indirect branch is unaligned.
  • Entry 4: Consistency - An invalid system control structure was loaded from memory, or an invalid value was written to a system register.
  • Entry 5: Debug Trap - Allows software-level debugging via the BREAKP instruction.
  • Entry 7: PIRQ - May be raised in response to a priority signal external to the processor that requires immediate resolution. This is handled like an IRQ, but uses priority 2 instead of priority 3.
  • Entries 8-15: Co-processor Unit n Error - The corresponding Coprocessor unit n signals an error after a CPIn instruction (n is Exception number - 4).
  • Entries6 and 16-31 are reserved.

The remaining entries (32-63), may be allocated as IRQ vectors.

Interrupt Checking

Interrupts are performed as follows:

subrountine TryInterruptProcessor(iv: u8, pri: u2):
    let intctl: u32 = ReadRegister(1, 0);
    if (intctl & 3) < pri:
        return;
    let addr: u32 = ReadRegister(1, 31) + (iv << 3);
    let retreg = IP | intctl & 3;
    if pri > 0:
        WriteRegister(1, pri, retreg);
    let iaddr = ReadMemory(addr);
    let rest = ReadMemory(addr + 4);
    CheckAndRaise(EX[2]);
    if (iaddr & 1) == 0:
        return;
    if rest != 0 or (iaddr & 2) != 0:
        Raise(EX[4]);
    let addr = iaddr & ~3;
    IP = addr;
    IL.valid = false;
    return;

subroutine InterruptProcessor(iv: u6, pri: u2):
    let intctl: u32 = ReadRegister(1, 0);
    if (intctl & 3) < pri:
        if pri == 1:
            if (intctl & 0x80000000) != 0:
                ResetProcessor();
            else:
                WriteRegister(1, 0, 0x80000000);
                InterruptProcessor(0, 0);
                return;
        else:
            return;

    let addr: u32 = ReadRegister(1, 31) + (iv << 3);
    let retreg = IP | intctl & 3;
    if pri > 0:
        WriteRegister(1, pri, retreg);
    let iaddr = ReadMemory(addr);
    let rest = ReadMemory(addr + 4);
    CheckAndRaise(EX[2]);
    if rest != 0 or (iaddr & 2) != 0:
        Raise(EX[4]);
    if (iaddr & 1) == 0:
        Raise(EX[4]);
    let addr = iaddr & ~3;
    IP = addr;
    IL.valid = false;
    return;

subroutine Raise(EX[n]: Except):
    CancelCurrentInstruction();
    InterruptProcessor(n, 1);
    return;

subroutine CheckAndRaise(EX[n]: Except):
    if AsynchronousExceptionPending(EX[n]):
        Raise(EX[n]);
        ClearPendingException(EX[n]);
    return;

subroutine CheckAsync():
    if AsynchronousExceptionPending(EX[7]):
        InterruptProcessor(7, 2);
        ClearPendingException(EX[7]);
    let cpe = ReadRegister(4, 30);
    for n in 0..8:
        if (cpe & (1 << n)) != 0 and AsynchronousExceptionPending(EX[8+n]):
            InterruptProcessor(8+n, 2);
        ClearPendingException(EX[8+n]);

    let irq, hasirq = PullPendingIrq();
    if hasirq:
        InterruptProcessor(32+irq, 3);

The processor behaves as if CheckAsync() is called after each instruction finishes writing to all memory and all registers.

Map 2: I/O Transfer Registers

Assembly Syntax: ion.

Map 2 defines a sequence of input and output shift registers for transfering data to external peripherals.

All Registers are defined and have no implied meaning.

Map 3: Information Registers

The Information Registers Map is a Read Only Map that contains information about the CPU. All Registers Presently Read 0. Writes are illegal and raise EX[2]

Map 4: Coprocessor Control

Each Co-processor has a 32-bit control word, which is defined by the Coprocessor.

Assembly Syntax: cpn or alias.

Register N (N < 8) in Map 4 is defined if Co-processor N is present and enabled. Additionally, Register 30 is the coprocessor enable (cpe) register, and Register 31 is the coprocessor present (cpp) register.

Reads and writes to an undefined register or a register corresponding to a not-present or disabled coprocessor results in EX[2]. Writing to cpp results in EX[2].

Where a MOV instruction writes to Register N (N < 8) in this map, the following guarantees are made about the ordering of surrounding instructions:

  • The MOV instruction will not begin executing until any Coprocessor Invocation instruction that references coprocessor N and occurs before it have been fully written-back,
  • Any Coprocessor Invocation instruction that references coprocessor N and occurs after it will not begin executing until the MOV instruction has been fully written-back

A coprocessor may enforce arbitrary validity requirements on writes to its corresponding control register. Violations of these constraints generates a coprocessor error.

Map 4, Register 30: Coprocessor Enable

The Coprocessor Enable register allows the system software to control what coprocessors are operating and usable from the CPU.

+31-----------------------------0+
|000000000000000000000000EEEEEEEE|
+--------------------------------+

The bits marked E may be set by the program when the corresponding bit of Register 31 is set. Setting the nth bit to 1 enables the coprocessor and setting it to 0 disables it.

A bit may only be written with 1 if the corresponding bit in cpp is 1. If a write violates this rule, EX[4] is raised. Bits marked as 0 can never be written with a 1.

Where a MOV instruction writes to this register the following guarantees are made about the ordering of surrounding instructions:

  • The MOV instruction will not begin executing until all off the following instructions that occur before it have been fully written-back:
    • Any MOV instruction that references a register in this map, other than an access to this register, cpp, or any undefined register,
    • Any MOV instruction that reads from a register in Maps 8-16 that does not refer to a not-present,
    • Any coprocessor invocation instruction that does not refer to a not-present coprocessor,
  • Any of the following instructions that occur after the MOV instruction will not begin executing until the MOV instruction has been fully written-back:
    • Any MOV instruction that references a register in this map, other than this register, cpp, or any undefined register,
    • Any MOV instruction that accesses a register in Maps 8-16 that does not refer to a not-present coprocessor,
    • Any coprocessor invocation instruction that does not refer to a not-present coprocessor.

Map 4, Register 31: Coprocessor Present

The Coprocessor Enable register allows the system software to determine what coprocessors are connected to the CPU. This register is read-only and cannot be written from the CPU.

+31-----------------------------0+
|000000000000000000000000PPPPPPPP|
+--------------------------------+

The nth bit is set to 1 if the nth coprocessor is present. Note that it is not guaranteed that the set of enabled coprocessors is contiguous or that the set of enabled coprocessors begins at 0.

Map 8-15: Co-processor Maps

Co-processors connected to the system may expose up to 32 registers each. Registers in map N are only defined if the coprocessor co-processor (Co-processor N-8) is enabled.

Writing to a register in map 8+N with cpe[bit N] clear raises Ex[2]. Violating a validity constraint enforced by the coprocessor may raise an asynchronous coprocessor error.

Reset State

On Reset (either hardware initiated, or initiated by an exception raised in an abort status), the CPU is initialized to the following state:

  • It is executing (Status = 0)
  • IP is initialized to 0xFF00.
  • cpe is set to 0.
  • ictl is set to m=0, a=0
  • r0 is 0.

All other registers, including flags, have undefined values. An undefined value means that any independent read from the register may return an unpredictable result. Undefined values are cleared when the register is written to, even if the register is assigned to itself.

Instructions

Undefined Instructions

MnemonicOpcodePayload
7------031---------------------8
UND00000000-
UND11111111-

(The Payload bits are ignored by both instructions)

  • EX[2]: Unconditionally

Unconditionally raises Invalid Instruction errors

instruction UND():
    Raise(EX[2])

Pause

MnemonicOpcodePayload
7------031---------------------8
PAUSE00000001000000000000000000kkkkkk
  • k: Total pause time
  • Ex[2]: If any reserved (fixed) bit is set to an invalid value.

Delays execution for k clock cycles, 0-63,

Any instruction that follows a PAUSE instruction will not begin executing until the k cycles have elapsed since the PAUSE instruction completed execution.

instruction PAUSE(k: u6):
    SuspendForClockTicks(k)

Move

MnemonicOpcodePayload
7------031---------------------8
MOV0000001000mmmmrl0sssss0ccccddddd
  • m: Map
  • r: Direction
  • l: Latency Control
  • s: Source Register
  • c: Condition Code (See Jump)
  • d: Destination Register
  • Ex[2]: If any reserved (fixed) bit is set to an invalid value
  • Ex[2]: If an undefined register map is referenced by the instruction
  • Ex[2]: If an undefined register in a map other than map 0 is referenced by the instruction
  • Ex[2]: If a read-only register in a map other than map 0 is referenced by the instruction
  • Ex[5]: If an invalid value is written to a register outside of map 0

Copies data between general purpose registers and to/from general purpose registers into other registers.

instruction MOV(d: u5, s: u5, m: u2, dir: u1, c: ConditionCode, l: bool):
    if m!=0:
        if dir==0:
            ValidateRegisterReadable(m,d);
        else:
            ValidateRegisterWritable(m,d);
    if CheckCondition(flags, c):
        let ms, md: u2;
        if dir==1:
            md = m;
            ms = 0;
        else:
            ms = m;
            md = 0;
        if md==3:
            Raise(EX[2]);
        let val: u32;
        val = ReadRegister(ms, s);
        if md == 1 or m > 3:
            ValidateConfigurationRegisterValue(d, val);
        WriteRegister(md, d, val);

LD/ST

MnemonicOpcodePayload
7------031---------------------8
ST00000011rrmm00000000wwsssssddddd
LD00000100rrmm00000000wwsssssddddd
LDI00000101iiiiiiiiiiiiiiii00xddddd
LRA00000110oooooooooooooooo00xddddd
  • r: Ordering
  • m: Update mode
  • w: Width
  • s: Source Register
  • x: Set sign bits (Higher half)
  • i: Immediate Value
  • o: Offset
  • d: Destination Register
  • Ex[2]: If any reserved (fixed) bit is set to an invalid value
  • ST: Ex[2]: If d is 0.
  • LD: Ex[2]: If s is 0.
  • ST, LD: Ex[2]: If w=3.
  • ST: Ex[2]: If r = 1
  • LD: Ex[2]: If r = 2
  • ST: Ex[1]: If d is not aligned to 2 ^ w bytes
  • LD: Ex[1]: If d is not aligned to 2 ^ w bytes.
  • ST, LD: Ex[1]: If accessing memory causes a bus error
  • ST: Stores 1 << w bytes from d to [s]
  • LD: Loads 1 << w bytes from [s] into d
  • LDI: Loads an immediate i (sign or zero extened) into the first (h=0) 16 bits of d
  • LRA: Loads the address IP + o (o is a signed immediate if x is true, and an unsigned immediate otherwise) into d. IP is taken from the beginning of the next instruction

Every ST, STIC, or STICW instruction that modifies a given memory region will become visible on every core of the system in the same order.


enum Ordering:
    Relaxed = 0,
    Acquire = 1,
    Release = 2,
    SeqCst = 3,

enum UpdateMode:
    None = 0,
    PostInc = 1,
    /* Illegal = 2 */,
    PreDec = 3,

instruction ST(s: u5, d: u5, w: u2 r: Ordering, m: UpdateMode):
    if r==1:
        Raise(EX[2])
    if m == 2:
        Raise(EX[2])
    if d==0:
       Raise(EX[2]);
    let val = ReadRegister(0,s);
    let addr: u32;
    let width = 2 << w;
    if (m&2)== 2:
        addr = ReadRegister(0, d) - width;
    else:
        addr = ReadRegister(0, d);
    if width == 8:
        Raise(EX[2]);
    if addr & (width - 1):
        Raise(EX[1])

    SynchronizeMemoryAccordingToStore(r, addr);
    WriteAlignedMemoryTruncate(addr, val, width);
    CheckAndRaisePending(EX[1]);
    let new_addr: u32;
    if m == 1:
        new_addr = addr + width;
    if m != 0:
        WriteRegister(0, d, new_addr);


instruction LD(s: u5, d: u5,w: u2, p: u2):
    if r==2:
        Raise(EX[2])
    if s==0:
        Raise(EX[2])
    let width = 2 << w;
    if width == 8:
            Raise(EX[2]);
    let addr: u32;
    if (m&2)== 2:
        addr = ReadRegister(0, s) - width;
    else:
        addr = ReadRegister(0, s);
    if addr & (width - 1):
        Raise(Ex[1])
    let val = ReadAlignedMemoryZeroExtend(addr, w+1);
    CheckAndRaisePending(EX[1]);
    SynchronizeMemoryAccordingToLoad(r, addr);
    WriteRegister(0,d,val);
    let new_addr: u32;
    if m == 1:
        new_addr = addr + width;
    if m != 0:
        WriteRegister(0, s, new_addr);

instruction LRA(d: u5, x: bool, i: u16):
    let val = SignExtendOrZeroExtend(i, x) + IP;
    WriteRegister(0,d,val);

Immediate Arithmetic

MnemonicOpcodePayload
7------031---------------------8
ADDI00001000iiiiiiiiiiiiiiiihfxddddd
  • i: Immediate
  • h: High half
  • f: Enable Flags Modification
  • x: Set sign bits (upper bits)
  • d: Destination Register
  • Ex[2]: If h and x are both set.

Sets P, N, and Z according to the result. Sets V and C according to the computation (signed overflow and carry)

Adds a 16-bit zero or sign-extended immediate to d.

instruction ADDI(d: u5, x: bool, f: bool, h: bool, i: u16):
    let imm: u32;

    if h and x:
        Raise(Ex[2]);

    if h:
        imm = ZeroExtend(i, 32) << 16;
    else if x:
        imm = SignExtend(i, 32);
    else:
        imm = ZeroExtend(i, 32);
    let r = ReadRegister(0,d);

    let result, flags_val = r + imm;

    WriteRegister(0, d, result);
    if f:
        flags = flags_val;

ALU Instructions

MnemonicOpcodePayload
7------031---------------------8
ADD00001001c0psssssfbbbbbaaaaaddddd
SUB00001010c0psssssfbbbbbaaaaaddddd
AND00001011jipsssssfbbbbbaaaaaddddd
OR00001100jipsssssfbbbbbaaaaaddddd
XOR00001101jipsssssfbbbbbaaaaaddddd
  • c: Carry in
  • j: Invert op 2
  • i: Invert op 1
  • p: Shift Polarity
  • s: Shift Quantity
  • f: Enable Flags Modification
  • b: Source Register 2
  • a: Source Register 1
  • d: Destination Register
  • Ex[2]: If any reserved (fixed) bit is set to an invalid value.
  • ADD/SUB: Sets P, N, and Z according to the result. Sets V and C according to the computation (signed overflow and carry)
  • AND/OR/XOR: Sets P, N, and Z according to the result. V and C are set to unspecified values.

Computes the ALU corresponding ALU operation between the values in GPRs b and a, writing the result to GPR d. The operand corresponding to p is first shifted left by s (p=1 shifts a, p=0 shifts b)

  • ADD: The result is a + b. If c is set, the carry flag is also added in.
  • SUB: The result is a - b. If c is set, the carry flag is borrowed by the subtraction.
  • AND: The result is (i)a & (j)b, a is inverted before being shifted if i is set, and b is inverted before being shifted if j is set.
  • OR: The result is (i)a | (j)b, a is inverted before being shifted if i is set, and b is inverted before being shifted if j is set.
  • XOR: The result is (i)a ^ (j)b, a is inverted before being shifted if i is set, and b is inverted before being shifted if j is set.
instruction {ADD, SUB}(a: u5, b: u5, d: u5, f: bool, s: u5, p: bool, c: bool):
    let src1, src2: u32;
    if p:
        src1 = ReadRegister(0, a) << s;
        src2 = ReadRegister(0,b);
    else:
        src1 = ReadRegister(0, a);
        src2 = ReadRegister(0,b) << s;
    let dest: u32;
    let flags_val, flags_mask: u4;
    switch (instruction):
        case ADD:
            dest, flags_val = src1 + src2 + (flags.c & c);
            flags_mask = 0xF;
        case SUB:
            dest, flags_val = src1 - src2 + (~flags.c & c);
            flags_mask = 0xF;
    if f:
        flags = flags_val & flags_mask | nondeterministic() & ~flags_mask;

instruction {AND, OR, XOR}(a: u5, b: u5, d: u5, f: bool, s: u5, p: bool, i: bool, j: bool):
    let src1, src2: u32;
    if p:
        src1 = ReadRegister(0, a) << s;
        src2 = ReadRegister(0,b);
    else:
        src1 = ReadRegister(0, a);
        src2 = ReadRegister(0,b) << s;

    let val1, val2: u32;

    if i:
        val1 = ~src1;
    else:
        val1 = src1;

    if j:
        val2 = ~src2;
    else:
        val2 = src2;

    let dest: u32;
    let flags_val, flags_mask: u5;
    switch (instruction):
        case AND:
            dest = val1 & val2;
            flags_val = LogicCondition(dest);
            flags_mask = 0x3;
        case OR:
            dest = val1 | val2;
            flags_val = LogicCondition(dest);
            flags_mask = 0x3;
        case XOR:
            dest = val1 ^ val2;
            flags_val = LogicCondition(dest);
            flags_mask = 0x3;
    if f:
        flags = flags_val & flags_mask | nondeterministic() & ~flags_mask;

Funnel Shifts

MnemonicOpcodePayload
7------031---------------------8
FSL00001110rrrrrw0xfqqqqqvvvvvddddd
FSR00001111rrrrrw0xfqqqqqvvvvvddddd
  • r: Shift Remainder (Input value)
  • w: Wrap Quantity
  • x: Invert by Sign
  • f: Enable Flags Modification
  • q: Shift Quantity
  • v: Input Value
  • d: Destination Register
  • Ex[2]: If any reserved (fixed) bit is set to an invalid value.

Sets P, Z, and N according to the result. Sets C if any 1 bit was shifted out of v. Sets V if q is greater than 32 (regardless of w)

Shifts v by q and places the value in d, filling the shifted in bits with bits taken from the corresponding high bits of r. q wraps at 32 if w is set. If w is clear, excess shift quanities shift r in fully first.

  • FSL: v is shifted left by q
  • FSR: v is shifted right by q
instruction FSL(d: u5, v: u5, q: u5, f: bool, x: bool, w: bool, r: u5):
    let val = ReadRegister(0, v);
    let quantity = ReadRegister(0, q);
    let remainder = ReadRegister(0, r);
    if x & SignBitOf(val):
        remainder = ~remainder;

    let overflow: u5;
    if quantity >= 32:
        overflow = 2;
    else:
        overflow = 0;

    if w:
        quantity = quantity & 31;


    let result, out = ShiftInLeft(val, remainder, quantity);
    WriteRegister(0, d);
    let carry: u5;

    if out != 0:
        carry = 1;
    else
        carry = 0;

    if quantity
    let flags_val = LogicCondition(result) | carry | overflow;
    if f:
        flags = flags_val;


instruction FSR(d: u5, v: u5, q: u5, c: bool, x: bool, w: bool, r: u5):
    let val = ReadRegister(0, v);
    let quantity = ReadRegister(0, q);
    let remainder = ReadRegister(0, r);
    if x & SignBitOf(val):
        remainder = ~remainder;

    let overflow: u5;
    if quantity >= 32:
        overflow = 2;
    else:
        overflow = 0;

    if w:
        quantity = quantity & 31;


    let result, out = ShiftInRight(val, remainder, quantity);
    WriteRegister(0, d);
    let carry: u5;

    if out != 0:
        carry = 1;
    else
        carry = 0;

    if quantity
    let flags_val = LogicCondition(result) | carry | overflow;
    if f:
        flags = flags_val;

Branches

MnemonicOpcodePayload
7------031---------------------8
JMP00010000ooooooooooooooocccclllll
JMPR00010001000000000rrrrr0cccclllll
IRET00010010000000000000pp0000000000
  • o: Destination Offset (Bits 2..17)
  • r: Destination Register
  • c: Condition Code
  • l: Link Register
  • p: Target Interrupt Priority
  • Ex[2]: If any reserved (fixed) bit is set to an invalid value
  • IRET: Ex[2]: if p = 0
  • JMPR: Ex[2]: If r = 0
  • JMPR: Ex[3]: If the destination address is not 4 byte aligned (even if the branch is not taken)
  • Ex[1]: If fetching the next instruction at the destination causes a bus error, if the branch is taken

Jumps to the destination, if the condition is satisfied, saving the return address in l if taken:

  • JMP: The offset is IP + o * 4 where o is a signed offset. IP is the same as the return address and points to the beginning of the next instruction
  • JMPR: The offset is read from r
  • IRET: The offset it read from register p (p!=0) in Map 1. intctl.m is also loaded from p.m and intctl.a is cleared
instruction JMP(c: ConditionCode, l: u5, o: u15):
    let disp = SignExtend(o) << 2;
    let curr_ip = IP;
    if CheckCondition(flags, c):
        if l != 0:
            WriteRegister(0,l, curr_ip);
        let dest_ip = curr_ip + disp;
        if not CheckBranchTarget(dest_ip):
            Raise(Ex[1])
        IP = dest_ip;


instruction JMPR(c: ConditionCode, l: u5, r: u5):
    if r == 0:
        Raise(Ex[2]);
    let addr = ReadRegister(0,r);
    if addr & 3 != 0:
        Raise(EX[3]);
    let curr_ip = IP;
    if CheckCondition(flags, c):
        if l != 0:
            WriteRegister(0,l, curr_ip);
        if not CheckBranchTarget(addr):
            Raise(Ex[1])
        IP = addr;

instruction IRET(p: u2):
    if p == 0:
        Raise(Ex[2]);
    let reg = p as u5;
    let val = ReadRegister(1, reg);
    let addr = val & !3;
    if not CheckBranchTarget(dest_ip):
        Raise(Ex[1])
    IP = addr;
    IL.valid = false;
    WriteRegister(1, 0, val & 3);

Condition Code

JMP, JMPR, and MOV all use a 4-bit condition code to encode the branch condition. This includes conditions for “Always” and “Never”.

enum ConditionCode is u4:
    Never = 0,
    Carry = 1,
    Zero = 2,
    Overflow = 3,
    CarryOrEqual = 4,
    SignedLess = 5,
    SignedLessOrEq = 6,
    Negative = 7,
    Positive = 8,
    SignedGreater = 9,
    SignedGreaterOrEq = 10,
    Above = 11,
    NotOverflow = 12,
    NotZero = 13,
    NotCarry = 14,
    Always = 15

function CheckCondition(flags: u32, cc: ConditionCode) is bool:
    switch (cc):
        case Never:
            return false;
        case Carry:
            return (flags & c) != 0;
        case Zero:
            return (flags & z) != 0;
        case Overflow:
            return (flags & v) != 0;
        case CarryOrEqual:
            return (flags & c|z) != 0;
        case SignedLess:
            return (((flags & v) != 0) == ((flags & n) != 0)) and (flags & z) == 0;
        case SignedLessOrEq:
            return (((flags & v) != 0) == ((flags & n) != 0)) or (flags & z) != 0;
        case Negative:
            return (flags & n) != 0;
        case Positive:
            return (flags & n) == 0;
        case SignedGreater:
            return not ((((flags & v) != 0) == ((flags & n) != 0)) or (flags & z) != 0);
        case SignedGreaterOrEq:
            return not ((((flags & v) != 0) == ((flags & n) != 0)) and (flags & z) == 0);
        case Above:
            return (flags & c|z) == 0;
        case NotOverflow:
            return (flags & v) == 0;
        case NotZero:
            return (flags & z) == 0;
        case NotCarry:
            return (flags & c) == 0;
        case Always:
            return true;

I/O Transfers

MnemonicOpcodePayload
7------031---------------------8
IN00010100wwwww000000ppppppppddddd
OUT00010101wwwww000000ppppppppsssss
  • w: Transfer Bit Width
  • p: Port Number
  • d: Destination Transfer Register
  • s: Source Transfer Register
  • Ex[2]: If any reserved (fixed bit) is set to an invalid value

Shift w (in 1..=32, mod 32) bits in an io transfer register in or out to an I/O Port. w=0 = 32

  • IN : Shifts bits into the high bits of the transfer register
  • OUT: Shifts bits out of the low bits of the transfer register

Any IN or OUT instruction will not begin executing on a CPU until all IN or OUT instructions that occur before it have been fully written back. Additionally

  • Any OUT instruction does not begin executing until all LD instructions that occur before it have been fully written back, and all ST instructions that occur before it have fully written to main memory
  • Any LD or ST instruction does not begin executing until any IN instruction that occurs before it has been fully written back
instruction IN(s: u5, p: u8, w: u5):
    let val = RotateRight(ReadBitsFromPort(p,w),ExtendWidth(w));
    let regval = ReadRegister(2, s);
    let resval, bitsout = ShiftRightInOut(regval, val, ExtendWidth(w));
    WriteRegister(2,s, resval);

instruction OUT(s: u5, p: u8, w: u5):
    let regval = ReadRegister(2, s);
    let resval, bitsout = ShiftRightInOut(regval, 0, ExtendWidth(w));
    WriteRegister(2,s, resval);
    WriteBitsToPort(p, ExtendWidth(w), bitsout);

function ExtendWidth(w: u5) is u6:
    if w==0:
        return 0x20;
    else:
        w;

Flags Manipulation

MnemonicOpcodePayload
7------031---------------------8
LDFLAGS0001100000000000000000fffffddddd
STFLAGS0001100100000000000000fffffsssss
XVP00011010000000000000000000000000
  • f: Flag modification mask
  • d: Destination Register
  • s: Source Register
  • Ex[2]: If any reserved (fixed) bit is set to an invalid value.
  • LDFLAGS loads the flags bits into the lower 5 bits of d (zero extended)
  • STFLAGS stores the lower 5 bits of s into the flags bits, overwriting only flags set to 1 in f
  • XVP exchanges the v and p flags

The flags bits are, in order

4---0
pznvc
  • p: Parity
  • z: Zero
  • n: Negative
  • v: Signed Overflow
  • c: Carry
instruction LDFL(d: u5, f: u5):
    let val = ZeroExtend(flags & f);
    WriteRegister(0,d, val);

instruction STFL(s: u5, f: u5)
    let val = ReadRegister(0, s);
    flags = (val & f) | (flags & ~f);

instruct XVP():
    let temp = flags.p;
    flags.p = flags.v;
    flags.v = temp;

Exchange Register Contents

MnemonicOpcodePayload
7------031---------------------8
XCHG00011100000000000bbbbblccccaaaaa
  • b: Register 2
  • l: Latency Control
  • c: Condition Code (See Jump)
  • a: Register 1
  • Ex[2]: If any reserved (fixed) bit is set to an invalid value.

Exchanges GPR values a and b, if the condition check succeeds.

instruction XCHG(a: u5, b: u5, l: bool, c: ConditionCode):
    let val1 = ReadRegister(0, a);
    let val2 = ReadRegister(0, b);
    if CheckCondtion(flags, c):
        WriteRegister(0, a, val2);
        WriteRegister(0, b, val1);

Extend Register Contents

MnemonicOpcodePayload
7------031---------------------8
EXT00011101wwwww00000000xsssssddddd
  • w: Value width
  • x: Extend Kind (sign/zero)
  • s: Source
  • d: Destination
  • Ex[2]: If any reserved (fixed) bit is set to an invalid value
  • Ex[2]: If w = 0.

Masks only the lower w bits of a register, and extends it according to x

enum ExtKind:
    Sign = 0,
    Zero = 1

instruction EXT(dest: u5, src: u5, x: ExtKind, w: u5):
    if w==0:
        Raise(Ex[2])
    let val = ReadRegister(0, src) & (1 << w)-1;
    let res: u32;
    switch(x):
        case Sign:
            res = SignExtend(val, w);
        case Zero:
            res = val;
    WriteRegister(0, dest, res);

Swap Byte order

MnemonicOpcodePayload
7------031---------------------8
BSWAP00011110000000000000000sssssddddd
  • s: Source operand
  • d: Destination Operand
  • Ex[2]: If any reserved (fixed) bit is set to an invalid value

Swaps the order of bytes from the source value.

instruction BSWAP(s: u5, d: u5):
    let sval: u32 = ReadRegister(0, s);
    let rval: u32;
    rval[8:0] = sval[32:24];
    rval[16:8] = sval[24:16];
    rval[24:16] = sval[16:8];
    rval[32:24] = sval[8:0];
    WriteRegister(0, d);

Random Bits

MnemonicOpcodePayload
7------031---------------------8
RBGEN00011111wwwww000100000eeeeeddddd
  • w: Poll width
  • e: Status Destination
  • d: Destination
  • Ex[2]: If any reserved (fixed) bit is set to an invalid value

Polls a hardware random bit generator. If successful, writes w (in 1..=32, mod 32) random bits to d and clears flags.z. If unsuccesful, writes 0 to d and sets flags.z. In all cases, the current status of the RBG is stored to e. (TODO: Write out status format). Note that flags.z is only set depending on success/failure. In particular, a successful poll that results in all 0s (Approximately a 2^-(w+1) chance) will still clear flags.z.

The Random Bit Generator polled by the instruction shall have at least the following properties:

  • Each complete output from the instruction is independant from other outputs on any core
  • Each output from the instruction is distinct from all other outputs, with 2^-((w)/2) probability of collision.
  • If this instruction is used to generate at least 128 bits of randomness, which is then processed by a Cryptographic Hash Function, the resulting output shall have at least 64 bits of enthropy.
instruction RBGEN(d: u5, e: u5, w: u5):
    let valid, result, status = PollRand(ExtendWidth(w));
    WriteRegister(0, e, status);
    if valid:
        WriteRegister(0, d, result);
        flags.z = 0;
    else:
        WriteRegister(0, d, 0);
        flags.z = 1;

Status format:

+31-----------------------------0+
|r0000000000000sseeeeeeeeeeeeeeee|
+--------------------------------+
BitNameDescription
rRepeatableIf set to 1, operation may be retried immediately
sStatus CodeStatus code (See Below)
eEnthropy AvailableTotal ratio of enthropy available (*2^16)

The following status code values are used

Status CodeNameDescription
0NORMALNormal status/spurious failure
1UNAVAILRequired minimum enthropy unavailable
2PAUSEGenerator Paused/Errored (Recoverable)
3FAULTUnrecoverable Generator Error

The CPU shall ensure that it automatically attempts a reset of the Random Bit Generator after reporting a PAUSE status in finite time. In the case of a FAULT status, the Generator is only reset after a RESET.

Invoke Coprocessor Unit

MnemonicOpcodePayload
7------031---------------------8
CPIx00100xxxppppppppppppppppppppffff
NCPIx00101xxxppppppppppppppppppppffff
CPIxEF00110xxxppppppppppppppppppffffff
NCPIxEF00111xxxppppppppppppppppppffffff

(x is a value from 0 to 7, representing the co-processor number to invoke, for example, CPI0 has opcode 0x20 and NCPI7 has opcode 0x2F)

  • p: Co-processor instruction payload
  • f: Co-processor function
  • CPIx, CPIxEF, Ex[8+x]: If executing the instruction raises a coprocessor error

Executes the specified Coprocessor function with the specified payload

  • CPIx/CPIxEF: Waits for the Co-processor to finish all operations, and raises the appropriate unit error if the Coprocessor reports it,
  • NCPIx/NCPIxEF: Finishes immediately.
  • CPIx/NCPIx: Allows specifying up to 16 functions with a 20-bit payload
  • CPIxEF/NCPIxEF: Allows specifying up to 64 functions with a 18-bit payload (bottom 18-bits of the 20-bit payload)

If a CPIx or CPIxEF instruction begins execution on a core, the following guarantees are made:

  • Any CPIx, CPIxEF, NCPIx, or NCPIxEF instruction (for the same x) that occurs after will not begin executing until the CPIx or CPIxEF and all NCPIx and NCPIxEF instructions that occur before the CPIx or CPIxEF have been fully written back, and
  • Any MOV instruction that loads from a register in map 8+x will not begin executing until the CPIx or CPIxEF and all NCPIx and NCPIxEF instructions that occur before the CPIx or CPIxEF have been fully written back.
instruction {CPI0, CPI1, CPI2, CPI3}(f: u4, p: u20):
    let coproc: u4;
    switch (instruction):
        case CPI0:
            coproc = 0;
        case CPI1:
            coproc = 1;
        case CPI2:
            coproc = 2;
        case CPI3:
            coproc = 3;
    if not IsCoprocessorEnabled(coproc):
        Raise(EX[3]);

    ExecuteCoprocessorInstruction(coproc, f, p);
    WaitOnCoprocessor(coproc);
    CheckAndRaisePending(EX[8+coproc]);

instruction {CPI0EF, CPI1EF, CPI2EF, CPI3EF}(f: u6, p: u18):
    let coproc: u4;
    switch (instruction):
        case CPI0EF:
            coproc = 0;
        case CPI1EF:
            coproc = 1;
        case CPI2EF:
            coproc = 2;
        case CPI3EF:
            coproc = 3;
    if not IsCoprocessorEnabled(coproc):
        Raise(EX[3]);

    ExecuteCoprocessorInstruction(coproc, f, p);
    WaitOnCoprocessor(coproc);
    CheckAndRaisePending(EX[8+coproc]);

instruction {NCPI0, NCPI1, NCPI2, NCPI3}(f: u4, p: u20):
    let coproc: u4;
    switch (instruction):
        case NCPI0:
            coproc = 0;
        case NCPI1:
            coproc = 1;
        case NCPI2:
            coproc = 2;
        case NCPI3:
            coproc = 3;
    if not IsCoprocessorEnabled(coproc):
        Raise(EX[3]);

    ExecuteCoprocessorInstruction(coproc, f, p);

instruction {NCPI0EF, NCPI1EF, NCPI2EF, NCPI3EF}(f: u6, p: u18):
    let coproc: u4;
    switch (instruction):
        case NCPI0EF:
            coproc = 0;
        case NCPI1EF:
            coproc = 1;
        case NCPI2EF:
            coproc = 2;
        case NCPI3EF:
            coproc = 3;
    if not IsCoprocessorEnabled(coproc):
        Raise(EX[3]);

    ExecuteCoprocessorInstruction(coproc, f, p);

Halt/Stop CPU

MnemonicOpcodePayload
7------031---------------------8
HALT010000000000000000000000000000mm

Places the CPU in a low-power state and stops executing. The CPU responds to interrupts as though ictl.m was set temporarily to m. The CPU resumes execution after receiving an interrupt that is valid at priority m (if m=0 then the CPU will never resume execution).

  • Ex[2]: If any reserved (fixed) bit is set to an invalid value

The HALT instruction will not begin executing until instructions that occur before it have been fully written back, and no instructions that occur after it will begin executing until the HALT instruction is fully written back.

instruction HALT(m: u2) {
    let saved_ictl: u32 = ReadRegister(1, 0);
    WriteRegister(1, 0, ZeroExtend(m) | (saved_ictl & (1 << 31)));

    if m != 0:
        SetStatus(2);
        WaitForInterrupt();
        WriteRegister(1, 0, saved_ictl);
    else:
        SetStatus(3);
        ShutdownCpu();
}

Debugging Hint

MnemonicOpcodePayload
7------031---------------------8
BREAKP01000001000000000000000000000000
  • Ex[2] If any undefined bit is set.
  • Ex[5]: If a handler for Ex[5] is present and priority 1 exceptions are not masked by intctl.

Hints that a debugger attached to the machine should take control of execution at this point. If a handler for Ex[5] is present and priority 1 interrupts are not masked, raises that exception, with a return address pointing to the next instruction.

instruction BREAKP():
    if CpuDebuggerPresent():
        BreakToDebugger();
    TryInterruptCpu(5, 1);

Interlocked instructions

MnemonicOpcodePayload
7------031---------------------8
FENCE01001000rr0000000000000000000000
STIC01001011rr0000001000wwsssssddddd
LDIL01001100rr0000000000wwsssssddddd
STICW01001101rr0000001bbbbbsssssddddd
LDILW01001110rr0000000bbbbbsssssddddd
  • r: Atomic Ordering
  • w: Width
  • b: Second source/destination register
  • s: Source Register
  • d: Destination Register
  • STIL, STILW: EX[1]: If d is unaligned
  • LDIL, LDILW: EX[1]: If s is unaligned
  • STIL, STILW: EX[2]: If d = 0
  • LDIL, LDILW: EX[2]: If s = 0
  • EX[1]: If a bus fault occurs
  • EX[2]: If w = 3
  • FENCE: EX[2]: if r = 0
  • LDIL, LDILW: EX[2]: if r = 2
  • STIC, STICW: EX[2]: if r = 1
  • STIC and STICW set z if the validation check fails. In this case, no memory write or synchronization occurs. All other flags are set to undefined values.
  • FENCE: Serializes memory between processors according to r.
  • STIC: Store s completing interlocked sequence on d. On success, the z flag is clear, and the store is guaranteed to be visible to any LDIL or LDILW instruction that is completed by a successful STIC instruction. It is also guaranteed that any ST instruction on any thread that was not observed by the LDIL instruction will not be overwritten by the STIC instruction.
  • LDIL: Load from s into d, starting an interlocked sequence on s with IL.width = w
  • STICW: Stores the 8-byte value in b:s into d, completing interlocked sequence on d. On success, the z flag is clear, and the store is guaranteed to be visible to any LDIL or LDILW instruction that is completed by a successful STIC instruction. It is also guaranteed that any ST instruction on any thread that was not observed by the LDILW instruction will not be overwritten by the STICW instruction.
  • LDILW: Loads from s into an 8-byte value in b:d, starting an interlocked sequence on s with IL.width = 3

Each a core acts as though it stores the following additional state:

  • Whether or not a valid interlocked transaction is occur (IL.valid)
  • The address of the most recently begun interlocked transaction (IL.addr) if IL.valid is true
  • The width of the most recently begun interlocked transaction (IL.width)

The validity of an interlocked transaction on a core is reset when any of the following occurs:

  • An STIC or STICW instruction completes execution, whether or not it was successful or failed
  • An interrupt or exception occurs
  • The iret instruction is issued
  • If transaction is invalidated by a memory write on any core that violates the guarantees of any subsequent STIC or STICW instruction.

An STIC instruction fails (sets z = 1 and does not modify any memory) if:

  • IL.valid is false
  • IL.addr refers to a different address than the STIC instruction
  • IL.width refers to a different width than the STIC instruction

An STICW instruction fails (sets z = 1 and does not modify any memory) if:

  • IL.valid is false
  • IL.addr refers to a different address than the STICW instruction
  • IL.width is not 3.

An STIC or STICW instruction does not modify any memory if it fails. It is unspecified whether the write access check is performed.

On a multicore system, any instruction that synchronizes memory according to the Acquire or SeqCst order guarantees that the instruction will not write back its result until, for each value loaded by any of the following instructions, all memory accesses visible to the corresponding store instruction will be observed by any instruction that occurs after that instruction:

  • For FENCE: Any LD, LDIL, or LDILW instruction that preceeds it
  • For LD, LDIL, or LDILW: That instruction.

On a multicore system, any instruction that synchronizes memory according to the Release or SeqCst order guarantees that the write will not be observed by any core until all memory operations that occur before it have completed and become visible to the instruction:

  • For FENCE; Any ST, STIC, or STICW instruction that occurs after it
  • For ST, STIC, or STICW: That insttruction.

On a multicore system, any instruction that synchronizes memory according to the SeqCst order guarantees that all cores will observe the same order of memory effects caused by all such instructions.

instruction FENCE(r: Ordering):
    if r == Relaxed:
        Raise(EX[2]);
    SynchronizeMemoryAccordingToFence(r);

instruction STIC(d: u5, s: u5, w: u2, r: Ordering):
    if r == Acquire:
        Raise(EX[2]);
    if w == 3:
        Raise(EX[2]);
    let dest = ReadRegister(0, d);
    if dest & (1 << w)-1 != 0:
        Raise(EX[1]);

    let failed: u5;
    let value = ReadRegister(0, s);
    if not IL.valid or IL.addr != dest or IL.width != w:
        failed = 8;
    else:
        failed = TryInterlockedMemoryWrite(dest, w, value);
        CheckAndRaisePending(EX[1]);

    IL.valid = false;

    flags = failed | nondeterministic() & ~8;

instruction STICW(d: u5, s: u5, b: u5, r: Ordering):
    if r == Acquire:
        Raise(EX[2]);
    let dest = ReadRegister(0, d);
    if dest & 7 != 0:
        Raise(EX[1]);

    let failed: u5;
    let value = ReadRegister(0, s);
    let value_hi = ReadRegister(0, b);
    SynchronizeMemoryAccordingToWrites(r);
    if not IL.valid or IL.addr != dest or IL.width != 3:
        failed = 8;
    else:
        failed = TryInterlockedMemoryWriteWide(dest, value, value_hi);
        CheckAndRaisePending(EX[1]);

    IL.valid = false;

    flags = failed | nondeterministic() & ~8;

instruction LDIL(d: u5, s: u5, w: u2, r: Ordering):
    if r == Release:
        Raise(EX[2]);
    if w == 3:
        Raise(EX[2]);

    let src = ReadRegister(0, s);
    if src & (1 << w)-1 != 0:
        Raise(EX[1]);
    let value = InterlockedMemoryRead(src, w);
    CheckAndRaisePending(EX[1]);
    IL.valid = true;
    IL.addr = src;
    IL.width = w;
    WriteRegister(0, d, value);

instruction LDILW(d: u5, s: u5, b: u5, r: Ordering):
    if r == Release:
        Raise(EX[2]);
    if w == 3:
        Raise(EX[2]);

    let src = ReadRegister(0, s);
    if src & (1 << w)-1 != 0:
        Raise(EX[1]);
    let value, value_hi = InterlockedMemoryReadWide(src);
    CheckAndRaisePending(EX[1]);
    IL.valid = true;
    IL.addr = src;
    IL.width = 3;
    WriteRegister(0, d, value);
    WriteRegister(0, b, value_hi);