// -------------------------------------------------------------------------------- // // CONTENTS // // Chapter 1 - INTRODUCTION // // Chapter 2 - NV_PPBDMA GPENTRY DATA FORMAT // // Chapter 3 - NV_PPBDMA PUSHBUFFER REGISTERS // // Chapter 4 - NV_PPBDMA INTERRUPT REGISTERS // // Chapter 5 - NV_PPBDMA LATENCY BUFFER REGISTERS // // Chapter 6 - HOST METHODS (NV_UDMA) // // Chapter 7 - RESERVED METHOD ADDRESSES // // Appendix A - KEY // 1 - INTRODUCTION // ================== // // Each PBDMA unit works on behalf of a single ESCHED (Engine Scheduler). The // PBDMA fetches pushbuffer data from memory, generates methods from the fetched // data, executes some of the generated methods itself, and sends the remainder of // the methods to one of the engines associated with the ESCHED. // This manual describes the arrayed NV_PPBDMA register space and all Host // methods. The NV_PPBDMA space defines registers that are contained within each // of Host's PBDMA units. Each PBDMA unit is allocated a 2KB address space for its // registers. // // The NV_UDMA space defines the Host methods. Examples include // SetObject, SEM_EXECUTE, and MemOps. A method consists of an // address doubleword and a data doubleword. The address specifies the operation // to be performed. The data is an operand. The NV_UDMA address space contains // the addresses of the methods that are executed by a PBDMA unit. // Note: The Host methods are defined using the register definition format. // Doo not confuse them with the many similarly named NV_PBDMA registers which // provide the internal Host state for these methods. // The software interface to the Host methods is specified separately in the // //hw/nvgpu/manuals/cl??6f.h header files. Together, the class header file and // this manual specify the Host class in lieu of a master .mfs file. // // Mnemonic Description Size Interface // -------- ----------- ---- --------- // PPBDMA Priv PBDMA Unit 128K HOST // UDMA Host Methods 256B // See the autogenerated dev_pbdma_NV_PPBDMA.refh include file built from dev_pbdma_NV_PBDMA.refh. #define NV_PPBDMA 0x0005ffff:0x00040000 /* RW--D */ // Local Devices // 2 - NV_PPBDMA GPENTRY DATA FORMAT // =================================== // // NV_PPBDMA_GP_ENTRY - GP-Entry Memory Format // // A pushbuffer contains the specifications of the operations that a GPU // context is to perform for a particular client. Pushbuffers are stored in // memory. A doubleword-sized (4-byte) unit of pushbuffer data is known as a // pushbuffer entry. GP entries indicate the location of the pushbuffer data in // memory. GP entries themselves are also stored in memory. // A GP entry specifies the location and size of a pushbuffer segment (a // contiguous block of PB entries) in memory. See "FIFO_DMA" in dev_ram.ref for // details about pushbuffer segments and the format of pushbuffer data. // // The NV_PPBDMA_GP_ENTRY0_GET and NV_PPBDMA_GP_ENTRY1_GET_HI fields of a GP // entry specify the 38-bit dword-address (which would make a 40-bit byte-address) // of the first pushbuffer entry of the GP entry's pushbuffer segment. Because // each pushbuffer entry (and by extension each pushbuffer segment) is doubleword // aligned (4-byte aligned), the least significant 2 bits of the 40-bit // byte-address are not stored. The byte-address of the first pushbuffer entry in // a GP entry's pushbuffer segment is // (NV_PPBDMA_GP_ENTRY1_GET_HI << 32) + (NV_PPBDMA_GP_ENTRY0_GET << 2). If the pushbuffer segment // address plus its length as specified by LENGTH*4 exceeds the bounds of // addressable virtual memory, the PBDMA raises the NV_PPBDMA_INTR_0_GPENTRY // interrupt. // The NV_PPBDMA_GP_ENTRY1_LENGTH field, when non-zero, indicates the number // of pushbuffer entries contained within the GP entry's pushbuffer segment. The // byte-address of the first pushbuffer entry beyond the pushbuffer segment is // (NV_PPBDMA_GP_ENTRY1_GET_HI << 32) + (NV_PPBDMA_GP_ENTRY0_GET << 2) + (NV_PPBDMA_GP_ENTRY1_LENGTH * 4). // If this address exceeds addressable virtual memory, the PBDMA raises the // NV_PPBDMA_INTR_0_GPENTRY interrupt. // If NV_PPBDMA_GP_ENTRY1_LENGTH is CONTROL (0), then the GP entry is a // "control" entry, meaning this GP entry will not cause any PB data to be fetched // or executed. In this case, the NV_PPBDMA_GP_ENTRY1_OPCODE field specifies an // operation to perform, and the NV_PPBDMA_GP_ENTRY0_OPERAND field contains the // operand. The available operations are as follows: // // * NV_PPBDMA_GP_ENTRY1_OPCODE_NOP: no operation will be performed, but note // that the SYNC field is still respected--see below. // // * NV_PPBDMA_GP_ENTRY1_OPCODE_GP_CRC: the ENTRY0_OPERAND field is compared // with the cyclic redundancy check value that was calculated over previous // GP entries (NV_PPBDMA_GP_CRC). After each comparison, the // NV_PPBDMA_GP_CRC is cleared, whether they match or differ. If they // differ, then Host initiates an interrupt (NV_PPBDMA_INTR_0_GPCRC). For // recovery, clearing the interrupt will cause the PBDMA to continue as if // the control entry was OPCODE_NOP. // // * NV_PPBDMA_GP_ENTRY1_OPCODE_PB_CRC: the ENTRY0_OPERAND is compared // with the CRC value that was calculated over the previous pushbuffer // segment (NV_PPBDMA_PB_CRC). The PB CRC resets to 0 with each pushbuffer // segment. If the two CRCs differ, Host will raise the // NV_PPBDMA_INTR_0_PBCRC interrupt. For recovery, clearing the interrupt // will continue as if the control entry was OPCODE_NOP. Note the PB_CRC is // indeterminate if an END_PB_SEGMENT PB control entry was used in the prior // segment or if SSDM disabled the device and the segment had conditional // fetching enabled. // // Host supports two privilege levels for channels: privileged and // non-privileged. The privilege level is determined by the // NV_PPBDMA_CONFIG_AUTH_LEVEL field set from the corresponding NV_RAMFC_CONFIG // dword in the RAMFC. Non-privileged channels cannot execute privileged methods, // but privileged channels can. Any attempt to run a privileged operation from a // non-privileged channel will result in PB raising NV_PPBDMA_INTR_0_METHOD. // // The NV_PPBDMA_GP_ENTRY1_SYNC field specifies whether a pushbuffer may be // fetched before Host has finished processing the preceding PB segment. If this // field is SYNC_PROCEED, then Host does not wait for the preceding PB segment to // be processed. If this field is SYNC_WAIT, then Host waits until the preceding // PB segment has been processed by Host before beginning to fetch the current PB // segment. // Host's processing of a PB segment consists of parsing PB entries into PB // instructions, decoding those instructions into control entries or method // headers, generating methods from method headers, determining whether methods are // to be executed by Host or by an engine, executing Host methods, and sending // non-Host methods and SetObject methods to engines. // Note that in the case where the final PB entry of the preceding PB segment // is a method header representing a PB compressed method sequence of nonzero // length--that is, the compressed method sequence is split across PB segments with // all of its method data entries in the PB segment for which SYNC_WAIT is // set--then Host is considered to have finished processing the preceding PB // segment once that method header is read. However, splitting a PB compressed // method sequence for software methods is not supported because Host will issue // the DEVICE interrupt indicating the SW method as soon as it processes the // method header, which happens prior to fetching the method data entries for that // compressed method sequence. Thus SW cannot actually execute any of the methods // in the sequence because the method data is not yet available, leaving the PBDMA // wedged. // When SYNC_WAIT is set, Host does not wait for any engine methods generated // from the preceding PB segment to complete. Host does not automatically wait // until an engine is done processing all methods generated from that PB segment. // If software desires that the engine finish processing all methods generated from // one PB segment before a second PB segment is fetched, then software may place // Host methods that wait until the engine is idle in the first PB segment (like // WFI, SET_REF, or SEM_EXECUTE with RELEASE_WFI_EN set). Alternatively, software // might put a semaphore acquire at the end of the first PB segment, and have an // engine release the semaphore. In both cases, SYNC_WAIT must be set on the // second PB segment. This field applies even if the NV_PPBDMA_GP_ENTRY1_LENGTH // field is zero; if SYNC_WAIT is specified in this case, no further GP entries // will be processed until the wait finishes. // // Some parts of a pushbuffer may not be executed depending on the value of // the NV_PPBDMA_SUBDEVICE_ID and SUBDEVICE_MASK. If an entire PB segment will not // be executed due to conditional execution, Host need not even bother fetching the // PB segment. // The NV_PPBDMA_GP_ENTRY0_FETCH field indicates whether the PB segment // specified by the GP entry should be fetched unconditionally or fetched // conditionally. If this field is FETCH_UNCONDITIONAL, then the PB segment is // fetched unconditionally. If this field is FETCH_CONDITIONAL, then the PB // segment is only fetched if the NV_PPBDMA_SUBDEVICE_STATUS field is // STATUS_ACTIVE. // // ******************************************************************************** // Warning: When using subdevice masking, one must take care to synchronize // properly with any later GP entries marked FETCH_CONDITIONAL. If GP fetching // gets too far ahead of PB processing, it is possible for a later conditional PB // segment to be discarded prior to reaching an SSDM command that sets // SUBDEVICE_STATUS to ACTIVE. This would cause Host to execute garbage data. One // way to avoid this would be to set the SYNC_WAIT flag on any FETCH_CONDITIONAL // segments following a subdevice reenable. // ******************************************************************************** // // If the PB segment is not fetched then it behaves as an OPCODE_NOP control // entry. If a PB segment contains a SET_SUBDEVICE_MASK PB instruction that Host // must see, then the GP entry for that PB segment must specify // FETCH_UNCONDITIONAL. // If the PB segment specifies FETCH_CONDITIONAL and the subdevice mask shows // STATUS_ACTIVE, but the PB segment contains a SET_SUBDEVICE_MASK PB instruction // that will disable the mask, the rest of the PB segment will be discarded. In // that case, an arbitrary number of entries past the SSDM may have already updated // the PB CRC, rendering the PB CRC indeterminate. // If Host must wait for a previous PB segment's Host processing to be // completed before examining NV_PPBDMA_SUBDEVICE_STATUS, then the GP entry should // also have its SYNC_WAIT field set. // A PB segment marked FETCH_CONDITIONAL must not have a PB compressed method // sequence that crosses a PB segment boundary (with its header in previous non- // conditional PB segment and its final valid data in a conditional PB segment)-- // doing so will cause a NV_PPBDMA_INTR_0_PBSEG. // // Software may monitor Host's progress through the pushbuffer by reading the // channel's NV_RAMUSERD_TOP_LEVEL_GET entry from USERD, which is backed by Host's // NV_PPBDMA_TOP_LEVEL_GET register. See "NV_RUNLIST_USERD_WRITEBACK" in // dev_runlist.ref for information about how frequently this information is written // back into USERD. If a PB segment occurs multiple times within a pushbuffer // (like a commonly used subroutine), then progress through that segment may be // less useful for monitoring, because software will not know which occurrence of // the segment is being processed. // The NV_PPBDMA_GP_ENTRY1_LEVEL field specifies whether progress through the // GP entry's PB segment should be indicated in NV_RAMUSERD_TOP_LEVEL_GET. If this // field is LEVEL_MAIN, then progress through the PB segment will be reported -- // NV_RAMUSERD_TOP_LEVEL_GET will equal NV_RAMUSERD_GET. If this field is // LEVEL_SUBROUTINE, then progress through this PB segment is not reported -- Host // will not alter NV_RAMUSERD_TOP_LEVEL_GET. If this field is LEVEL_SUBROUTINE, // reads of NV_RAMUSERD_TOP_LEVEL_GET will return the last value of NV_RAMUSERD_GET // from a PB segment at LEVEL_MAIN. // // If the GP entry's opcode is OPCODE_ILLEGAL or an invalid opcode, Host will // initiate an interrupt (NV_PPBDMA_INTR_0_GPENTRY). If a GP entry specifies a PB // segment that crosses the end of the virtual address space (0xFFFFFFFFFF), then // Host will initiate an interrupt (NV_PPBDMA_INTR_0_GPENTRY). Invalid GP entries // are treated like traps: they will set the interrupt and freeze the PBDMA, but // the invalid GP entry is discarded. Once the interrupt is cleared, the PBDMA // unit will simply continue with the next GP entry. // Note a corner case exists where the PB segment described by a GP entry is // at the end of the virtual address space, or in other words, the last PB entry in // the described PB segment is the last dword in the virtual address space. This // type of GP entry is not valid and will generate a GPENTRY interrupt. The // PBDMA's PUT pointer describes the address of the first dword beyond the PB // segment, thus making the last dword in the virtual address space unusable for // storing pbentry. #define NV_PPBDMA_GP_ENTRY__SIZE 8 /* */ #define NV_PPBDMA_GP_ENTRY0 0x10000000 /* */ #define NV_PPBDMA_GP_ENTRY0_OPERAND 31:0 /* */ #define NV_PPBDMA_GP_ENTRY0_FETCH 0:0 /* */ #define NV_PPBDMA_GP_ENTRY0_FETCH_UNCONDITIONAL 0x00000000 /* */ #define NV_PPBDMA_GP_ENTRY0_FETCH_CONDITIONAL 0x00000001 /* */ #define NV_PPBDMA_GP_ENTRY0_GET 31:2 /* */ #define NV_PPBDMA_GP_ENTRY1 0x10000004 /* */ #define NV_PPBDMA_GP_ENTRY1_GET_HI 7:0 /* */ #define NV_PPBDMA_GP_ENTRY1_LEVEL 9:9 /* */ #define NV_PPBDMA_GP_ENTRY1_LEVEL_MAIN 0x00000000 /* */ #define NV_PPBDMA_GP_ENTRY1_LEVEL_SUBROUTINE 0x00000001 /* */ #define NV_PPBDMA_GP_ENTRY1_LENGTH 30:10 /* */ #define NV_PPBDMA_GP_ENTRY1_LENGTH_CONTROL 0x00000000 /* */ #define NV_PPBDMA_GP_ENTRY1_SYNC 31:31 /* */ #define NV_PPBDMA_GP_ENTRY1_SYNC_PROCEED 0x00000000 /* */ #define NV_PPBDMA_GP_ENTRY1_SYNC_WAIT 0x00000001 /* */ #define NV_PPBDMA_GP_ENTRY1_OPCODE 7:0 /* */ #define NV_PPBDMA_GP_ENTRY1_OPCODE_NOP 0x00000000 /* */ #define NV_PPBDMA_GP_ENTRY1_OPCODE_ILLEGAL 0x00000001 /* */ #define NV_PPBDMA_GP_ENTRY1_OPCODE_GP_CRC 0x00000002 /* */ #define NV_PPBDMA_GP_ENTRY1_OPCODE_PB_CRC 0x00000003 /* */ // // 3 - NV_PPBDMA PUSHBUFFER REGISTERS // ==================================== // // Note: As most of these registers directly reflect the current state of the PBDMA // this means that while a Host channel switch is in progress the registers may be // in an inconsistent state until the channel switch is complete. See // NV_PPBDMA_STATUS_SCHED for more information on how to tell if a chsw is in // progress. // // Number of NOPs for self-modifying gpfifo // // This is a formula for SW to estimate the number of NOPs needed to pad the gpfifo // such that the modification of a gp entry by the engine or by the CPU can take // effect. Here, NV_PPBDMA_LB_GPBUF_CONTROL_SIZE refers to the SIZE field in the // NV_PPBDMA_LB_GPBUF_CONTROL(pbdma) register. // // NUM_GP_NOPS = ((NV_PPBDMA_LB_GPBUF_CONTROL_SIZE + 1) * NV_PPBDMA_LB_ENTRY_SIZE)/ NV_PPBDMA_GP_ENTRY__SIZE // // // // // GP_BASE - Base and Limit of the Circular Buffer of GP Entries // // GP entries are stored in a buffer in memory. The NV_PPBDMA_GP_BASE_OFFSET // and NV_PPBDMA_GP_BASE_HI_OFFSET fields specify the 37-bit address in 8-byte // granularity of the start of a circular buffer that contains GP entries (GPFIFO). // This address is a virtual (not a physical) address. GP entries are always // NV_PPBDMA_GP_ENTRY__SIZE-byte aligned, so the least significant three bits of the byte // address are not stored. The byte address of the GPFIFO base pointer is thus: // // gpfifo_base_ptr = GP_BASE + (GP_BASE_HI_OFFSET << 32) // // The number of GP entries in the circular buffer is always a power of 2. // The NV_PPBDMA_GP_BASE_HI_LIMIT2 field specifies the number of bits used to count // the memory allocated to the GP FIFO. The LIMIT2 value specified in these // registers is Log base 2 of the number of entries in the GP FIFO. For example, // if the number of entries is 2^16--indicating a memory area of // (2^16)*NV_PPBDMA_GP_ENTRY__SIZE bytes--then the value written in LIMIT2 is 16. // The circular buffer containing GP entries cannot cross the maximum address. // If OFFSET + (1< 0xFFFFFFFFFF, then Host will // initiate a CPU interrupt (NV_PPBDMA_INTR_0_GPFIFO). // The NV_PPBDMA_GP_PUT, NV_PPBDMA_GP_GET, and NV_PPBDMA_GP_FETCH registers // (and their associated NV_RAMFC and NV_RAMUSERD entries) are relative to the // value of this register. // These registers are part of a GPU context's state. On a switch, the values // of these registers are saved to, and restored from, the NV_RAMFC_GP_BASE and // NV_RAMFC_GP_BASE_HI entries in the RAMFC part of the GPU context's GPU-instance // block. // Typically, software initializes the information in NV_RAMFC_GP_BASE and // NV_RAMFC_GP_BASE_HI when the GPU context's GPU-instance block is first created. // These registers are available to software only for debug. Software should use // them only if the GPU context is assigned to a PBDMA unit and that PBDMA unit is // stalled. While a GPU context's Host context is not contained within a PBDMA // unit, software should use the RAMFC entries to access this information. // A pair of these registers exists for each of Host's PBDMA units. These // registers run on Host's internal bus clock. #define NV_PPBDMA_GP_BASE(i) (0x00040048+(i)*2048) /* RW-4A */ #define NV_PPBDMA_GP_BASE__SIZE_1 32 /* */ #define NV_PPBDMA_GP_BASE_OFFSET 31:3 /* RWXUF */ #define NV_PPBDMA_GP_BASE_OFFSET_ZERO 0x00000000 /* RW--V */ #define NV_PPBDMA_GP_BASE_RSVD 2:0 /* RWXUF */ #define NV_PPBDMA_GP_BASE_RSVD_ZERO 0x00000000 /* RW--V */ #define NV_PPBDMA_GP_BASE_HI(i) (0x0004004c+(i)*2048) /* RW-4A */ #define NV_PPBDMA_GP_BASE_HI__SIZE_1 32 /* */ #define NV_PPBDMA_GP_BASE_HI_OFFSET 7:0 /* RWXUF */ #define NV_PPBDMA_GP_BASE_HI_OFFSET_ZERO 0x00000000 /* RW--V */ #define NV_PPBDMA_GP_BASE_HI_LIMIT2 20:16 /* RWXUF */ #define NV_PPBDMA_GP_BASE_HI_LIMIT2_ZERO 0x00000000 /* RW--V */ #define NV_PPBDMA_GP_BASE_HI_RSVDA 15:8 /* RWXUF */ #define NV_PPBDMA_GP_BASE_HI_RSVDA_ZERO 0x00000000 /* RW--V */ #define NV_PPBDMA_GP_BASE_HI_RSVDB 31:21 /* RWXUF */ #define NV_PPBDMA_GP_BASE_HI_RSVDB_ZERO 0x00000000 /* RW--V */ // // GP_FETCH - Pointer to the next GP-Entry to be Fetched // // Host does not fetch all GP entries with a single request to the memory // subsystem. Host fetches GP entries in batches. The NV_PPBDMA_GP_FETCH register // indicates the index of the next GP entry to be fetched by Host. The actual 40-bit // virtual address of the specified GP entry is computed as follows: // fetch address = GP_FETCH_ENTRY * NV_PPBDMA_GP_ENTRY__SIZE + GP_BASE // If NV_PPBDMA_GP_PUT==NV_PPBDMA_GP_FETCH, then requests to fetch the entire // GP circular buffer have been issued, and Host cannot make more requests until // NV_PPBDMA_GP_PUT is changed. Host may finish fetching GP entries long before it // has finished processing the PB segments specified by those entries. // Software should not use NV_PPBDMA_GP_FETCH (it should use NV_PPBDMA_GP_GET), to // determine whether the GP circular buffer is full. NV_PPBDMA_GP_FETCH represents // the current extent of prefetching of GP entries; prefetched entries may be // discarded and refetched later. // This register is part of a GPU context's state. On a switch, the value of // this register is saved to, and restored from, the NV_RAMFC_GP_FETCH entry of // the RAMFC part of the GPU context's GPU-instance block. // A PBDMA unit maintains this register. Typically, software does not need to // access this register. This register is available to software only for debug. // Because Host may fetch GP entries long before it is ready to process the // entries, and because Host may discard GP entries that it has fetched, software // should not use NV_PPBDMA_GP_FETCH to monitor Host's progress (software should // use NV_PPBDMA_GP_GET for monitoring). Software should use this register only if // the GPU context is assigned to a PBDMA unit and that PBDMA unit is stalled. // While a GPU context's Host context is not contained within a PBDMA unit, // software should use NV_RAMFC_GP_FETCH to access this information. // If after a PRI write, or after this register has been restored from RAMFC // memory, the value equals or exceeds the size of the circular buffer that stores // GP entries (1<LOAD->VALID->INVALID. // // NEXT_TSGID field: // // This field specifies the TSG ID of the next context to be loaded on the // PBDMA. This field is valid just before or during a PBDMA channel switch. // NEXT_TSGID is valid when the CHANLOAD field alias reports IN_PROGRESS. // // CHAN field alias: // // This field indicates whether the TSGID field is valid. It overlays a // single bit of the CHAN_STATUS field so as to report VALID when CHAN_STATUS // reports VALID, CHSW_SAVE, or CHSW_SWITCH, but INVALID when CHAN_STATUS reports // INVALID. // // CHANLOAD field alias: // // This field indicates whether a load is in in-progress. It overlays a // single bit of the CHAN_STATUS field so as to report VALID exactly when // CHAN_STATUS reports CHSW_LOAD or CHSW_SWITCH. When CHANLOAD reports valid, // the NEXT_TSGID field is valid, but not otherwise. // // CHSW field alias: // // CHSW gives the current channel switch state of the PBDMA and aliases the // top bit of the CHAN_STATUS field. CHSW is IN_PROGRESS if we're currently // loading a new channel, saving a channel, or switching between two valid // channels. Otherwise, the value is NOT_IN_PROGRESS. CHSW reports IN_PROGRESS // exactly when CHAN_STATUS reports CHSW_SAVE, CHSW_LOAD, or CHSW_SWITCH, and // NOT_IN_PROGRESS otherwise. #define NV_PPBDMA_STATUS_SCHED(i) (0x0004015c+(i)*2048) /* R--4A */ #define NV_PPBDMA_STATUS_SCHED__SIZE_1 32 /* */ #define NV_PPBDMA_STATUS_SCHED_TSGID 11:0 /* */ #define NV_PPBDMA_STATUS_SCHED_TSGID_HW 10:0 /* R-XUF */ #define NV_PPBDMA_STATUS_SCHED_CHAN_STATUS 15:13 /* R-EVF */ #define NV_PPBDMA_STATUS_SCHED_CHAN_STATUS_INVALID 0x00000000 /* R-E-V */ #define NV_PPBDMA_STATUS_SCHED_CHAN_STATUS_VALID 0x00000001 /* R---V */ #define NV_PPBDMA_STATUS_SCHED_CHAN_STATUS_CHSW_SAVE 0x00000005 /* R---V */ #define NV_PPBDMA_STATUS_SCHED_CHAN_STATUS_CHSW_LOAD 0x00000006 /* R---V */ #define NV_PPBDMA_STATUS_SCHED_CHAN_STATUS_CHSW_SWITCH 0x00000007 /* R---V */ #define NV_PPBDMA_STATUS_SCHED_NEXT_TSGID 27:16 /* */ #define NV_PPBDMA_STATUS_SCHED_NEXT_TSGID_HW 26:16 /* R-XUF */ // SW convenience macros #define NV_PPBDMA_STATUS_SCHED_CHAN 13:13 /* */ #define NV_PPBDMA_STATUS_SCHED_CHAN_INVALID 0x00000000 /* */ #define NV_PPBDMA_STATUS_SCHED_CHAN_VALID 0x00000001 /* */ #define NV_PPBDMA_STATUS_SCHED_CHANLOAD 14:14 /* */ #define NV_PPBDMA_STATUS_SCHED_CHANLOAD_NOT_IN_PROGRESS 0x00000000 /* */ #define NV_PPBDMA_STATUS_SCHED_CHANLOAD_IN_PROGRESS 0x00000001 /* */ #define NV_PPBDMA_STATUS_SCHED_CHSW 15:15 /* */ #define NV_PPBDMA_STATUS_SCHED_CHSW_NOT_IN_PROGRESS 0x00000000 /* */ #define NV_PPBDMA_STATUS_SCHED_CHSW_IN_PROGRESS 0x00000001 /* */ // STATUS_INST and STATUS_INST_HI registers: // // These debug registers reflect the instance pointer associated with the // channel last loaded on the PBDMA. // // The channel's instance pointer contains the RAMFC esched state for the // channel, the MMU state for the channel such as the page directory base (PDB) // pointer, and a pointer to the engine context(s) associated with the channel. // See NV_RAMFC in dev_ram.ref for further information. // // The instance pointer information gets populated into these registers from // STATUS_NEXT_INST registers at the point when both the RAMFC load completes and // the channel instance bind acks return. This means it will report the previous // channel value during the NV_PPBDMA_STATUS_SCHED_CHAN_STATUS CHSW_SWITCH and // CHSW_SAVE states, and will report the previous channel's instance during the // CHSW_LOAD and INVALID states. It will report the current channel's value while // CHAN_STATUS is VALID. // // STATUS_INST_TARGET field: // // This field indicates the memory space aperature in which the channel's // instance block resides. // // STATUS_INST_PTR_LO field: // // This field contains the low 20 bits of the 4k-block granularity instance // pointer address. This field is aligned such that the low 32 bits of the // byte-addressed instance pointer can be obtained by simply masking off the other // fields. // // STATUS_INST_HI_PTR_HI field: // // The PTR_HI field in the register immediately following STATUS_INST contains // the upper 32 bits of the currently loaded channel's instance pointer. Because // this field is adjacent to the STATUS_INST register, SW can simply perform a 64 // bit read of STATUS_INST and mask off the VALID and TARGET fields to obtain the // full byte-addressed instance pointer. (Assuming, of course, a little-endian // system.) // // STATUS_INST_VALID field (deprecated): // // This field is deprecated and will be removed in a future architecture. See // bug 2455401. VALID can only ever report a snapshot of whether the INST pointer // fields correspond to a currently loaded channel. Because other registers will // have to be read to gather further relevant information about esched state, the // INST information can became stale. Consistent state across multiple registers // can only be guaranteed when the PBDMA is frozen, in which case there is no need // for a VALID field in this register: one can simply read the STATUS_SCHED // register for the CHAN_STATUS field for the relevant information when the PBDMA // is frozen. // // The behavior of the STATUS_INST VALID field matches NV_PPBDMA_CHANNEL_VALID // field - see the documentation there for the detailed description. // // These registers are maintained by Hardware and are available for debug // purposes. A set of these registers exists for each of Host's PBDMA units. // These registers are not context switched. #define NV_PPBDMA_STATUS_INST(i) (0x00040160+(i)*2048) /* R--4A */ #define NV_PPBDMA_STATUS_INST__SIZE_1 32 /* */ #define NV_PPBDMA_STATUS_INST_TARGET 1:0 /* R-XUF */ #define NV_PPBDMA_STATUS_INST_TARGET_VID_MEM 0x00000000 /* R---V */ #define NV_PPBDMA_STATUS_INST_TARGET_SYS_MEM_COHERENT 0x00000002 /* R---V */ #define NV_PPBDMA_STATUS_INST_TARGET_SYS_MEM_NONCOHERENT 0x00000003 /* R---V */ #define NV_PPBDMA_STATUS_INST_VALID 11:11 /* R-EVF */ #define NV_PPBDMA_STATUS_INST_VALID_FALSE 0x00000000 /* R-E-V */ #define NV_PPBDMA_STATUS_INST_VALID_TRUE 0x00000001 /* R---V */ #define NV_PPBDMA_STATUS_INST_PTR_LO 31:12 /* R-XUF */ #define NV_PPBDMA_STATUS_INST_HI(i) (0x00040164+(i)*2048) /* R--4A */ #define NV_PPBDMA_STATUS_INST_HI__SIZE_1 32 /* */ #define NV_PPBDMA_STATUS_INST_HI_PTR_HI 31:0 /* R-XUF */ // STATUS_NEXT_INST and STATUS_NEXT_INST_HI registers: // // These debug registers reflect the instance pointer associated with the // channel currently scheduled to be loaded on the PBDMA if there is one. When // VALID is TRUE, this pointer matches the pointer in the loading channel's // corresponding runlist entry. // // The channel's instance pointer contains the RAMFC esched state for the // channel, the MMU state for the channel such as the page directory base (PDB) // pointer, and a pointer to the engine context(s) associated with the channel. // See NV_RAMFC in dev_ram.ref for further information. // // STATUS_NEXT_INST_VALID field: // // When TRUE, a channel switch has been queued, and the INST pointer reported // in the TARGET, PTR_LO, and STATUS_NEXT_INST_HI PTR_HI fields corresponds to the // INST pointer of the channel that is scheduled to load on the PBDMA. When FALSE, // either no channel has yet loaded on the PBDMA, or the INST pointer has already // been loaded into the current STATUS_INST registers. // // STATUS_NEXT_INST_TARGET field: // // This field indicates the memory space aperature in which the channel's // instance block resides. // // STATUS_NEXT_INST_PTR_LO field: // // This field contains the low 20 bits of the 4k-block granularity instance // pointer address. This field is aligned such that the low 32 bits of the // byte-addressed instance pointer can be obtained by simply masking off the other // fields. // // STATUS_NEXT_INST_HI_PTR_HI field: // // The PTR_HI field in the register immediately following STATUS_INST contains // the upper 32 bits of the currently loaded channel's instance pointer. Because // this field is adjacent to the STATUS_INST register, SW can simply perform a 64 // bit read of STATUS_INST and mask off the VALID and TARGET fields to obtain the // full byte-addressed instance pointer. (Assuming, of course, a little-endian // system.) // // These registers are maintained by Hardware and are available for debug // purposes. A set of these registers exists for each of Host's PBDMA units. // These registers are not context switched. #define NV_PPBDMA_STATUS_NEXT_INST(i) (0x00040138+(i)*2048) /* R--4A */ #define NV_PPBDMA_STATUS_NEXT_INST__SIZE_1 32 /* */ #define NV_PPBDMA_STATUS_NEXT_INST_TARGET 1:0 /* R-XUF */ #define NV_PPBDMA_STATUS_NEXT_INST_TARGET_VID_MEM 0x00000000 /* R---V */ #define NV_PPBDMA_STATUS_NEXT_INST_TARGET_SYS_MEM_COHERENT 0x00000002 /* R---V */ #define NV_PPBDMA_STATUS_NEXT_INST_TARGET_SYS_MEM_NONCOHERENT 0x00000003 /* R---V */ #define NV_PPBDMA_STATUS_NEXT_INST_VALID 11:11 /* R-EVF */ #define NV_PPBDMA_STATUS_NEXT_INST_VALID_FALSE 0x00000000 /* R-E-V */ #define NV_PPBDMA_STATUS_NEXT_INST_VALID_TRUE 0x00000001 /* R---V */ #define NV_PPBDMA_STATUS_NEXT_INST_PTR_LO 31:12 /* R-XUF */ #define NV_PPBDMA_STATUS_NEXT_INST_HI(i) (0x0004013c+(i)*2048) /* R--4A */ #define NV_PPBDMA_STATUS_NEXT_INST_HI__SIZE_1 32 /* */ #define NV_PPBDMA_STATUS_NEXT_INST_HI_PTR_HI 31:0 /* R-XUF */ // STATUS_USERD and STATUS_USERD_HI registers: // // These debug registers reflect the USERD pointer associated with the // channel last loaded on the PBDMA. // // The channel's USERD pointer specifies the location of the channel's // usermode driver memory area. This area contains the GP_PUT pointer which the // driver writes to indicate how far into the GPFIFO work has been submitted. The // area also contains status information about the esched's progress on the // channel. See NV_RAMUSERD in dev_ram.ref for further information. // // The USERD pointer information gets populated into these registers from the // STATUS_NEXT_USERD registers at the point when the PBDMA decides to issue a bind // for the new channel. This means that in the NV_PPBDMA_STATUS_SCHED_CHAN_STATUS // CHSW_SAVE state, it will report the expected previous channel value. In // CHSW_SWITCH, it will switch to next channel value in the middle of the // CHSW_SWITCH state. Throughout the CHSW_LOAD state, it reports the next value. // It will be stable and report the current channel value during the CHAN_STATUS // VALID state. In the CHAN_STATUS INVALID state, it will report the USERD of the // channel last loaded on the PBDMA. // // STATUS_USERD_TARGET field: // // This field indicates the memory space aperture in which the channel's USERD // area resides. // // STATUS_USERD_PTR_LO field: // // This field contains the low 20 bits of the 4k-block granularity USERD // pointer address. This field is aligned such that the low 32 bits of the // byte-addressed USERD pointer can be obtained by simply masking off the other // fields. // // STATUS_USERD_HI_PTR_HI field: // // The PTR_HI field in the register immediately following STATUS_USERD contains // the upper 32 bits of the currently loaded channel's USERD pointer. Because // this field is adjacent to the STATUS_USERD register, SW can simply perform a 64 // bit read of STATUS_USERD and mask off the VALID and TARGET fields to obtain the // full byte-addressed USERD pointer. (Assuming, of course, a little-endian // system.) // // STATUS_USERD_VALID field (deprecated): // // This field is deprecated and will be removed in a future architecture. See // bug 2455401. VALID can only ever report a snapshot of whether the USERD pointer // fields correspond to a currently loaded channel. Because other registers will // have to be read to gather further relevant information about esched state, the // USERD information can became stale. Consistent state across multiple registers // can only be guaranteed when the PBDMA is frozen, in which case there is no need // for a VALID field in this register: one can simply read the STATUS_SCHED // register for the CHAN_STATUS field for the relevant information when the PBDMA // is frozen. // // The behavior of the STATUS_USERD VALID field matches NV_PPBDMA_CHANNEL_VALID // field - see the documentation there for the detailed description. // // These registers are maintained by Hardware and are available for debug // purposes. A set of these registers exists for each of Host's PBDMA units. // These registers are not context switched. #define NV_PPBDMA_STATUS_USERD(i) (0x00040168+(i)*2048) /* R--4A */ #define NV_PPBDMA_STATUS_USERD__SIZE_1 32 /* */ #define NV_PPBDMA_STATUS_USERD_TARGET 1:0 /* R--UF */ #define NV_PPBDMA_STATUS_USERD_TARGET_VID_MEM 0x00000000 /* R---V */ #define NV_PPBDMA_STATUS_USERD_TARGET_VID_MEM_NVLINK_COHERENT 0x00000001 /* R---V */ #define NV_PPBDMA_STATUS_USERD_TARGET_SYS_MEM_COHERENT 0x00000002 /* R---V */ #define NV_PPBDMA_STATUS_USERD_TARGET_SYS_MEM_NONCOHERENT 0x00000003 /* R---V */ #define NV_PPBDMA_STATUS_USERD_VALID 7:7 /* R-EVF */ #define NV_PPBDMA_STATUS_USERD_VALID_FALSE 0x00000000 /* R-E-V */ #define NV_PPBDMA_STATUS_USERD_VALID_TRUE 0x00000001 /* R---V */ #define NV_PPBDMA_STATUS_USERD_PTR_LO 31:8 /* R--UF */ #define NV_PPBDMA_STATUS_USERD_HI(i) (0x0004016c+(i)*2048) /* R--4A */ #define NV_PPBDMA_STATUS_USERD_HI__SIZE_1 32 /* */ #define NV_PPBDMA_STATUS_USERD_HI_PTR_HI 31:0 /* R--UF */ // // STATUS_NEXT_USERD and STATUS_NEXT_USERD_HI registers: // // These debug registers reflect the USERD pointer associated with the last // channel scheduled to load on the PBDMA. When VALID is TRUE, this pointer // matches the USERD pointer in the loading channel's corresponding runlist entry. // // The channel's USERD pointer specifies the location of the channel's // usermode driver memory area. This area contains the GP_PUT pointer which the // driver writes to indicate how far into the GPFIFO work has been submitted. The // area also contains status information about the esched's progress on the // channel. See NV_RAMUSERD in dev_ram.ref for further information. // // STATUS_NEXT_USERD_VALID field: // // When TRUE, a channel switch has been queued, and the USERD pointer reported // in the TARGET, PTR_LO, and STATUS_NEXT_USERD_HI PTR_HI fields corresponds to the // USERD pointer of the channel that is scheduled to load on the PBDMA. When // FALSE, either no channel has yet loaded on the PBDMA, or the USERD pointer has // already been loaded into the current STATUS_USERD registers. // // STATUS_NEXT_USERD_TARGET field: // // This field indicates the memory space aperature in which the channel's // USERD area resides. // // STATUS_NEXT_USERD_PTR_LO field: // // This field contains the low 20 bits of the 4k-block granularity USERD // pointer address. This field is aligned such that the low 32 bits of the // byte-addressed USERD pointer can be obtained by simply masking off the other // fields. // // STATUS_NEXT_USERD_HI_PTR_HI field: // // The PTR_HI field in the register immediately following STATUS_USERD contains // the upper 32 bits of the currently loaded channel's USERD pointer. Because // this field is adjacent to the STATUS_USERD register, SW can simply perform a 64 // bit read of STATUS_USERD and mask off the VALID and TARGET fields to obtain the // full byte-addressed USERD pointer. (Assuming, of course, a little-endian // system.) // // These registers are maintained by Hardware and are available for debug // purposes. A set of these registers exists for each of Host's PBDMA units. // These registers are not context switched. #define NV_PPBDMA_STATUS_NEXT_USERD(i) (0x00040140+(i)*2048) /* R--4A */ #define NV_PPBDMA_STATUS_NEXT_USERD__SIZE_1 32 /* */ #define NV_PPBDMA_STATUS_NEXT_USERD_TARGET 1:0 /* R--UF */ #define NV_PPBDMA_STATUS_NEXT_USERD_TARGET_VID_MEM 0x00000000 /* R---V */ #define NV_PPBDMA_STATUS_NEXT_USERD_TARGET_VID_MEM_NVLINK_COHERENT 0x00000001 /* R---V */ #define NV_PPBDMA_STATUS_NEXT_USERD_TARGET_SYS_MEM_COHERENT 0x00000002 /* R---V */ #define NV_PPBDMA_STATUS_NEXT_USERD_TARGET_SYS_MEM_NONCOHERENT 0x00000003 /* R---V */ #define NV_PPBDMA_STATUS_NEXT_USERD_VALID 7:7 /* R-EVF */ #define NV_PPBDMA_STATUS_NEXT_USERD_VALID_FALSE 0x00000000 /* R-E-V */ #define NV_PPBDMA_STATUS_NEXT_USERD_VALID_TRUE 0x00000001 /* R---V */ #define NV_PPBDMA_STATUS_NEXT_USERD_PTR_LO 31:8 /* R--UF */ #define NV_PPBDMA_STATUS_NEXT_USERD_HI(i) (0x00040144+(i)*2048) /* R--4A */ #define NV_PPBDMA_STATUS_NEXT_USERD_HI__SIZE_1 32 /* */ #define NV_PPBDMA_STATUS_NEXT_USERD_HI_PTR_HI 31:0 /* R--UF */ // CHANNEL register: // // This register contains the channel ID and GFID of the channel currently // loaded on the PBDMA. If the PBDMA has been preempted and no channel is loaded // or loading, the register contains information about the previously loaded // channel. // // // CHID field: // // CHID contains the esched-local system channel ID of the channel currently // loaded on the PBDMA, or the channel last loaded on the PBDMA during a channel // save. The chid gets populated with value from NEXT_CHANNEL_CHID as soon as the // RAMFC load completes. Note the esched does not wait for the channel's bind acks // to return. // // The update of the chid field corresponds with the STATUS_SCHED_CHAN_STATUS // field as follows: // // * CHAN_STATUS==VALID: CHID contains the current channel loaded on PBDMA // * CHAN_STATUS==INVALID or CHSW_SAVE: CHID specifies the channel that was last // loaded on the PBDMA if any. If no channel has been loaded on the PBDMA, the // value is unspecified. // * CHAN_STATUS==CHSW_LOAD or CHSW_SWITCH: prior to completing the RAMFC load, // CHID contains the prior channel loaded if any. After the RAMFC load // completes, CHID transitions to the channel ID reported in the // NEXT_CHANNEL_CHID register field. // // The CHID field identifies the RAMFC that is currently accessible via the // NV_PPBDMA registers. However, note that the RAMFC requires multiple cycles to // read, so it is possible that the PBDMA register state is not consistent during // channel load. // // // GFID field: // // The GFID associated with the channel identifier reported in the CHID field. // // // VALID field (deprecated): // // This field is deprecated and will be removed in a future architecture. See // bug 2455401. VALID can only ever report a snapshot of whether the CHID and GFID // fields correspond to a currently loaded channel. Because other registers will // have to be read to gather further relevant information about esched state, the // CHID information can became stale. Consistent state across multiple registers // can only be guaranteed when the PBDMA is frozen, in which case there is no need // for a VALID field in this register: one can simply read the STATUS_SCHED // register for the CHAN_STATUS field for the relevant information when the PBDMA // is frozen. // // VALID is set as follows based on the state of the channel switcher: // // * during CHSW_VALID after CHSW_LOAD and CHSW_SWITCH - VALID is set TRUE at the // point that all three of the following conditions are satisfied: the LB drain // completes, the RAMFC reads return, and the esched receives the bind ack from // MMU. This occurs some time after STATUS_SCHED_CHAN_STATUS switches from // CHSW_LOAD to VALID, and therefore after the CHID and GFID fields have already // been updated from NEXT_CHANNEL. // // * during CHSW_SAVE and CHSW_SWITCH - VALID is set FALSE immediately after the // RAMFC save writes are issued, without waiting for the write acks to return. // // * during CHSW_INVALID - VALID is FALSE as a consequence of having passed through // the CHSW_SAVE state. // // The VALID field roughly corresponds to whether or not the PBDMA registers // contain valid state: when TRUE, the registers are valid. When FALSE, the // registers may not contain valid state due to possible incoming RAMFC read data // from a channel load. // // More specifically, VALID==TRUE indicates that all sequences necessary to // begin processing the pushbuffer have completed. // // This register is maintained by Hardware and is available for debug // purposes. One of these registers exists for each of Host's PBDMA units. This // register is not context switched. #define NV_PPBDMA_CHANNEL(i) (0x00040120+(i)*2048) /* R--4A */ #define NV_PPBDMA_CHANNEL__SIZE_1 32 /* */ #define NV_PPBDMA_CHANNEL_CHID 11:0 /* */ #define NV_PPBDMA_CHANNEL_CHID_HW 10:0 /* R-XUF */ #define NV_PPBDMA_CHANNEL_VALID 13:13 /* R-IVF */ #define NV_PPBDMA_CHANNEL_VALID_FALSE 0x00000000 /* R-I-V */ #define NV_PPBDMA_CHANNEL_VALID_TRUE 0x00000001 /* R---V */ #define NV_PPBDMA_CHANNEL_GFID 21:16 /* R-XVF */ // NEXT_CHANNEL - Next chid and gfid // // The NV_PPBDMA_NEXT_CHANNEL register contains the channel id and gfid of the // channel that is scheduled to be loaded or is currently being loaded onto the // PBDMA. If the VALID field is FALSE, then this PBDMA unit is not scheduled to // switch or switching to a new channel. // This information is maintained by Hardware. This register is available for // debug purposes. // One of these registers exists for each of Host's PBDMA units. This // register is not context switched. This register runs on the internal-domain // clock. // // CHID field: // // The ID of the next channel that will be loaded on the PBDMA. Until another // channel switch is queued, this field will retain the ID of the last loaded // channel. // // GFID field: // // The GFID associated with the channel identifier reported in the CHID field. // // VALID field: // // The VALID field in NEXT_CHANNEL is TRUE during a channel switch while the // CHID field reports the ID of esched-local system channel ID of the channel being // loaded on the PBDMA. This field transitions to FALSE when the CHID is loaded // into the current CHID field in the NV_PPBDMA_CHANNEL register, which occurs after // the STATUS_SCHED_CHAN_STATUS field has transitioned from CHSW_LOAD or CHSW_SWITCH // to INVALID. // // This register is maintained by Hardware and is available for debug // purposes. One of these registers exists for each of Host's PBDMA units. This // register is not context switched. #define NV_PPBDMA_NEXT_CHANNEL(i) (0x00040124+(i)*2048) /* R--4A */ #define NV_PPBDMA_NEXT_CHANNEL__SIZE_1 32 /* */ #define NV_PPBDMA_NEXT_CHANNEL_CHID 11:0 /* */ #define NV_PPBDMA_NEXT_CHANNEL_CHID_HW 10:0 /* R-XUF */ #define NV_PPBDMA_NEXT_CHANNEL_VALID 13:13 /* R-IVF */ #define NV_PPBDMA_NEXT_CHANNEL_VALID_FALSE 0x00000000 /* R-I-V */ #define NV_PPBDMA_NEXT_CHANNEL_VALID_TRUE 0x00000001 /* R---V */ #define NV_PPBDMA_NEXT_CHANNEL_GFID 21:16 /* R-XVF */ // // GP_SHADOW_* - Last Received GP-Entry Header // // The NV_PPBDMA_GP_SHADOW_* registers contain the last GP entry that was // received by the PBDMA unit. This is the data at NV_PPBDMA_GP_GET-8. If the // PBDMA unit is indicating an invalid GP entry (NV_PPBDMA_INTR_0_GPENTRY), then // this register will contain that entry. // See Chapter 2: GPENTRY DATA FORMAT for a description of the GPFIFO format. // One of these registers exists for each of Host's PBDMA units. This // register is not context switched. This register runs on the internal-domain // clock. #define NV_PPBDMA_GP_SHADOW_0(i) (0x00040110+(i)*2048) /* RW-4A */ #define NV_PPBDMA_GP_SHADOW_0__SIZE_1 32 /* */ #define NV_PPBDMA_GP_SHADOW_0_VALUE 31:0 /* RWXUF */ #define NV_PPBDMA_GP_SHADOW_1(i) (0x00040114+(i)*2048) /* RW-4A */ #define NV_PPBDMA_GP_SHADOW_1__SIZE_1 32 /* */ #define NV_PPBDMA_GP_SHADOW_1_VALUE 31:0 /* RWXUF */ // // HDR_SHADOW - The Last Pushbuffer-Entry Header Processed // // The NV_PPBDMA_HDR_SHADOW register contains the raw PB instruction // corresponding to the information in NV_PPBDMA_PB_HEADER. If the PBDMA unit is // indicating an invalid PB entry (NV_PPBDMA_INTR_0_PBENTRY), then this register // or NV_PPBDMA_PB_DATA0 will contain the raw data for that entry which triggers // the interrupt. // One of these registers exists for each of Host's PBDMA units. This // register is not context switched. This register runs on the internal-domain // clock. // NOTE: - the NV_PPBDMA_HDR_SHADOW register may not always be // reliable when the PBENTRY interrupt fires due to a address wrap error. Because // this register is not stored in RAMFC, if the channel switches out after the // header is already stored in this register it will not be restored when the // channel switches back. So, if the address wrap error occurs after the switch // back, NV_PPBDMA_HDR_SHADOW will contain the incorrect data. SW should not rely // on this register to always be accurate. #define NV_PPBDMA_HDR_SHADOW(i) (0x00040118+(i)*2048) /* RW-4A */ #define NV_PPBDMA_HDR_SHADOW__SIZE_1 32 /* */ #define NV_PPBDMA_HDR_SHADOW_VALUE 31:0 /* RWXUF */ // MEM_OP_* [registers] - Memory-Operation Operand Backing Registers // // The NV_PPBDMA_MEM_OP_* registers contain bits 95:0 of the operands // to a memory management operation. Memory management operations are // triggered by NV_UDMA_MEM_OP_D methods; see NV_UDMA_MEM_OP* below for the method // documentation. // This register is part of a GPU context's state. On a switch, the value of // these registers are saved to, and restored from, the NV_RAMFC_MEM_OP_A, // NV_RAMFC_MEM_OP_B, and NV_RAMFC_MEM_OP_C fields of the RAMFC part of the GPU // context's GPU-instance block. // Software uses NV_UDMA_MEM_OP_* methods to alter this information. // Typically, software does not access this register directly. This register is // available to software only for debug. Software should use this register only if // the GPU context is assigned to a PBDMA unit and that PBDMA unit is stalled. // While a GPU context's Host state is not contained within a PBDMA unit, software // should use NV_RAMFC_MEM_OP_C to access this information. // One of these registers exists for each of Host's PBDMA units. This // register runs on Host's internal domain clock. These registers were added // and/or moved for Pascal (MEM_OP_A used to exist at offsets 400a0 + // i*NV_HOST_NV_PPBDMA_STRIDE). #define NV_PPBDMA_MEM_OP_A(i) (0x00040004+(i)*2048) /* RW-4A */ #define NV_PPBDMA_MEM_OP_A__SIZE_1 32 /* */ #define NV_PPBDMA_MEM_OP_A_DATA 31:0 /* RWXUF */ #define NV_PPBDMA_MEM_OP_B(i) (0x00040064+(i)*2048) /* RW-4A */ #define NV_PPBDMA_MEM_OP_B__SIZE_1 32 /* */ #define NV_PPBDMA_MEM_OP_B_DATA 31:0 /* RWXUF */ #define NV_PPBDMA_MEM_OP_C(i) (0x000400a0+(i)*2048) /* RW-4A */ #define NV_PPBDMA_MEM_OP_C__SIZE_1 32 /* */ #define NV_PPBDMA_MEM_OP_C_DATA 31:0 /* RWXUF */ // SIGNATURE - RAMFC Signature Register // // This register contains a value that specifies which Host class ID software // expects the hardware to support, and indicates if the RAMFC might be valid. It // is intended for debug and as a runtime check that RM is exposing the proper Host // class ID for the chip. // When the RAMFC part of a GPU context's instance block is restored into // Host, if the HW field does not contain the class ID specified by // HW_HOST_CLASS_ID or the value HW_VALID, then Host will freeze and initiate an // NV_PPBDMA_INTR_*_SIGNATURE interrupt. Host's class ID can be queried at runtime // from NV_CTRL_ESCHED_CONFIG_CLASS_ID; see dev_ctrl.ref. Note the Host class is also // known as "channel_gpfifo". HW_VALID (0xface) is meant to be used by RM to ease // transitions between Host classes for new architectures. The HW field does not // provide a direct check for Host methods sent by a given user mode driver; // attempting to send methods from a mismatching Host class may or may not work // depending on the method. // The SW field is for use by software. Host is not affected by the value. // This register is part of a GPU context's state. On a switch, the value of // this register is saved to, and restored from, the NV_RAMFC_SIGNATURE field of // the RAMFC part of the GPU context's GPU-instance block. // One of these registers exists for each of Host's PBDMA units. This // register runs on Host's internal domain clock. This register was added for // Fermi. #define NV_PPBDMA_SIGNATURE(i) (0x00040010+(i)*2048) /* RW-4A */ #define NV_PPBDMA_SIGNATURE__SIZE_1 32 /* */ #define NV_PPBDMA_SIGNATURE_HW 15:0 /* RWXUF */ #define NV_PPBDMA_SIGNATURE_HW_VALID 0x0000face /* RW--V */ #define NV_PPBDMA_SIGNATURE_HW_HOST_CLASS_ID 50543 /* RW--V */ #define NV_PPBDMA_SIGNATURE_SW 31:16 /* RWXUF */ #define NV_PPBDMA_SIGNATURE_SW_ZERO 0x00000000 /* RW--V */ // CONFIG - Miscellaneous Configuration Register // // The CONFIG register is used to configure miscellaneous functions of a PBDMA on a // per-channel basis. Software can configure these bits via the corresponding // NV_RAMFC_CONFIG dword in each channel's RAMFC. // The L2_EVICT field controls the l2_class field for memory requests from a PBDMA unit. // The CE_SPLIT field controls Host taking large copies and splitting them into smaller // copies to allow fast Copy Engine (CE) switching. If the field value is ENABLE, Host will analyze // each copy command to determine if the copy should be split into smaller copies, and may // modify the commands sent to the CE. // // // If the field value is DISABLE, Host will not modify the copy commands sent to the CE. // If the field is written from ENABLE to DISABLE while Host is in the middle of splitting a copy, // Host will continue splitting the current copy until the whole copy has been split. Future // copies, however, will not be split while the field remains set to DISABLE. // The THROTTLE_MODE field controls how much work Host sends to the CE. The // goal is to send enough work to keep the Copy Engine busy while Host switches // away to another channel to check on a semaphore, while at the same time // maintaining the CE preemption latency below 10 microseconds. When the field is // set to THROTTLE, Host will limit the number of copies it sends to the CE. This // is legacy behavior and is needed on PCIE GEN3 systems. Setting the field to // NO_THROTTLE will prevent Host from limiting the amount of work that Host sends // to the CE. NVLINK2 and PCIE GEN4_LITE systems should have the field set to // NO_THROTTLE. // Note: Because this is a static setting, if a system slowdown occurs and the link // is downgraded, preemption latency may exceed 10 microseconds. // The AUTH_LEVEL field specifies the authorization level of the channel. // When AUTH_LEVEL is NON_PRIVILEGED, the channel will not be able to execute // privileged operations via Host methods on its pushbuffer. Any attempt to do so // will result in the NV_PPBDMA_INTR_*_METHOD interrupt being raised. When // AUTH_LEVEL is PRIVILEGED, the channel will be able to execute all methods. // The following PBDMA methods are considered privileged for execution and // require AUTH_LEVEL_PRIVILEGED: // * MEM_OP_D methods with one of the following operations: // * MMU_TLB_INVALIDATE // * ACCESS_COUNTER_CLR // // The USERD_WRITEBACK field controls whether USERD will be written back to // memory. Regardless of the setting here, USERD is always written back to memory // when the channel switches off of the PBDMA. When USERD_WRITEBACK is ENABLE, // USERD will also be written back to memory whenever the PBDMA falls idle or the // writeback timer configured via NV_RUNLIST_USERD_WRITEBACK_TIMER expires. When the // field value is DISABLE, the writeback only occurs on channel save. Note GP_PUT // does not get written back to memory because it is written by software; // otherwise, GP_PUT updates could be lost on writeback. // This register is part of a GPU context's state. On a switch, the value of // this register is saved to, and restored from, the NV_RAMFC_CONFIG field of the RAMFC // part of the GPU context's GPU-instance block. // // // One of these registers exists for each of Host's PBDMA units. This // register runs on Host's internal domain clock. This register was added for // Fermi. #define NV_PPBDMA_CONFIG(i) (0x000400f4+(i)*2048) /* R--4A */ #define NV_PPBDMA_CONFIG__SIZE_1 32 /* */ #define NV_PPBDMA_CONFIG_L2_EVICT 0:0 /* R--VF */ #define NV_PPBDMA_CONFIG_L2_EVICT_FIRST 0x00000000 /* R---V */ #define NV_PPBDMA_CONFIG_L2_EVICT_NORMAL 0x00000001 /* R---V */ #define NV_PPBDMA_CONFIG_CE_SPLIT 4:4 /* R--VF */ #define NV_PPBDMA_CONFIG_CE_SPLIT_ENABLE 0x00000000 /* R---V */ #define NV_PPBDMA_CONFIG_CE_SPLIT_DISABLE 0x00000001 /* R---V */ #define NV_PPBDMA_CONFIG_CE_THROTTLE_MODE 5:5 /* R--VF */ #define NV_PPBDMA_CONFIG_CE_THROTTLE_MODE_THROTTLE 0x00000000 /* R---V */ #define NV_PPBDMA_CONFIG_CE_THROTTLE_MODE_NO_THROTTLE 0x00000001 /* R---V */ #define NV_PPBDMA_CONFIG_AUTH_LEVEL 8:8 /* R--VF */ #define NV_PPBDMA_CONFIG_AUTH_LEVEL_NON_PRIVILEGED 0x00000000 /* R---V */ #define NV_PPBDMA_CONFIG_AUTH_LEVEL_PRIVILEGED 0x00000001 /* R---V */ #define NV_PPBDMA_CONFIG_USERD_WRITEBACK 12:12 /* R--VF */ #define NV_PPBDMA_CONFIG_USERD_WRITEBACK_DISABLE 0x00000000 /* R---V */ #define NV_PPBDMA_CONFIG_USERD_WRITEBACK_ENABLE 0x00000001 /* R---V */ // INTR_NOTIFY // // INTR_NOTIFY is used to set the interrupt vector the PBDMA will use // when executing an NV_UDMA_NON_STALL_INT method. This vector is part of the // RAMFC image of the channel, and each channel can have a different vector #define NV_PPBDMA_INTR_NOTIFY(i) (0x000400f8+(i)*2048) /* RW-4A */ #define NV_PPBDMA_INTR_NOTIFY__SIZE_1 32 /* */ #define NV_PPBDMA_INTR_NOTIFY_VECTOR 11:0 /* RWXUF */ #define NV_PPBDMA_INTR_NOTIFY_CTRL_GSP 30:30 /* RWXUF */ #define NV_PPBDMA_INTR_NOTIFY_CTRL_GSP_DISABLE 0 /* R---V */ #define NV_PPBDMA_INTR_NOTIFY_CTRL_GSP_ENABLE 1 /* R---V */ #define NV_PPBDMA_INTR_NOTIFY_CTRL_CPU 31:31 /* RWXUF */ #define NV_PPBDMA_INTR_NOTIFY_CTRL_CPU_DISABLE 0 /* R---V */ #define NV_PPBDMA_INTR_NOTIFY_CTRL_CPU_ENABLE 1 /* R---V */ // After a channel switch, the first method Host will send to the graphics or // copy engine is a NV_PMETHOD_SET_CHANNEL_INFO method. The lower 16 bits of the // payload of this method (defined in internal_methods.ref) will consist of the // lower 16 bit value from this register. The upper 16 bits of the payload will // be populated by Host with the channel ID. // The lower 16 bits of the value of this method is expected to be set in // RAMFC by writing 32 bits to the offset specified as NV_RAMFC_SET_CHANNEL_INFO // at channel allocation. When generating the method, Host will ignore the upper // 16 bits of the register value and populate the upper 16 bits of the method // payload with the channel ID. The register value should only change if the // channel's TSG is preempted and the channel is not loaded on a PBDMA. // The VEID field is used to specify the Virtual Engine ID (VEID) for the // channel. A VEID is a collection of independent compute or graphics state which // shares execution resources and a context image. Each channel in a TSG can be // for a different VEID, any channels sharing a VEID will share WFI behavior. // The RESERVED field is reserved for Host and any value written in these // upper 16 bits by SW is ignored by Host when generating the internal method // NV_PMETHOD_SET_CHANNEL_INFO. // The SET_CHANNEL_INFO data should be set in RAMFC via the // NV_RAMFC_SET_CHANNEL_INFO entry rather than through this register. // This register is part of a GPU context's state. On a switch, the value of // this register is saved to and restored from the NV_RAMFC_SET_CHANNEL_INFO // field of the RAMFC part of the GPU context's GPU-instance block. // One of these registers exists for each of Host's PBDMA units. This // register runs on Host's internal domain clock. #define NV_PPBDMA_SET_CHANNEL_INFO(i) (0x000400fc+(i)*2048) /* RW-4A */ #define NV_PPBDMA_SET_CHANNEL_INFO__SIZE_1 32 /* */ #define NV_PPBDMA_SET_CHANNEL_INFO_VALUE 31:0 /* RWXUF */ #define NV_PPBDMA_SET_CHANNEL_INFO_SCG_TYPE 0:0 /* */ #define NV_PPBDMA_SET_CHANNEL_INFO_SCG_TYPE_GRAPHICS_COMPUTE0 0x00000000 /* */ #define NV_PPBDMA_SET_CHANNEL_INFO_SCG_TYPE_COMPUTE1 0x00000001 /* */ #define NV_PPBDMA_SET_CHANNEL_INFO_VEID ((6-1)+8):8 /* */ #define NV_PPBDMA_SET_CHANNEL_INFO_CHID 27:16 /* */ #define NV_PPBDMA_SET_CHANNEL_INFO_RESERVED 31:27 /* */ // HCI_CTRL - Misc Additional HCE State // HCE_CTRL is used for misc. HCE state that needs to be channel swapped // in addition to the normal CE CLASS state. // Some of the state bits are part of the MP/SP blocks' interactions with the // HCE Handling logic. // SP_AWAITS_HCEH indicates that the SP block is waiting for HCEH to finish // processing an HCE trigger method. // HCE_RENDER_DISABLED indicates that CE class rendering has been turned off. // HCE_SUBCHSW indicates that methods have been sent to HCE, and thus GR // will need to flush its caches when the next GR method in this channel // flows down to GR (indicated by interface bit). // HCE_PRIV_MODE indicates that physical launchDMA copies are allowed. // NOP_RCVD indicates that HCE logic has decoded a NOP method, and will // send the NOP to CE when permitted.(see launch_dma_rcvd description) // LAUNCH_DMA_RCVD indicates that the HCE logic has decoded a launchdma // method from MP, and it will be sent to CE when CE has returned enough // credits, and other criteria are met. // PM_TRIGGER_RCVD indicates that HCE logic has decoded a pm_trigger method // and wants to send it to CE. // SET_RENDER_ENABLE_C_RCVD indicates that HCE logic has decoded a // set_render_enable method, and is in the process of updating the render enable // state for CE. Note, this is not strictly necessary as channel state, but it // is useful for debug while the channel is loaded. #define NV_PPBDMA_HCE_CTRL(i) (0x000400e4+(i)*2048) /* RW-4A */ #define NV_PPBDMA_HCE_CTRL__SIZE_1 32 /* */ #define NV_PPBDMA_HCE_CTRL_SP_AWAITS_HCEH 0:0 /* RWXUF */ #define NV_PPBDMA_HCE_CTRL_SP_AWAITS_HCEH_NO 0x00000000 /* RW--V */ #define NV_PPBDMA_HCE_CTRL_SP_AWAITS_HCEH_YES 0x00000001 /* RW--V */ #define NV_PPBDMA_HCE_CTRL_HCE_RENDER_DISABLED 2:2 /* RWXUF */ #define NV_PPBDMA_HCE_CTRL_HCE_RENDER_DISABLED_NO 0x00000000 /* RW--V */ #define NV_PPBDMA_HCE_CTRL_HCE_RENDER_DISABLED_YES 0x00000001 /* RW--V */ #define NV_PPBDMA_HCE_CTRL_HCE_SUBCHSW 4:4 /* RWXUF */ #define NV_PPBDMA_HCE_CTRL_HCE_SUBCHSW_NO 0x00000000 /* RW--V */ #define NV_PPBDMA_HCE_CTRL_HCE_SUBCHSW_YES 0x00000001 /* RW--V */ #define NV_PPBDMA_HCE_CTRL_HCE_PRIV_MODE 5:5 /* RWXUF */ #define NV_PPBDMA_HCE_CTRL_HCE_PRIV_MODE_NO 0x00000000 /* RW--V */ #define NV_PPBDMA_HCE_CTRL_HCE_PRIV_MODE_YES 0x00000001 /* RW--V */ #define NV_PPBDMA_HCE_CTRL_LAUNCH_DMA_RCVD 16:16 /* RWXUF */ #define NV_PPBDMA_HCE_CTRL_LAUNCH_DMA_RCVD_NO 0x00000000 /* RW--V */ #define NV_PPBDMA_HCE_CTRL_LAUNCH_DMA_RCVD_YES 0x00000001 /* RW--V */ #define NV_PPBDMA_HCE_CTRL_NOP_RCVD 17:17 /* RWXUF */ #define NV_PPBDMA_HCE_CTRL_NOP_RCVD_NO 0x00000000 /* RW--V */ #define NV_PPBDMA_HCE_CTRL_NOP_RCVD_YES 0x00000001 /* RW--V */ #define NV_PPBDMA_HCE_CTRL_PM_TRIGGER_RCVD 18:18 /* RWXUF */ #define NV_PPBDMA_HCE_CTRL_PM_TRIGGER_RCVD_NO 0x00000000 /* RW--V */ #define NV_PPBDMA_HCE_CTRL_PM_TRIGGER_RCVD_YES 0x00000001 /* RW--V */ #define NV_PPBDMA_HCE_CTRL_PM_TRIGGER_END_RCVD 19:19 /* RWXUF */ #define NV_PPBDMA_HCE_CTRL_PM_TRIGGER_END_RCVD_NO 0x00000000 /* RW--V */ #define NV_PPBDMA_HCE_CTRL_PM_TRIGGER_END_RCVD_YES 0x00000001 /* RW--V */ #define NV_PPBDMA_HCE_CTRL_SET_RENDER_ENABLE_C_RCVD 20:20 /* RWXUF */ #define NV_PPBDMA_HCE_CTRL_SET_RENDER_ENABLE_C_RCVD_NO 0x00000000 /* RW--V */ #define NV_PPBDMA_HCE_CTRL_SET_RENDER_ENABLE_C_RCVD_YES 0x00000001 /* RW--V */ // 4 - NV_PPBDMA INTERRUPT REGISTERS // =================================== // // The interrupt registers NV_PPBDMA_INTR_* contain and control the interrupt // state for each PBDMA. Interrupts are set by events and are cleared by software // running on the CPU or GSP. // // Interrupts in the PBDMA are divided into two interrupt trees: // // NV_RUNLIST_INTR_0_PBDMAn_INTR_TREE_0 NV_RUNLIST_INTR_0_PBDMAn_INTR_TREE_1 // | | // ______^______ ______^______ // / \ / \ // | OR | | OR | // '_______________' '_______________' // ||||||| | | ||||||| // other tree0 | | other tree1 // ANDed intr bits ^ ^ ANDed intr bits // AND AND // | | | | // _______. .______ _______. .________ // / \ / \ // | \ / | // PPBDMA_INTR_0/1_EN_SET_TREE(p,0)_intr Y PPBDMA_INTR_0/1_EN_SET_TREE(p,1)_intr // | // NV_PPBDMA_INTR_0/1_intr_bit // // ^ // Note this 0/1 indicates the 1st or 2nd // intr register, each containing up to 32 // interrupt bits. INTR_2 will be defined // if needed in the future. // // where "n" in the NV_RUNLIST_INTR_0_PBDMAn_INTR_TREE_* interrupt bits is the // per-runlist PBDMA id of the PBDMA, and "p" refers to the global PBDMA ID (an // implementation detail due to the arrayed structure of the registers in // NV_PPBDMA). Since a runlist can serve at most NV_HOST_MAX_NUM_PBDMAS_PER_RUNLIST, // and there are NV_HOST_NUM_INTR_TREES_PER_RUNLIST interrupt trees, // the NV_RUNLIST_INTR_0 register will have // NV_HOST_MAX_NUM_PBDMAS_PER_RUNLIST * NV_HOST_NUM_INTR_TREES_PER_PBDMA such bits. // // The NV_PPBDMA_INTR_0/1_EN_SET/CLEAR_TREE register array determines which PBDMA // interrupt bits will route to which tree for each PBDMA. For each interrupt // tree, the separate interrupt bits in INTR_0 are first anded with their // corresponding bits in the INTR_0_EN_SET_TREE(pbdma_id, tree) register. Similarly, // the bits in INTR_1 are anded with those in the INTR_1_EN_SET_TREE(pbdma_id, tree) // register. Those bits are then rolled up in a reduction OR to determine whether // the corresponding NV_RUNLIST_INTR_0_PBDMAn_INTR_TREE_tree bit will be set or // not - the resulting bits are level driven by the rolled-up PBDMA intr bits. // // It is expected that SW will configure the next higher enables in // NV_RUNLIST_INTR_0_EN_SET/CLEAR_TREE(0) such that its // NV_RUNLIST_INTR_0_PBDMAn_INTR_TREE_0 bits will propagate through the runlist's // tree 0 vector specified in NV_RUNLIST_INTR_VECTORID(0), and similarly for // tree 1: that is, SW should set NV_RUNLIST_INTR_0_EN_SET/CLEAR_TREE(1) such that // its NV_RUNLIST_INTR_0_PBDMAn_INTR_TREE_1 bits propagate to // NV_RUNLIST_INTR_VECTORID(1). // // Note this diagram is effectively a continuation below the diagram in // Chapter 5 - Interrupt Registers in dev_runlist.ref. Simply consider each of the // NV_RUNLIST_INTR_0_PBDMAn_INTR_TREE_t outputs of the NV_PPBDMA_INTR diagram as // corresponding to a NV_RUNLIST_INTR_0_intr_bit at the bottom of the // NV_RUNLIST_INTR diagram. // INTR_0 - PBDMA Unit Interrupt Register // // The NV_PPBDMA_INTR_* registers are a PBDMA unit's interrupt register. The // logical-OR of this register feeds into the NV_RUNLIST_INTR_* register. If a bit // in this register is PENDING, then the corresponding interrupt condition has // occurred, and software has not yet indicated to hardware that the exception has // been handled. If a field is NON_PENDING then there are no exceptions of the // corresponding type that have not been handled. Software writes RESET to one of // these fields to indicate that a pending interrupt has been handled. // Software cannot set bits in this register. Attempting to write a bit to a // one actually clears the interrupt source. In this way, software can clear // individual bits in this register. When software recognizes an interrupt, and // services it, it can then clear the individual source by writing that single bit // in this register to RESET. Then it can read the register and see if all bits // are clear. If not, it can service other interrupts in this reg. This is // especially important since some of these bits are asynchronous to others in this // register. While an interrupt service routine (ISR) is clearing an interrupt, // other interrupts may occur. // Interrupts differ in severity. Some interrupts (like software interrupts) // are expected in the normal operation of of the GPU, and do not indicate that any // GPU context has been damaged, or hung. Some interrupts (like timeouts) do not // indicate damage, but indicate that deadlock might have occurred. Some interrupts // indicate that an error has occurred that might have damaged a GPU context, but // has not damaged any of the others. Finally some interrupts indicate that any // or all of the active GPU contexts have been damaged. // This register is for interrupts that cause a PBDMA unit to stall // (non-stalling non-switching interrupts are stored on a per-channel bias) Bits in // this register being set to PENDING will prevent the contents of the PBDMA unit // from being switched out. Until software handles these interrupts and writes the // bits to RESET, the PBDMA will be frozen. // One of these registers exists for each of Host's PBDMA units. This // register is not context switched. This register runs on Host's internal domain // clock. This register is new for Fermi. // // Interrupt field summary for INTR_0 and INTR_0_EN_*: #define NV_PPBDMA_INTR_0(i) (0x00040108+(i)*2048) /* RW-4A */ #define NV_PPBDMA_INTR_0__SIZE_1 32 /* */ // The NV_PPBDMA_INTR_*_GPFIFO field indicates that a PBDMA unit encountered // an invalid GPFIFO (circular buffer of GP-Entries). A GPFIFO that crosses the // end of the memory address space (0xFFFFFFFFFF) is invalid. The invalid value // will be in NV_PPBDMA_GP_BASE and NV_PPBDMA_GP_BASE_HI. Fixing this and clearing // the interrupt will allow the PBDMA unit to continue. The error is limited to // the channel. #define NV_PPBDMA_INTR_0_GPFIFO 13:13 /* RWIUF */ #define NV_PPBDMA_INTR_0_GPFIFO_NOT_PENDING 0x00000000 /* R-I-V */ #define NV_PPBDMA_INTR_0_GPFIFO_PENDING 0x00000001 /* R---V */ #define NV_PPBDMA_INTR_0_GPFIFO_RESET 0x00000001 /* -W--C */ // The NV_PPBDMA_INTR_*_GPPTR field indicated that a PBDMA unit encountered // invalid GP pointers (either NV_PPBDMA_GP_PUT, NV_PPBDMA_GP_FETCH, or // NV_PPBDMA_GP_GET). These pointers are invalid if they are not between zero and // one less than the size of the circular buffer that contains GP entries: // 1<= PV), // where SV is the semaphore value in memory, PV is the payload value, and >= is // an unsigned greater-than-or-equal-to comparison. // If OPERATION is ACQ_CIRC_GEQ, the acquire succeeds when the two's // complement signed representation of the semaphore value minus the payload value // is non-negative; that is, when the semaphore value is within half a range // greater than or equal to the payload value, modulo that range. The // PAYLOAD_SIZE field determines if Host is doing a 32 bit comparison or a 64 bit // comparison. So in other words, the condition is met when the PAYLOAD_SIZE is // 32BIT and the semaphore value is within the range [payload, // ((payload+(2^(32-1)))-1)], modulo 2^32, or when the PAYLOAD_SIZE is 64BIT and // the semaphore value is within the range [payload, ((payload+(2^(64-1)))-1)], // modulo 2^64. // If OPERATION is ACQ_AND, the acquire succeeds when the bitwise-AND of the // semaphore value and the payload value is not zero. The PAYLOAD_SIZE field // determines if a 32 bit or 64 bit value is read from memory, and compared to. // If OPERATION is ACQ_NOR, the acquire succeeds when the bitwise-NOR of the // semaphore value and the payload value is not zero. PAYLOAD_SIZE determines if // a 32 bit or 64 bit value is read from memory, and compared to. // If OPERATION is RELEASE, then Host simply writes the payload value to the // semaphore structure in memory at the SEM_ADDR_LO/_HI address. The exact value // written depends on the operation defined. If PAYLOAD_SIZE is 32BIT then a 32 // bit payload value from PAYLOAD_LO is used. If PAYLOAD_SIZE is 64BIT then a 64 // bit payload specified by PAYLOAD_LO/_HI is used. // If OPERATION is REDUCTION, then Host sends the memory system an instruction // to perform the atomic reduction operation specified in the REDUCTION field on // the memory value, using the PAYLOAD_LO/_HI payload value as the operand. The // OPERATION_PAYLOAD_SIZE field determines if a 32 bit or 64 bit reduction is // performed. Note that if the semaphore address refers to a page whose PTE has // ATOMIC_DISABLE set, the operation will result in an ATOMIC_VIOLATION fault; see // dev_mmu.ref. // Note that if the PAYLOAD_SIZE is 64BIT, the semaphore address is required // to be 8-byte aligned. If RELEASE_TIMESTAMP is EN while the operation is a // RELEASE or REDUCTION operation, the semaphore address is required to be 16-byte // aligned. The semaphore address is not required to be 16-byte aligned during an // acquire operation. If the semaphore address is not aligned according to the // field values Host will raise the NV_PBDMA_INTR_0_SEMAPHORE interrupt. // // Semaphore switch option: // // The NV_UDMA_SEM_EXECUTE_ACQUIRE_SWITCH_TSG field specifies whether or not // Host should switch to processing another TSG as soon as the acquire fails. If a // channel within a given TSG is waiting on a semaphore acquire and all other // channels in that TSG have no work (each channel either is waiting on a semaphore // acquire, is idle, is unbound, or is disabled), the entire TSG can make no // further progress until one of the relevant semaphores is released. Because it // may be a long time before the release, it may be more efficient for the PBDMA // unit to switch off the blocked TSG prior to the runqueue timeslice expiring, so // that it can serve a different TSG that is not waiting, or so that it can poll // other semaphores on other TSGs whose channels are waiting on acquires (for // similar reasons to have a kernel process yield when it likely will accomplish // nothing by remaining scheduled). // When a semaphore acquire fails, the PBDMA unit will always switch to // another channel within the same TSG, provided that it has not already completed // a traversal through all the TSG's channels. If every pending channel in the TSG // is waiting on a semaphore acquire, the Host scheduler is able identify a lack // of progress for the entire TSG by the time it has completed a traversal through // all of its channels. In this case the value of ACQUIRE_SWITCH_TSG for each of // these channels determines whether the PBDMA will switch to another TSG or start // another traversal through the same TSG. // If ACQUIRE_SWITCH_TSG is DIS for any of the pending channels in the TSG, // the Host scheduler will ignore any lack of progress and continue processing the // TSG, until either every channel in the TSG runs out of work or the timeslice // expires. // If ACQUIRE_SWITCH_TSG is EN for every pending channel in the TSG, the Host // scheduler will recognize a lack of progress for the whole TSG, and will switch // to the next serviceable TSG on the runqueue, if possible. // In the case described above, if there isn't a different serviceable TSG // on the runlist, then the current channel's TSG will continue to be scheduled // and the acquire retry will be naturally delayed by the time it takes for Host's // runlist processing to return to the same channel. This retry delay may be too // short, in which case the runlist search can be throttled to increase the delay // by configuring NV_PFIFO_ACQ_PRETEST; see dev_fifo.ref. Note that if the // channel remains switched in, the prefetched pushbuffer data is not discarded, // so setting ACQUIRE_SWITCH_TSG_EN cannot deterministically be depended on to // cause the discarding of prefetched pushbuffer data. // Also note that when switching between channels within a TSG, Host does not // wait on any timer (such as NV_PFIFO_ACQ_PRETEST or NV_PBDMA_ACQUIRE_RETRY), // but is instead throttled by the time it takes to switch channels. Host will // honor the ACQUIRE_RETRY time, but only if the same channel is rescheduled // without a channel switch. // If, at the end of the traversal of any given runlist, every schedulable TSG // in the runlist has switched away due to all pending channels in each TSG having // made no forward progress on a failed a semaphore acquire with ACQUIRE_SWITCH_TSG // enabled (meaning the semaphore acquire had failed the previous time the channel // was scheduled), Host will fire the NV_PFIFO_INTR_0_RUNLIST_ACQUIRE interrupt for // that runlist. // // Semaphore wait-for-idle option: // // The NV_UDMA_SEM_EXECUTE_RELEASE_WFI field applies only to releases and // reductions. It specifies whether Host should wait until the engine to which // the channel last sent methods is idle (in other words, until all previous // methods in the channel have been completed) before writing to memory as part of // the release or reduction operation. If this field is RELEASE_WFI_EN, then Host // waits for the engine to be idle, inserts a SYSMEMBAR (system memory barrier), // and then updates the value in memory. If this field is RELEASE_WFI_DIS, Host // performs the semaphore operation on the memory without waiting for the engine to // be idle, and without using a system memory barrier. // // Semaphore timestamp option: // // The NV_UDMA_SEM_EXECUTE_RELEASE_TIMESTAMP specifies whether a timestamp // should be written by a release in addition to the payload. If // RELEASE_TIMESTAMP is DIS, then only the semaphore payload will be written. If // the field is EN then both the semaphore payload and a nanosecond timestamp will // be written. In this case, the semaphore address must be 16-byte aligned; see // the related note at NV_UDMA_SEM_ADDR_LO. If RELEASE_TIMESTAMP is EN and // SEM_ADDR_LO is not 16-byte aligned, then Host will fire the // NV_PBDMA_INTR_0_SEMAPHORE interrupt. When a 16-byte semaphore is written, the // semaphore timestamp will be written before the semaphore payload so that when // an acquire succeeds, the timestamp write will have completed. This ensures SW // will not get an out-of-date timestamp on platforms which guarantee ordering // within a 16-byte aligned region. The timestamp value is snapped from the // NV_PTIMER_TIME_1/0 registers; see dev_timer.ref. // // Below is the little endian format of 16-byte semaphores in memory: // // ---- ------------------- ------------------- // byte Data(Little endian) Data(Little endian) // PAYLOAD_SIZE=32BIT PAYLOAD_SIZE=64BIT // ---- ------------------- ------------------- // 0 Payload[ 7: 0] Payload[ 7: 0] // 1 Payload[15: 8] Payload[15: 8] // 2 Payload[23:16] Payload[23:16] // 3 Payload[31:24] Payload[31:24] // 4 0 Payload[39:32] // 5 0 Payload[47:40] // 6 0 Payload[55:48] // 7 0 Payload[63:56] // 8 timer[ 7: 0] timer[ 7: 0] // 9 timer[15: 8] timer[15: 8] // 10 timer[23:16] timer[23:16] // 11 timer[31:24] timer[31:24] // 12 timer[39:32] timer[39:32] // 13 timer[47:40] timer[47:40] // 14 timer[55:48] timer[55:48] // 15 timer[63:56] timer[63:56] // ---- ------------------- ------------------- // // // Semaphore reduction operations: // // The NV_UDMA_SEM_EXECUTE_REDUCTION field specifies the reduction operation // to perform on the semaphore memory value, using the semaphore payload from // SEM_PAYLOAD_LO/HI as an operand, when the OPERATION field is // OPERATION_REDUCTION. Based on the PAYLOAD_SIZE field the semaphore value and // the payload are interpreted as 32bit or 64bit integers and the reduction // operation is performed according to the signedness specified via the // REDUCTION_FORMAT field described below. The reduction operation leaves the // modified value in the semaphore memory according to the operation as follows: // // REDUCTION_IMIN - the minimum of the value and payload // REDUCTION_IMAX - the maximum of the value and payload // REDUCTION_IXOR - the bitwise exclusive or (XOR) of the value and payload // REDUCTION_IAND - the bitwise AND of the value and payload // REDUCTION_IOR - bitwise OR of the value and payload // REDUCTION_IADD - the sum of the value and payload // REDUCTION_INC - the value incremented by 1, or reset to 0 if the incremented // value would exceed the payload // REDUCTION_DEC - the value decremented by 1, or reset back to the payload // if the original value is already 0 or exceeds the payload // // Note that INC and DEC are somewhat surprising: they can be used to repeatedly // loop the semaphore value when performed successively with the same payload p. // INC repeatedly iterates from 0 to p inclusive, resetting to 0 once exceeding p. // DEC repeatedly iterates down from p to 0 inclusive, resetting back to p once // the value would otherwise underflow. Therefore, an INC or DEC reduction with // payload 0 effectively releases a semaphore by setting its value to 0. // // The reduction opcode assignment matches the enumeration in the XBAR translator // (to avoid extra remapping of hardware), but this does not match the graphics FE // reduction opcodes used by graphics backend semaphores. The reduction operation // itself is performed by L2. // // Semaphore signedness option: // // The NV_UDMA_SEM_EXECUTE_REDUCTION_FORMAT field specifies whether the // values involved in a reduction operation will be interpreted as signed or // unsigned. // // The following table summarizes each reduction operation, and the signedness and // payload size supported for each operation: // // signedness // r op 32b 64b function (v = memory value, p = semaphore payload) // -----+-----+-----+--------------------------------------------------- // IMIN U,S U,S v = (v < p) ? v : p // IMAX U,S U,S v = (v > p) ? v : p // IXOR N/A N/A v = v ^ p // IAND N/A N/A v = v & p // IOR N/A N/A v = v | p // IADD U,S U v = v + p // INC U inv v = (v >= p) ? 0 : v + 1 // DEC U inv v = (v == 0 || v > p) ? p : v - 1 (from L2 IAS) // // An operation with signedness "N/A" will ignore the value of REDUCTION_FORMAT // when executing, and either value of REDUCTION_FORMAT is valid. If an operation // is "U only" this means a signed version of this operation is not supported, and // if it is marked "inv" then it is unsupported for any signedness. If Host sees // an unsupported reduction op (in other words, is expected to run a reduction op // while PAYLOAD_SIZE and REDUCTION_FORMAT are set to unsupported values for that // op), Host will fire the NV_PBDMA_INTR_0_SEMAPHORE interrupt. // // Example: A signed 32-bit IADD reduction operation is valid. A signed 64-bit // IADD reduction operation is unsupported and will trigger an interrupt if sent to // Host. A 64-bit INC (or DEC) operation is not supported and will trigger an // interrupt if sent to Host. // // Legal semaphore operation combinations: // // The following table diagrams the types of semaphore operations that are // possible. In the columns, "x" matches any field value. ACQ refers to any of // the ACQUIRE, ACQ_STRICT_GEQ, ACQ_CIRC_GEQ, ACQ_AND, and ACQ_NOR operations. // REL refers to either a RELEASE or a REDUCTION operation. // // OP SWITCH WFI PAYLOAD_SIZE TIMESTAMP Description // --- ------ --- ------------ --------- -------------------------------------------------------------- // ACQ 0 x 0 x acquire; 4B (32 bit comparison); retry on fail // ACQ 0 x 1 x acquire; 8B (64 bit comparison); retry on fail // ACQ 1 x 0 x acquire; 4B (32 bit comparison); switch on fail // ACQ 1 x 1 x acquire; 8B (64 bit comparison); switch on fail // REL x 0 0 1 WFI & release 4B payload + timestamp semaphore // REL x 0 1 1 WFI & release 8B payload + timestamp semaphore // REL x 1 0 1 do not WFI & release 4B payload + timestamp semaphore // REL x 1 1 1 do not WFI & release 8B payload + timestamp semaphore // REL x 0 0 0 WFI & release doubleword (4B) semaphore payload // REL x 0 1 0 WFI & release quadword (8B) semaphore payload // REL x 1 0 0 do not WFI & release doubleword (4B) semaphore payload // REL x 1 1 0 do not WFI & release quadword (8B) semaphore payload // --- ------ --- ------------ --------- -------------------------------------------------------------- // // While the channel is loaded on a PBDMA unit, information from this method // is stored in the NV_PBDMA_SEM_EXECUTE register. Otherwise, this information // is stored in the NV_RAMFC_SEM_EXECUTE field of the RAMFC part of the channel's // instance block. // // Undefined bits: // // Bits in the NV_UDMA_SEM_EXECUTE method data that are not used by the // specified OPERATION should be set to 0. When non-zero, their behavior is // undefined. #define NV_UDMA_SEM_EXECUTE 0x0000006C /* -W-4R */ #define NV_UDMA_SEM_EXECUTE_OPERATION 2:0 /* -W-VF */ #define NV_UDMA_SEM_EXECUTE_OPERATION_ACQUIRE 0x00000000 /* -W--V */ #define NV_UDMA_SEM_EXECUTE_OPERATION_RELEASE 0x00000001 /* -W--V */ #define NV_UDMA_SEM_EXECUTE_OPERATION_ACQ_STRICT_GEQ 0x00000002 /* -W--V */ #define NV_UDMA_SEM_EXECUTE_OPERATION_ACQ_CIRC_GEQ 0x00000003 /* -W--V */ #define NV_UDMA_SEM_EXECUTE_OPERATION_ACQ_AND 0x00000004 /* -W--V */ #define NV_UDMA_SEM_EXECUTE_OPERATION_ACQ_NOR 0x00000005 /* -W--V */ #define NV_UDMA_SEM_EXECUTE_OPERATION_REDUCTION 0x00000006 /* -W--V */ #define NV_UDMA_SEM_EXECUTE_ACQUIRE_SWITCH_TSG 12:12 /* -W-VF */ #define NV_UDMA_SEM_EXECUTE_ACQUIRE_SWITCH_TSG_DIS 0x00000000 /* -W--V */ #define NV_UDMA_SEM_EXECUTE_ACQUIRE_SWITCH_TSG_EN 0x00000001 /* -W--V */ #define NV_UDMA_SEM_EXECUTE_RELEASE_WFI 20:20 /* -W-VF */ #define NV_UDMA_SEM_EXECUTE_RELEASE_WFI_DIS 0x00000000 /* -W--V */ #define NV_UDMA_SEM_EXECUTE_RELEASE_WFI_EN 0x00000001 /* -W--V */ #define NV_UDMA_SEM_EXECUTE_PAYLOAD_SIZE 24:24 /* -W-VF */ #define NV_UDMA_SEM_EXECUTE_PAYLOAD_SIZE_32BIT 0x00000000 /* -W--V */ #define NV_UDMA_SEM_EXECUTE_PAYLOAD_SIZE_64BIT 0x00000001 /* -W--V */ #define NV_UDMA_SEM_EXECUTE_RELEASE_TIMESTAMP 25:25 /* -W-VF */ #define NV_UDMA_SEM_EXECUTE_RELEASE_TIMESTAMP_DIS 0x00000000 /* -W--V */ #define NV_UDMA_SEM_EXECUTE_RELEASE_TIMESTAMP_EN 0x00000001 /* -W--V */ #define NV_UDMA_SEM_EXECUTE_REDUCTION 30:27 /* -W-VF */ #define NV_UDMA_SEM_EXECUTE_REDUCTION_IMIN 0x00000000 /* -W--V */ #define NV_UDMA_SEM_EXECUTE_REDUCTION_IMAX 0x00000001 /* -W--V */ #define NV_UDMA_SEM_EXECUTE_REDUCTION_IXOR 0x00000002 /* -W--V */ #define NV_UDMA_SEM_EXECUTE_REDUCTION_IAND 0x00000003 /* -W--V */ #define NV_UDMA_SEM_EXECUTE_REDUCTION_IOR 0x00000004 /* -W--V */ #define NV_UDMA_SEM_EXECUTE_REDUCTION_IADD 0x00000005 /* -W--V */ #define NV_UDMA_SEM_EXECUTE_REDUCTION_INC 0x00000006 /* -W--V */ #define NV_UDMA_SEM_EXECUTE_REDUCTION_DEC 0x00000007 /* -W--V */ #define NV_UDMA_SEM_EXECUTE_REDUCTION_FORMAT 31:31 /* -W-VF */ #define NV_UDMA_SEM_EXECUTE_REDUCTION_FORMAT_SIGNED 0x00000000 /* -W--V */ #define NV_UDMA_SEM_EXECUTE_REDUCTION_FORMAT_UNSIGNED 0x00000001 /* -W--V */ // // NON_STALL_INT [method] - Non-Stalling Interrupt Method // // The NON_STALL_INT method causes an interrupt to be raised in the CTRL tree // specified by INTR_NOTIFY register. // Note that if the corresponding interrupt's enable bit is not enabled, the // interrupt will not propagate past the leaf interrupt register. This method does // not cause the PBDMA to stall the execution of the channel, switch out the // channel from the PBDMA, or disable channel switching for the PBDMA. // The NON_STALL_INT method's data (NV_UDMA_NON_STALL_INT_HANDLE) is ignored. // Software should handle all of a channel's non-stalling interrupts before it // unbinds the channel from the GPU context. #define NV_UDMA_NON_STALL_INT 0x00000020 /* -W-4R */ #define NV_UDMA_NON_STALL_INT_HANDLE 31:0 /* -W-VF */ // // // // MEM_OP methods: membars, and cache and TLB management. // // MEM_OP_A, MEM_OP_B, and MEM_OP_C set up state for performing a memory // operation. MEM_OP_D sets additional state, specifies the type of memory // operation to perform, and triggers sending the mem op to HUB. To avoid // unexpected behavior for future revisions of the MEM_OP methods, all 4 methods // should be sent for each requested mem op, with irrelevant fields set to 0. // Note that hardware does not enforce the requirement that unrelated fields be set // to 0, but ignoring this advice could break forward compatibility. // Host does not wait until an engine is idle before beginning to execute // this method. // While a GPU context is bound to a channel and assigned to a PBDMA unit, // the NV_UDMA_MEM_OP_A-C values are stored in the NV_PBDMA_MEM_OP_A-C registers // respectively. While the GPU context is not assigned to a PBDMA unit, these // values are stored in the respective NV_RAMFC_MEM_OP_A-C fields of the RAMFC part // of the GPU context's instance block in memory. // // Usage, operations, and configuration: // // MEM_OP_D_OPERATION specifies the type of memory operation to perform. This // field determines the value of the opcode on the Host/FB interface. When Host // encounters the MEM_OP_D method, Host sends the specified request to the FB and // waits for an indication that the request has completed before beginning to // process the next method. To issue a memory operation, first issue the 3 // MEM_OP_A-C methods to configure the operation as documented below. Then send // MEM_OP_D to complete the configuration and trigger the operation. The // operations available for MEM_OP_D_OPERATION are as follows: // MEMBAR - perform a memory barrier; see below. // MMU_TLB_INVALIDATE - invalidate page translation and attribute data from // the given page directory that are cached in the Memory-Management Unit TLBs. // MMU_TLB_INVALIDATE_TARGETED - invalidate page translation and attributes // data corresponding to a specific page in a given page directory. // L2_SYSMEM_INVALIDATE - invalidate data from system memory cached in L2. // L2_PEERMEM_INVALIDATE - invalidate peer-to-peer data in the L2 cache. // L2_CLEAN_COMPTAGS - clean the L2 compression tag cache. // L2_FLUSH_DIRTY - flush dirty lines from L2. // L2_WAIT_FOR_SYS_PENDING_READS - ensure all sysmem reads are past the point // of being modified by a write through a reflected mapping. To do this, L2 drains // all sysmem reads to the point where they cannot be modified by future // non-blocking writes to reflected sysmem. L2 will block any new sysmem read // requests and drain out all read responses. Note VC's with sysmem read requests // at the head would stall any request till the flush is complete. The niso-nb vc // does not have sysmem read requests so it would continue to flow. L2 will ack // that the sys flush is complete and unblock all VC's. Note this operation is a // NOP on tegra chips. // ACCESS_COUNTER_CLR - clear page access counters. // // Depending on the operation given in MEM_OP_D_OPERATION, the other fields of // all four MEM_OP methods are interpreted differently: // // MMU_TLB_INVALIDATE* // ------------------- // // When the operation is MMU_TLB_INVALIDATE or MMU_TLB_INVALIDATE_TARGETED, // then Host will initiate a TLB invalidate as described above. The MEM_OP // configuration fields specify what to invalidate, where to perform the // invalidate, and optionally trigger a replay or cancel event for replayable // faults buffered within the TLBs as part of UVM page management. // When the operation is MMU_TLB_INVALIDATE_TARGETED, // MEM_OP_C_TLB_INVALIDATE_PDB must be ONE, and the TLB_INVALIDATE_TARGET_ADDR_LO // and HI fields must be filled in to specify the target page. // To control whether to invalidate link TLB only, or non-link TLB only or all // TLBs, INVAL_SCOPE field (aliased with cancel_gpc_id) should be used with // MEM_OP_C_TLB_INVALIDATE_REPLAY set to NONE. // These operations are privileged and can only be executed from channels // with NV_PBDMA_CONFIG_AUTH_LEVEL set to PRIVILEGED. This is configured via the // NV_RAMFC_CONFIG dword in the channel's RAMFC during channel setup. // // MEM_OP_A_TLB_INVALIDATE_CANCEL_TARGET_GPC_ID and // MEM_OP_A_TLB_INVALIDATE_CANCEL_TARGET_CLIENT_UNIT_ID identify the GPC and uTLB // within that GPC respectively that should perform the cancel operation when // MEM_OP_C_TLB_INVALIDATE_REPLAY is CANCEL_TARGETED. These field values should be // copied from the GPC_ID and CLIENT fields from the associated // NV_UVM_FAULT_BUF_ENTRY packet or NV_PFB_PRI_MMU_FAULT_BUFFER_* entry. The // CLIENT_UNIT_ID corresponds to the values specified by NV_PFAULT_CLIENT_GPC_* in // dev_fault.ref. These fields are used with the CANCEL_TARGETED operation. The // fields also overlap with CANCEL_MMU_ENGINE_ID, and are interpreted as // CANCEL_MMU_ENGINE_ID during reply of type REPLAY_CANCEL_VA_GLOBAL. For other // replay operations, these fields must be 0. // // MEM_OP_A_TLB_INVALIDATE_CANCEL_MMU_ENGINE_ID specifies the associated // MMU_ENGINE_ID of the requests targeted by a REPLAY_CANCEL_VA_GLOBAL // operation. The field is ignored if the replay operation is not // REPLAY_CANCEL_VA_GLOBAL. This field overlaps with CANCEL_TARGET_GPC_ID and // CANCEL_TARGET_CLIENT_UNIT_ID field. // // MEM_OP_A_TLB_INVALIDATE_INVALIDATION_SIZE is aliased/repurposed // with MEM_OP_A_TLB_INVALIDATE_CANCEL_TARGET_CLIENT_UNIT_ID field // when MEM_OP_C_TLB_INVALIDATE_REPLAY (below) is anything other // than CANCEL_TARGETED or CANCEL_VA_GLOBAL or // CANCEL_VA_TARGETED. In the invalidation size enabled replay type // cases, actual region to be invalidated is calculated as // 4K*(2^INVALIDATION_SIZE) i.e., // 4K*(2^CANCEL_TARGET_CLIENT_UNIT_ID); client unit id and gpc id // are not applicable. // // MEM_OP_A_TLB_INVALIDATE_SYSMEMBAR controls whether a Hub SYSMEMBAR // operation is performed after waiting for all outstanding acks to complete, after // the TLB is invalidated. Note if ACK_TYPE is ACK_TYPE_NONE then this field is // ignored and no MEMBAR will be performed. This is provided as a SW optimization // so that SW does not need to perform a NV_UDMA_MEM_OP_D_OPERATION_MEMBAR op with // MEMBAR_TYPE SYS_MEMBAR after the TLB_INVALIDATE. This field must be 0 if // TLB_INVALIDATE_GPC is DISABLE. // // MEM_OP_B_TLB_INVALIDATE_TARGET_ADDR_HI:MEM_OP_A_TLB_INVALIDATE_TARGET_ADDR_LO // specifies the 4k aligned virtual address of the page whose translation to // invalidate within the TLBs. These fields are valid only when OPERATION is // MMU_TLB_INVALIDATE_TARGETED; otherwise, they must be set to 0. // // MEM_OP_C_TLB_INVALIDATE_PDB controls whether a TLB invalidate should apply // to a particular page directory or to all of them. If PDB is ALL, then all page // directories are invalidated. If PDB is ONE, then the PDB address and aperture // are specified in the PDB_ADDR_LO:PDB_ADDR_HI and PDB_APERTURE fields. // Note that ALL does not make sense when OPERATION is MMU_TLB_INVALIDATE_TARGETED; // the behavior in that case is undefined. [> [TODO,smueller,2014/04/07,31820709,TOREVIEW] or is it? <] // // MEM_OP_C_TLB_INVALIDATE_GPC controls whether the GPC-MMU and uTLB entries // should be invalidated in addition to the Hub-MMU TLB (Note: the Hub TLB is // always invalidated). Set it to INVALIDATE_GPC_ENABLE to invalidate the GPC TLBs. // The REPLAY, ACK_TYPE, and SYSMEMBAR fields are only used by the GPC TLB and so // are ignored if INVALIDATE_GPC is DISABLE. // // MEM_OP_C_TLB_INVALIDATE_REPLAY specifies the type of replay to perform in // addition to the invalidate. A replay causes all replayable faults outstanding // in the TLB to attempt their translations again. Once a TLB acks a replay, that // TLB may start accepting new translations again. The replay flavors are as // follows: // NONE - do not replay any replayable faults on invalidate. // START - initiate a replay across all TLBs, but don't wait for completion. // The replay will be acked as soon as the invalidate is processed, but // replays themselves are in flight and not necessarily translated. // START_ACK_ALL - initiate the replay and wait until it completes. // The replay will be acked after all pending transactions in the replay // fifo have been translated. New requests will remain stalled in the // gpcmmu until all transactions in the replay fifo have completed and // there are no pending faults left in the replay fifo. // CANCEL_TARGETED - initiate a cancel-replay on a targeted uTLB, causing any // replayable translations buffered in that uTLB to become non-replayable // if they fault again. In this case, the first faulting translation // will be reported in the NV_PFIFO_INTR_MMU_FAULT registers and will // raise PFIFO_INTR_0_MMU_FAULT. The specific TLB to target for the // cancel is specified in the CANCEL_TARGET fields. Note the TLB // invalidate still applies globally to all TLBs. // CANCEL_GLOBAL - like CANCEL_TARGETED, but all TLBs will cancel-replay. // CANCEL_VA_GLOBAL - initiates a cancel operation that cancels all requests // with the matching mmu_engine_id and access_type that land in the // specified 4KB aligned virtual address within the scope of specified // PDB. All other requests are replayed. If the specified engine is not // bound, or if the PDB of the specified engine does not match the // specified PDB, all requests will be replayed and none will be canceled. // // MEM_OP_C_TLB_INVALIDATE_ACK_TYPE controls which sort of ACK the uTLBs wait // for after having issued a membar to L2. ACK_TYPE_NONE does not perform any sort // of membar. ACK_TYPE_INTRANODE waits for an ack from the XBAR. // ACK_TYPE_GLOBALLY waits for an L2 ACK. ACK_TYPE_GLOBALLY is equivalent to a // MEMBAR operation from the engine, or a SYS_MEMBAR if // MEM_OP_A_TLB_INVALIDATE_SYSMEMBAR is EN. // // MEM_OP_C_TLB_INVALIDATE_PAGE_TABLE_LEVEL specifies which levels in the page // directory hierarchy of the TLB cache to invalidate. The levels are numbered // from the bottom up, with the PTE being at the bottom with level 1. The // specified level and all those below it in the hierarchy -- that is, all those // with a lower numbered level -- are invalidated. ALL (the 0 default) is // special-cased to indicate the top level; this causes the invalidate to apply to // the entire page mapping structure. The field is ignored if the replay operation // is REPLAY_CANCEL_VA_GLOBAL. // // MEM_OP_C_TLB_INVALIDATE_ACCESS_TYPE specifies the associated ACCESS_TYPE of // the requests targeted by a REPLAY_CANCEL_VA_GLOBAL operation. This field // overlaps with the INVALIDATE_PAGE_TABLE_LEVEL field, and is ignored if the // replay operation is not REPLAY_CANCEL_VA_GLOBAL. The ACCESS_TYPE field can get // one of the following values: // READ - the cancel_va_global should be performed on all pending read requests. // WRITE - the cancel_va_global should be performed on all pending write requests. // ATOMIC_STRONG - the cancel_va_global should be performed on all pending // strong atomic requests. // ATOMIC_WEAK - the cancel_va_global should be performed on all pending // weak atomic requests. // ATOMIC_ALL - the cancel_va_global should be performed on all pending atomic // requests. // WRITE_AND_ATOMIC - the cancel_va_global should be performed on all pending // write and atomic requests. // ALL - the cancel_va_global should be performed on all pending requests. // // // MEM_OP_C_TLB_INVALIDATE_PDB_APERTURE specifies the target aperture of the // page directory for which TLB entries should be invalidated. This field must be // 0 when TLB_INVALIDATE_PDB is ALL. // // MEM_OP_C_TLB_INVALIDATE_PDB_ADDR_LO specifies the low 20 bits of the // 4k-block-aligned PDB (base address of the page directory) when // TLB_INVALIDATE_PDB is ONE; otherwise this field must be 0. The PDB byte address // should be 4k aligned and right-shifted by 12 before being split and packed into // the ADDR fields. Note that the PDB_ADDR_LO field starts at bit 12, so it is // possible to set MEM_OP_C to the low 32 bits of the byte address, mask off the // low 12, and then or in the rest of the configuration fields. // // MEM_OP_D_TLB_INVALIDATE_PDB_ADDR_HI contains the high bits of the PDB when // TLB_INVALIDATE_PDB is ONE. Otherwise this field must be 0. // // UVM handling of replayable faults: // // The following example illustrates how TLB invalidate may be used by the // UVM driver: // 1. When the TLB invalidate completes, all memory accesses using the old // TLB entries prior to the invalidate will finish translation (but not // completion), and any new virtual accesses will trigger new // translations. The outstanding in-flight translations are allowed to // fault but will not indefinitely stall the invalidate. // 2. When the TLB invalidate completes, in-flight memory accesses using the // old physical translations may not yet be visible to other GPU clients // (such as CopyEngine) or to the CPU. Accesses coming from clients that // support recoverable faults (such as TEX and GCC) can be made visible by // requesting the MMU to perform a membar using the ACK_TYPE and SYSMEMBAR // fields. // a. If ACK_TYPE is NONE the SYSMEMBAR field is ignored and no membar // is performed. // b. If ACK_TYPE is INTRANODE the invalidate will wait until all // in-flight physical accesses using the old translations are visible // to XBAR clients on the blocking VC. // c. If ACK_TYPE is GLOBALLY the invalidate will wait until all // in-flight physical accesses using the old translations are at the // point of coherence in L2, meaning writes will be visible to all // other GPU clients and reads will not be mutable by them. // d. If the SYSMEMBAR field is set to EN then a Hub SYSMEMBAR will also // be performed following the ACK_TYPE membar. This is the equivalent // of performing a NV_UDMA_MEM_OP_C_MEMBAR_TYPE_SYS_MEMBAR. // 3. If fault replay was requested then all pending recoverable faults in // the TLB replay list will be retranslated. This includes all faults // discovered while the invalidate was pending. This replay may generate // more recoverable faults. // 4. If fault replay cancel was requested then another replay is attempted of // all pending replayable faults on the targeted TLB(s). If any of these // re-fault they are discarded (sticky NACK or ACK/TRAP sent back to the // client depending on the setting of NV_PGPC_PRI_MMU_DEBUG_CTRL). // // // // MEMBAR // ------ // // When the operation is MEMBAR, Host will perform a memory barrier operation. // All other fields must be set to 0 except for MEM_OP_C_MEMBAR_TYPE. When // MEMBAR_TYPE is MEMBAR, then a memory barrier will be performed with respect to // other clients on the GPU. When it is SYS_MEMBAR, the memory barrier will also be // performed with respect to the CPU and peer GPUs. // // MEMBAR - This issues a MEMBAR operation following all reads, writes, and // atomics currently in flight from the PBDMA. The MEMBAR operation will push all // such accesses already in flight on the same VC as the PBDMA to a point of GPU // coherence before proceeding. After this operation is complete, reads from any // GPU client will see prior writes from this PBDMA, and writes from any GPU client // cannot modify the return data of earlier reads from this PBDMA. This is true // regardless of whether those accesses target vidmem, sysmem, or peer mem. // WARNING: This only guarantees accesses from the same VC as the PBDMA that // are already in flight are coherent. Accesses from clients such as SM or a // non-PBDMA engine need already be at some point of coherency before this // operation to be coherent. // // SYS_MEMBAR - This implies the MEMBAR type above but in addition to having // accesses reach coherence with all GPU clients, this further waits for accesses // to be coherent with respect to the CPU and peer GPUs as well. After this // operation is complete, reads from the CPU or peer GPUs will see prior writes // from this PBDMA, and writes from the CPU or peer GPUs cannot modify the return // data of earlier reads from this PBDMA (with the exception of CPU reflected // writes, which can modify earlier reads). Note SYS_MEMBAR is really only needed // to guarantee ordering with off-chip clients. For on-chip clients such as the // graphics engine or copy engine, accesses to sysmem will be coherent with just a // MEMBAR operation. SYS_MEMBAR provides the same function as // OPERATION_SYSMEMBAR_FLUSH on previous architectures. // WARNING: As described above, SYS_MEMBAR will not prevent CPU reflected // writes issued after the SYS_MEMBAR from clobbering the return data of reads // issued before the SYS_MEMBAR. To handle this case, the invalidate must be // followed with a separate L2_WAIT_FOR_SYS_PENDING_READS mem op. // // // // L2* // --- // // These values initiate a cache management operation -- see above. All other // fields must be 0; there are no configuration options. // // // // // The ACCESS_COUNTER_CLR operation // -------------------------------- // When MEM_OP_D_OPERATION is ACCESS_COUNTER_CLR, Host will request to clear // the the page access counters. There are two types of access counters - MIMC and // MOMC. This operation can be issued to clear all counters of all types, all // counters of a specified type (MIMC or MOMC), or a specific counter indicated by // counter type, bank and notify tag. // This operation is privileged and can only be executed from channels with // NV_PBDMA_CONFIG_AUTH_LEVEL set to PRIVILEGED. This is configured via the // NV_RAMFC_CONFIG dword in the channel's RAMFC during channel setup. // // The operation uses the following fields in the MEM_OP_* methods: // ACCESS_COUNTER_CLR_TYPE (TY) : type of the access counter clear // operation // ACCESS_COUNTER_CLR_TARGETED_TYPE (T) : type of the access counter for // targeted operation // ACCESS_COUNTER_CLR_TARGETED_NOTIFY_TAG : 20 bits notify tag of the access // counter for targeted operation // ACCESS_COUNTER_CLR_TARGETED_BANK : 4 bits bank number of the access // counter for targeted operation // // // // // // MEM_OP method field defines: // // MEM_OP_A [method] - Memory Operation Method 1/4 - see above for documentation #define NV_UDMA_MEM_OP_A 0x00000028 /* -W-4R */ #define NV_UDMA_MEM_OP_A_TLB_INVALIDATE_CANCEL_TARGET_CLIENT_UNIT_ID 5:0 /* -W-VF */ #define NV_UDMA_MEM_OP_A_TLB_INVALIDATE_INVALIDATION_SIZE 5:0 /* -W-VF */ #define NV_UDMA_MEM_OP_A_TLB_INVALIDATE_CANCEL_TARGET_GPC_ID 10:6 /* -W-VF */ #define NV_UDMA_MEM_OP_A_TLB_INVALIDATE_INVAL_SCOPE 7:6 /* -W-VF */ #define NV_UDMA_MEM_OP_A_TLB_INVALIDATE_INVAL_SCOPE_ALL_TLBS 0 /* -W--V */ #define NV_UDMA_MEM_OP_A_TLB_INVALIDATE_INVAL_SCOPE_LINK_TLBS 1 /* -W--V */ #define NV_UDMA_MEM_OP_A_TLB_INVALIDATE_INVAL_SCOPE_NON_LINK_TLBS 2 /* -W--V */ #define NV_UDMA_MEM_OP_A_TLB_INVALIDATE_INVAL_SCOPE_RSVRVD 3 /* -W--V */ #define NV_UDMA_MEM_OP_A_TLB_INVALIDATE_CANCEL_MMU_ENGINE_ID 6:0 /* -W-VF */ #define NV_UDMA_MEM_OP_A_TLB_INVALIDATE_SYSMEMBAR 11:11 /* -W-VF */ #define NV_UDMA_MEM_OP_A_TLB_INVALIDATE_SYSMEMBAR_EN 0x00000001 /* -W--V */ #define NV_UDMA_MEM_OP_A_TLB_INVALIDATE_SYSMEMBAR_DIS 0x00000000 /* -W--V */ #define NV_UDMA_MEM_OP_A_TLB_INVALIDATE_TARGET_ADDR_LO 31:12 /* -W-VF */ // MEM_OP_B [method] - Memory Operation Method 2/4 - see above for documentation #define NV_UDMA_MEM_OP_B 0x0000002c /* -W-4R */ #define NV_UDMA_MEM_OP_B_TLB_INVALIDATE_TARGET_ADDR_HI 31:0 /* -W-VF */ // MEM_OP_C [method] - Memory Operation Method 3/4 - see above for documentation #define NV_UDMA_MEM_OP_C 0x00000030 /* -W-4R */ // Membar configuration field. Note: overlaps MMU_TLB_INVALIDATE* config fields. #define NV_UDMA_MEM_OP_C_MEMBAR_TYPE 2:0 /* -W-VF */ #define NV_UDMA_MEM_OP_C_MEMBAR_TYPE_SYS_MEMBAR 0x00000000 /* -W--V */ #define NV_UDMA_MEM_OP_C_MEMBAR_TYPE_MEMBAR 0x00000001 /* -W--V */ // Invalidate TLB entries for ONE page directory base, or for ALL of them. #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_PDB 0:0 /* -W-VF */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_PDB_ONE 0x00000000 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_PDB_ALL 0x00000001 /* -W--V */ // Invalidate GPC MMU TLB entries or not (Hub-MMU entries are always invalidated). #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_GPC 1:1 /* -W-VF */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_GPC_ENABLE 0x00000000 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_GPC_DISABLE 0x00000001 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_REPLAY 4:2 /* -W-VF */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_REPLAY_NONE 0x00000000 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_REPLAY_START 0x00000001 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_REPLAY_START_ACK_ALL 0x00000002 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_REPLAY_CANCEL_TARGETED 0x00000003 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_REPLAY_CANCEL_GLOBAL 0x00000004 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_REPLAY_CANCEL_VA_GLOBAL 0x00000005 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_ACK_TYPE 6:5 /* -W-VF */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_ACK_TYPE_NONE 0x00000000 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_ACK_TYPE_GLOBALLY 0x00000001 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_ACK_TYPE_INTRANODE 0x00000002 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_ACCESS_TYPE 9:7 /* -W-VF */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_ACCESS_TYPE_VIRT_READ 0 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_ACCESS_TYPE_VIRT_WRITE 1 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_ACCESS_TYPE_VIRT_ATOMIC_STRONG 2 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_ACCESS_TYPE_VIRT_RSVRVD 3 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_ACCESS_TYPE_VIRT_ATOMIC_WEAK 4 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_ACCESS_TYPE_VIRT_ATOMIC_ALL 5 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_ACCESS_TYPE_VIRT_WRITE_AND_ATOMIC 6 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_ACCESS_TYPE_VIRT_ALL 7 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_PAGE_TABLE_LEVEL 9:7 /* -W-VF */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_PAGE_TABLE_LEVEL_ALL 0x00000000 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_PAGE_TABLE_LEVEL_PTE_ONLY 0x00000001 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_PAGE_TABLE_LEVEL_UP_TO_PDE0 0x00000002 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_PAGE_TABLE_LEVEL_UP_TO_PDE1 0x00000003 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_PAGE_TABLE_LEVEL_UP_TO_PDE2 0x00000004 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_PAGE_TABLE_LEVEL_UP_TO_PDE3 0x00000005 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_PAGE_TABLE_LEVEL_UP_TO_PDE4 0x00000006 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_PAGE_TABLE_LEVEL_UP_TO_PDE5 0x00000007 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_PDB_APERTURE 11:10 /* -W-VF */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_PDB_APERTURE_VID_MEM 0x00000000 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_PDB_APERTURE_SYS_MEM_COHERENT 0x00000002 /* -W--V */ #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_PDB_APERTURE_SYS_MEM_NONCOHERENT 0x00000003 /* -W--V */ // Address[31:12] of page directory for which TLB entries should be invalidated. #define NV_UDMA_MEM_OP_C_TLB_INVALIDATE_PDB_ADDR_LO 31:12 /* -W-VF */ #define NV_UDMA_MEM_OP_C_ACCESS_COUNTER_CLR_TARGETED_NOTIFY_TAG 19:0 /* -W-VF */ // MEM_OP_D [method] - Memory Operation Method 4/4 - see above for documentation // (Must be preceded by MEM_OP_A-C.) #define NV_UDMA_MEM_OP_D 0x00000034 /* -W-4R */ // Address[58:32] of page directory for which TLB entries should be invalidated. #define NV_UDMA_MEM_OP_D_TLB_INVALIDATE_PDB_ADDR_HI 26:0 /* -W-VF */ #define NV_UDMA_MEM_OP_D_OPERATION 31:27 /* -W-VF */ #define NV_UDMA_MEM_OP_D_OPERATION_MEMBAR 0x00000005 /* -W--V */ #define NV_UDMA_MEM_OP_D_OPERATION_MMU_TLB_INVALIDATE 0x00000009 /* -W--V */ #define NV_UDMA_MEM_OP_D_OPERATION_MMU_TLB_INVALIDATE_TARGETED 0x0000000a /* -W--V */ #define NV_UDMA_MEM_OP_D_OPERATION_L2_PEERMEM_INVALIDATE 0x0000000d /* -W--V */ #define NV_UDMA_MEM_OP_D_OPERATION_L2_SYSMEM_INVALIDATE 0x0000000e /* -W--V */ #define NV_UDMA_MEM_OP_D_OPERATION_L2_CLEAN_COMPTAGS 0x0000000f /* -W--V */ #define NV_UDMA_MEM_OP_D_OPERATION_L2_FLUSH_DIRTY 0x00000010 /* -W--V */ #define NV_UDMA_MEM_OP_D_OPERATION_L2_WAIT_FOR_SYS_PENDING_READS 0x00000015 /* -W--V */ #define NV_UDMA_MEM_OP_D_OPERATION_ACCESS_COUNTER_CLR 0x00000016 /* -W--V */ #define NV_UDMA_MEM_OP_D_ACCESS_COUNTER_CLR_TYPE 1:0 /* -W-VF */ #define NV_UDMA_MEM_OP_D_ACCESS_COUNTER_CLR_TYPE_MIMC 0x00000000 /* -W--V */ #define NV_UDMA_MEM_OP_D_ACCESS_COUNTER_CLR_TYPE_MOMC 0x00000001 /* -W--V */ #define NV_UDMA_MEM_OP_D_ACCESS_COUNTER_CLR_TYPE_ALL 0x00000002 /* -W--V */ #define NV_UDMA_MEM_OP_D_ACCESS_COUNTER_CLR_TYPE_TARGETED 0x00000003 /* -W--V */ #define NV_UDMA_MEM_OP_D_ACCESS_COUNTER_CLR_TARGETED_TYPE 2:2 /* -W-VF */ #define NV_UDMA_MEM_OP_D_ACCESS_COUNTER_CLR_TARGETED_TYPE_MIMC 0x00000000 /* -W--V */ #define NV_UDMA_MEM_OP_D_ACCESS_COUNTER_CLR_TARGETED_TYPE_MOMC 0x00000001 /* -W--V */ #define NV_UDMA_MEM_OP_D_ACCESS_COUNTER_CLR_TARGETED_BANK 6:3 /* -W-VF */ // // SET_REF [method] - Set Reference Count Method // // The SET_REF method allows the user to set the reference count // (NV_PBDMA_REF_CNT) to a value. The reference count may be monitored to track // Host's progress through the pushbuffer. Instead of monitoring // NV_RAMUSERD_TOP_LEVEL_GET, software may put into the method stream SET_REF // methods that set the reference count to ever increasing values, and then read // NV_RAMUSERD_REF to determine how far in the stream Host has gone. // Before the reference count value is altered, Host waits for the engine to // be idle (to have completed executing all earlier methods), issues a SysMemBar // flush, and waits for the flush to complete. // While the GPU context is bound to a channel and assigned to a PBDMA unit, // the reference count value is stored in the NV_PBDMA_REF register. While the // GPU context is not assigned to a PBDMA unit, the reference count value is stored // in the NV_RAMFC_REF field of the RAMFC portion of the GPU context's GPU-instance // block. #define NV_UDMA_SET_REF 0x00000050 /* -W-4R */ #define NV_UDMA_SET_REF_CNT 31:0 /* -W-VF */ // // // // // YIELD [method] - Yield Method // // The YIELD method causes a channel to yield the remainder of its timeslice. // The method's OP field specifies whether the channels' PBDMA timeslice, the // channel's runlist timeslice, or no timeslice is yielded. // If YIELD_OP_RUNLIST_TIMESLICE, then Host will act as if the channel's // runlist or TSG timeslice expired. Host will exit the TSG and switch to the next // channel after the TSG on the runlist. If there is no such channel to switch to, // then YIELD_OP_RUNLIST_TIMESLICE will not cause a switch. // When the PBDMA executes a YIELD_OP_RUNLIST_TIMESLICE method, it guarantees // that it will not execute further methods from the same channel or TSG until the // channel is restarted by the scheduler. However, note that this does not yield // the engine timeslice; if the engine is preemptable, the context will continue // to run on the engine until the remainder of its timeslice expires before Host // will attempt to preempt it. Also if there is an outstanding ctx load either // due to ctx_reload or from the other PBDMA in the SCG case, then yielding won't // take place until the outstanding ctx load finishes or aborts due to a preempt. // When the ctx load does complete on the other PBDMA, it is possible for that // PBDMA to execute some small number of additional methods before the runlist // yield takes effect and that PBDMA halts work for its channel. // // Note: In case of two PBDMA serving the same engine, when one PBDMA // has finished loading its context on the engine and the second PBDMA has an // outstanding context load request to the same engine, if one of the PBDMAs // executes a RUNLIST YIELD method, the engine timeslice for engine is also yielded. // The implication being software while using RUNLIST YIELD method may notice that // the engine is sometimes not getting its full timeslice. Thus, there are times // when PBDMA is yielded and the engine also loses its timeslice whereas other times // only PBDMA is yielded and engine retains its timeslice. // // If NV_UDMA_YIELD_OP_TSG, and if the channel is part of a TSG, then Host // will switch to the next channel in the same TSG, and if the channel is not part // of the TSG then this will be treated similar to YIELD_OP_NOP. If there is only // one channel with work in the TSG, Host will simply reschedule the same channel // in the TSG. YIELD_OP_TSG does not cause the scheduler to leave the TSG. The TSG // timeslice (TSG timeslice is equivalent to runlist timeslice for TSGs) counter // continues to increment through the channel switch and does not restart after // executing the yield method. When the PBDMA executes a Yield method, it // guarantees that it will not execute the method following that Yield until the // channel is restarted by the scheduler. // YIELD_OP_NOP and OP_NOP1 are simply a NOP. Neither timeslice is yielded. // This was kept for compatibility with existing tests; NV_UDMA_NOP is the // preferred NOP, but also see the universal NOP PB instruction. See the // description of NV_FIFO_DMA_NOP in the "FIFO_DMA" section of dev_ram.ref. #define NV_UDMA_YIELD 0x00000080 /* -W-4R */ #define NV_UDMA_YIELD_OP 1:0 /* -W-VF */ #define NV_UDMA_YIELD_OP_NOP 0x00000000 /* -W--V */ #define NV_UDMA_YIELD_OP_NOP1 0x00000001 /* -W--V */ #define NV_UDMA_YIELD_OP_RUNLIST_TIMESLICE 0x00000002 /* -W--V */ #define NV_UDMA_YIELD_OP_TSG 0x00000003 /* -W--V */ // // WFI [method] - Wait-for-Idle Method // // The WFI (Wait-For-Idle) method will stall Host from processing any more // methods on the channel until the engine to which the channel last sent methods // is idle. Note that the subchannel encoded in the method header is ignored (as // it is for all Host-only methods) and does NOT specify which engine to idle. In // Kepler, this is only relevant on runlists that serve multiple engines // (specifically, the graphics runlist, which also serves GR COPY). // The WFI method has a single field SCOPE which specifies the level of WFI // the Host method performs. ALL waits for all work in the engine from the same // context to be idle across all classes and subchannels. CURRENT_VEID causes the // WFI to only apply to work from the same VEID as the current channel. Note for // engines that do not support VEIDs, CURRENT_VEID works identically to ALL. // Note that Host methods ignore the subchannel field in the method. A Host // WFI method always applies to the engine the channel last sent methods to. If a // WFI with ALL is specified and the channel last sent work to the GRCE, this will // only guarantee that GRCE has no work in progress. It is possible that the GR // context will have work in progress from other VEIDs, or even the current VEID if // the current channel targets GRCE and has never sent FE methods before. This // means that if SW wants to idle the graphics pipe for all VEIDs, SW must send a // method to GR immediately before the WFI method. A GR_NOP is sufficient. // Note also that even if the current NV_PBDMA_TARGET is GRAPHICS and not // GRCE, there are cases where Host can trivially complete a WFI without sending // the NV_PMETHOD_HOST_WFI internal method to FE. This can happen when // // 1. the runlist timeslices to a different TSG just before the WFI method, // 2. the other TSG does a ctxsw request due to methods for FE, and // 3. FECS reports non-preempted in the ctx ack, so CTX_RELOAD doesn't get set. // // In that case, when the channel switches back onto the PBDMA, the PBDMA rightly // concludes that there is no way the context could be non-idle for that channel, // and therefore filters out the WFI, even if the other PBDMA is sending work to // other VEIDs. As in the subchannel case, a GR_NOP preceding the WFI is // sufficient to ensure that a SCOPE_ALL_VEID WFI will be sent to FE regardless of // timeslicing as long as the NOP and the WFI are submitted as part of the same // GP_PUT update. This is ensured by the semantics of the channel state // SHOULD_SEND_HOST_TSG_EVENT behaving like CTX_RELOAD: the GR_NOP causes the PBDMA // to set the SHOULD_SEND_HOST_TSG_EVENT state, so even a channel or context switch // will still result in the PBDMA having the engine context loaded. Thus the WFI // will cause the HOST_WFI internal method to be sent to FE. #define NV_UDMA_WFI 0x00000078 /* -W-4R */ #define NV_UDMA_WFI_SCOPE 0:0 /* -W-VF */ #define NV_UDMA_WFI_SCOPE_CURRENT_VEID 0x00000000 /* -W--V */ #define NV_UDMA_WFI_SCOPE_ALL 0x00000001 /* -W--V */ #define NV_UDMA_WFI_SCOPE_ALL_VEID 0x00000001 /* */ // // // // // CLEAR_FAULTED [method] - Clear Faulted Method // // The CLEAR_FAULTED method acts like SW method and raises the METHOD interrupt // in INTR_0 register. The method data is ignored. SW may use this to initiate the // clearing the appropriate faulted bit in a faulted channel's NV_CHRAM_CHANNEL // entry. However, this is not the recommended way. // See the section "PBDMA_FAULTED and ENG_FAULTED common behavior" in // dev_runlist.ref #define NV_UDMA_CLEAR_FAULTED 0x00000084 /* -W-4R */ // SW should fill in this ID using the channel handle provided by RM. #define NV_UDMA_CLEAR_FAULTED_HANDLE 30:0 /* */ #define NV_UDMA_CLEAR_FAULTED_TYPE 31:31 /* -W-VF */ #define NV_UDMA_CLEAR_FAULTED_TYPE_PBDMA_FAULTED 0x00000000 /* -W--V */ #define NV_UDMA_CLEAR_FAULTED_TYPE_ENG_FAULTED 0x00000001 /* -W--V */ // Addresses that are not defined in this device are reserved. Those below // 0x100 are reserved for future Host methods. Addresses 0x100 and beyond are // reserved for the engines served by Host.