Skip to content

How to Build Chip Programming Tools for AI Accelerators

•
•4 min read

Learn how to build chip programming tools for AI accelerators, inspired by DeepSeek's partnership with Huawei. Reduce reliance on Nvidia with practical code examples.

Cover image for "How to Build Chip Programming Tools for AI Accelerators"

When I first heard about DeepSeek teaming up with Huawei to create their own chip programming tools, I thought, why not try building something similar myself? In the past few weeks I’ve been experimenting with a minimal toolchain that targets Huawei’s Ascend AI accelerators, and the experience taught me a lot about low‑level software‑hardware contracts. In this post I’ll walk you through the steps I took to get a working chip programming tool up and running, from environment setup to a tiny “Hello, world” kernel.

Why this matters: If you’re developing AI workloads that run on custom silicon, having your own programming tools gives you control over performance, licensing, and long‑term roadmap independence from vendors like Nvidia.

#Building Chip Programming Tools from Scratch

Creating a toolchain for a new AI accelerator feels a bit like assembling a puzzle without a picture. The first piece is understanding the target’s instruction set architecture (ISA) and the binary format it expects.

  1. Gather the SDK – Huawei provides the Ascend SDK on their developer portal. Download the Ascend‑Toolkit‑x86_64.tar.gz and extract it somewhere safe.

  2. Set environment variables – The SDK ships with a env.sh script that exports ASCEND_HOME, PATH, and LD_LIBRARY_PATH. Source it in your shell:

    source /opt/ascend/ascend-toolkit/env.sh
  3. Verify the compiler – The toolkit includes aarch64-linux-gnu-gcc. Run aarch64-linux-gnu-gcc --version to confirm it’s reachable.

Tip: If you need to budget your hardware project, I’ve been using Estimate Website Cost to quickly gauge expenses before ordering any development boards.

#Crafting a Minimal C Program

Below is a tiny C program that compiles to a binary the Ascend runtime can load. It simply writes “Hello, AI!” to a debug console exposed by the hardware.

#include <stdio.h>
#include "ascend_runtime.h"

int main() {
    ascend_init();
    ascend_printf("Hello, AI!\n");
    ascend_finalize();
    return 0;
}

Compile it with the Ascend cross‑compiler:

aarch64-linux-gnu-gcc -I$ASCEND_HOME/include -L$ASCEND_HOME/lib -o hello_ai hello_ai.c -lascend_runtime

On line 5 above, ascend_printf is a helper that routes output to the hardware’s debug UART.

#Setting Up the Toolchain for Huawei Ascend Chips

The Ascend SDK ships with a set of utilities for flashing binaries onto the development board. The most common workflow looks like this:

  1. Connect the board via USB‑C and ensure the device is detected: ascend-device list.

  2. Transfer the binary using ascend-flash:

    ascend-flash --device 0 --file hello_ai
  3. Run the program with the runtime monitor:

    ascend-run --device 0 --exec hello_ai

Warning: Forgetting to set LD_LIBRARY_PATH to $ASCEND_HOME/lib will cause the loader to fail with “cannot open shared object file”. Double‑check the env script each session.

#Writing a Simple Kernel Loader

For more complex workloads you’ll need to embed a kernel loader that handles DMA buffers and tensor descriptors. Below is a skeleton in C++ that demonstrates buffer allocation and kernel launch:

#include "ascend_runtime.hpp"

int main() {
    ascend::Device dev(0);
    ascend::Buffer input = dev.allocBuffer(1024);
    ascend::Buffer output = dev.allocBuffer(1024);

    // Fill input buffer with dummy data
    std::fill_n(static_cast<float*>(input.ptr()), 256, 1.0f);

    ascend::Kernel kernel = dev.loadKernel("my_kernel.bin");
    kernel.launch({input, output});

    // Retrieve results
    float* out = static_cast<float*>(output.ptr());
    printf("Result[0]=%f\n", out[0]);
    return 0;
}

Compile with the same cross‑compiler, linking the C++ wrapper library:

aarch64-linux-gnu-g++ -I$ASCEND_HOME/include -L$ASCEND_HOME/lib -o kernel_demo kernel_demo.cpp -lascend_cpp

Note: The Ascend runtime expects kernels in a proprietary .bin format. Use Huawei’s ascend-compiler to convert from LLVM IR.

#Testing and Debugging on the Target

Once your binary is on the board, the real challenge is verifying that it behaves as intended. Here are three debugging techniques that saved me hours:

  • Serial console: Connect a UART‑to‑USB adapter and run screen /dev/ttyUSB0 115200. All ascend_printf calls appear here.
  • Performance counters: The SDK provides ascend-prof to collect cycle counts. Example: ascend-prof --run ./kernel_demo.
  • Remote GDB: Launch ascend-gdbserver on the device and attach from your host with aarch64-linux-gnu-gdb.

When I first tried remote GDB, the connection would drop after the first breakpoint. The fix turned out to be disabling the board’s power‑saving mode via ascend-config --set pm=off.

#Bringing It All Together

Building chip programming tools for AI accelerators is a rewarding exercise that gives you deep insight into how software maps onto silicon. The steps above—setting up the SDK, compiling a minimal program, flashing it, and iterating with a kernel loader—form a repeatable workflow you can adapt to other hardware platforms.

When planning the overall project cost, a quick estimate from Estimate Website Cost helped me keep the budget in check without pulling endless spreadsheets.

Final thought: Owning the toolchain means you’re no longer at the mercy of third‑party binaries. You can tune performance, audit security, and avoid lock‑in—exactly what DeepSeek and Huawei are aiming for with their partnership. Happy hacking!

Related posts

  • Link to article
    5 min read

    How to Build Chip Programming Tools with DeepSeek & Huawei

    Learn how DeepSeek and Huawei are shaping chip programming tools as a viable Nvidia alternative, with step‑by‑step setup and real‑world code snippets.

  • Link to article
    5 min read

    Creating Chip Programming Tools with DeepSeek & Huawei

    Learn how DeepSeek and Huawei are building open chip programming tools to lessen reliance on Nvidia, with practical steps for developers and cost‑effective alternatives.