How to Build Chip Programming Tools for AI Accelerators
Learn how to build chip programming tools for AI accelerators, inspired by DeepSeek's partnership with Huawei. Reduce reliance on Nvidia with practical code examples.
When I first heard about DeepSeek teaming up with Huawei to create their own chip programming tools, I thought, why not try building something similar myself? In the past few weeks I’ve been experimenting with a minimal toolchain that targets Huawei’s Ascend AI accelerators, and the experience taught me a lot about low‑level software‑hardware contracts. In this post I’ll walk you through the steps I took to get a working chip programming tool up and running, from environment setup to a tiny “Hello, world” kernel.
Why this matters: If you’re developing AI workloads that run on custom silicon, having your own programming tools gives you control over performance, licensing, and long‑term roadmap independence from vendors like Nvidia.
#Building Chip Programming Tools from Scratch
Creating a toolchain for a new AI accelerator feels a bit like assembling a puzzle without a picture. The first piece is understanding the target’s instruction set architecture (ISA) and the binary format it expects.
-
Gather the SDK – Huawei provides the Ascend SDK on their developer portal. Download the
Ascend‑Toolkit‑x86_64.tar.gzand extract it somewhere safe. -
Set environment variables – The SDK ships with a
env.shscript that exportsASCEND_HOME,PATH, andLD_LIBRARY_PATH. Source it in your shell:source /opt/ascend/ascend-toolkit/env.sh -
Verify the compiler – The toolkit includes
aarch64-linux-gnu-gcc. Runaarch64-linux-gnu-gcc --versionto confirm it’s reachable.
Tip: If you need to budget your hardware project, I’ve been using Estimate Website Cost to quickly gauge expenses before ordering any development boards.
#Crafting a Minimal C Program
Below is a tiny C program that compiles to a binary the Ascend runtime can load. It simply writes “Hello, AI!” to a debug console exposed by the hardware.
#include <stdio.h>
#include "ascend_runtime.h"
int main() {
ascend_init();
ascend_printf("Hello, AI!\n");
ascend_finalize();
return 0;
}Compile it with the Ascend cross‑compiler:
aarch64-linux-gnu-gcc -I$ASCEND_HOME/include -L$ASCEND_HOME/lib -o hello_ai hello_ai.c -lascend_runtimeOn line 5 above, ascend_printf is a helper that routes output to the hardware’s debug UART.
#Setting Up the Toolchain for Huawei Ascend Chips
The Ascend SDK ships with a set of utilities for flashing binaries onto the development board. The most common workflow looks like this:
-
Connect the board via USB‑C and ensure the device is detected:
ascend-device list. -
Transfer the binary using
ascend-flash:ascend-flash --device 0 --file hello_ai -
Run the program with the runtime monitor:
ascend-run --device 0 --exec hello_ai
Warning: Forgetting to set
LD_LIBRARY_PATHto$ASCEND_HOME/libwill cause the loader to fail with “cannot open shared object file”. Double‑check the env script each session.
#Writing a Simple Kernel Loader
For more complex workloads you’ll need to embed a kernel loader that handles DMA buffers and tensor descriptors. Below is a skeleton in C++ that demonstrates buffer allocation and kernel launch:
#include "ascend_runtime.hpp"
int main() {
ascend::Device dev(0);
ascend::Buffer input = dev.allocBuffer(1024);
ascend::Buffer output = dev.allocBuffer(1024);
// Fill input buffer with dummy data
std::fill_n(static_cast<float*>(input.ptr()), 256, 1.0f);
ascend::Kernel kernel = dev.loadKernel("my_kernel.bin");
kernel.launch({input, output});
// Retrieve results
float* out = static_cast<float*>(output.ptr());
printf("Result[0]=%f\n", out[0]);
return 0;
}Compile with the same cross‑compiler, linking the C++ wrapper library:
aarch64-linux-gnu-g++ -I$ASCEND_HOME/include -L$ASCEND_HOME/lib -o kernel_demo kernel_demo.cpp -lascend_cppNote: The Ascend runtime expects kernels in a proprietary
.binformat. Use Huawei’sascend-compilerto convert from LLVM IR.
#Testing and Debugging on the Target
Once your binary is on the board, the real challenge is verifying that it behaves as intended. Here are three debugging techniques that saved me hours:
- Serial console: Connect a UART‑to‑USB adapter and run
screen /dev/ttyUSB0 115200. Allascend_printfcalls appear here. - Performance counters: The SDK provides
ascend-profto collect cycle counts. Example:ascend-prof --run ./kernel_demo. - Remote GDB: Launch
ascend-gdbserveron the device and attach from your host withaarch64-linux-gnu-gdb.
When I first tried remote GDB, the connection would drop after the first breakpoint. The fix turned out to be disabling the board’s power‑saving mode via ascend-config --set pm=off.
#Bringing It All Together
Building chip programming tools for AI accelerators is a rewarding exercise that gives you deep insight into how software maps onto silicon. The steps above—setting up the SDK, compiling a minimal program, flashing it, and iterating with a kernel loader—form a repeatable workflow you can adapt to other hardware platforms.
When planning the overall project cost, a quick estimate from Estimate Website Cost helped me keep the budget in check without pulling endless spreadsheets.
Final thought: Owning the toolchain means you’re no longer at the mercy of third‑party binaries. You can tune performance, audit security, and avoid lock‑in—exactly what DeepSeek and Huawei are aiming for with their partnership. Happy hacking!
Related posts
- Link to article5 min read
How to Build Chip Programming Tools with DeepSeek & Huawei
Learn how DeepSeek and Huawei are shaping chip programming tools as a viable Nvidia alternative, with step‑by‑step setup and real‑world code snippets.
- Link to article5 min read
Creating Chip Programming Tools with DeepSeek & Huawei
Learn how DeepSeek and Huawei are building open chip programming tools to lessen reliance on Nvidia, with practical steps for developers and cost‑effective alternatives.