Yanzhen Zhu

AI Compilers / VLSI EDA / Hardware and Software System Design

Full Stack Engineer for computing systems

Yanzhen Zhu

AI compiler, VLSI EDA algorithms, hardware and software design.

Full-stack engineer specializing in algorithm design and low-level implementation. Core competencies include AI compiler graph optimization and code generation, VLSI EDA placement and partitioning, FPGA acceleration design, circuit design, and embedded software development.

Projects

Research and engineering with measurable system outcomes.

Diffusion EDA Pybind11

Diffusion-Based Macro Placement

Infrastructure and model components for macro placement research.

  • Developed LEF/DEF and bookshelf backend parsers with Pybind11.
  • Used ILP workflows for automated dataset generation.
  • Designed a random-walk and cross-attention encoder for large heterogeneous netlist graphs.
  • Introduced a legality evaluation network to guide diffusion outputs with auxiliary loss.
Partitioning Timing STA

Netlist Partition with Cell Replication

Timing-driven optimization for multi-FPGA partitioning.

  • Proposed register-boundary replication to avoid cutting critical paths.
  • Built decision-tree bottleneck identification and overlap-aware path merging.
  • Combined STA-driven register pruning to preserve correctness while reducing replication cost.
  • Improved average TNS by 38.2%, max TNS by 94%, and runtime by 62x versus TritonPart.
3D-IC Placement Hypergraph

3D VLSI Macro Placement

Wirelength-driven 3D placement algorithm for macros.

  • Proposed an iterative optimization method for in-die macro placement.
  • Designed DFS and BFS heuristic search with pruning and quantization.
  • Implemented multithreaded hypergraph partitioning with anchor mapping.
  • Reduced hierarchical placement wirelength by 21.9% compared with NTUPlace.
Zynq-7020 FPGA Vision

Heterogeneous AMP Target Detector

Embedded vision acceleration platform on Zynq-7020.

  • Constructed a heterogeneous acceleration framework across Linux and bare-metal cores.
  • Implemented Sobel and Bayer-array hardware accelerators through AXI4.
  • Integrated image processing and guidance algorithms for target detection.
  • Reduced latency to 12.6% of the original baseline.

Publication

Work on MFSs partitioning, macro placement, and computational modeling.

Papers

  1. Yanzhen Zhu, et al., "MHEncoder: A Multi-Level Heterogeneous Encoder Framework for Scalable VLSI Netlist Representation," IEEE/ACM International Conference on Computer-Aided Design.
  2. Yanzhen Zhu, et al., "WIMPlace: An Wire Length Driven Tension Refine Based Macro Placer," Journal of Electronics & Information Technology.
  3. Coauthors, Yanzhen Zhu, et al., "An Improved Genetic Algorithm with Wavelet Packet and Low-Pass Filters for Reducing Pressure and Flow Pulsations in Axial Piston Motors," International Journal of Modelling and Simulation.
  4. Coauthors, Yanzhen Zhu, et al., "MFSPart: A Generalized Partitioning Framework for Multi-FPGA Systems and Its Ensemble-Based Extension," IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems.

Patents

  • Yanzhen Zhu, et al., Intelligent bricklaying robot invention patent.
  • Yanzhen Zhu, et al., VCD parsing method based on perfect hashing and post-processing logic.
  • Yanzhen Zhu, et al., Grouping method for automatic planning of memory built-in self-test.

Education

IC design, EDA algorithms, numerical methods and computer architecture.

2024.09 - 2025.09

Hong Kong University of Science and Technology

Visiting research in Prof. Zhiyao Xie's group, focused on deep learning for VLSI EDA placement algorithms.

2023.09 - 2026.06

Guangdong University of Technology, M.S. Research

Prof. Shuting Cai's group, focused on 3D-IC placement, VLSI EDA placement, and partitioning algorithms.

2019.09 - 2023.06

Guangdong University of Technology, B.S. in Integrated Circuit

Coursework included embedded systems, data structures and algorithms, computer architecture, SoC design, analog IC, digital IC, DSP, and information theory.

Skills

A stack forged in compiler, EDA algorithm, Linux, circuit and software design.

AI Compiler Graph Optimization Code Generation Operator Fusion VLSI Backend Algorithms Multi-FPGA Partitioning Macro Placement ILP Workflows Hypergraph Methods STA-Driven Optimization C C++ Python Verilog HDL Shell Linux Docker Kubernetes Nginx FreeRTOS RISC-V Zynq-7020 AXI4

Awards

Competition results in IC EDA, electronics, intelligent vehicles, and robotics.