<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/"><channel><title>Design · HJ4IT</title><link>https://hj4.it/categories/design/</link><description>Design · HJ4IT</description><generator>Hugo</generator><language>ko-KR</language><copyright>© CC BY 4.0</copyright><lastBuildDate>Sat, 22 Feb 2025 00:00:00 +0000</lastBuildDate><atom:link href="https://hj4.it/categories/design/index.xml" rel="self" type="application/rss+xml"/><item><title>[Design] Escaping Routing Hell: Partitioning 1.55B Gates (2)</title><link>https://hj4.it/posts/atom-chip-emulation-fpga-part2/</link><pubDate>Sat, 22 Feb 2025 00:00:00 +0000</pubDate><guid isPermaLink="true">https://hj4.it/posts/atom-chip-emulation-fpga-part2/</guid><dc:creator>Hyunje Jo</dc:creator><category>Design</category><category domain="tag">FPGA</category><category domain="tag">Atom Chip</category><category domain="tag">ASIC</category><category domain="tag">Emulation</category><description>Building on Part 1, this post digs into the 128-bit NOC AXI bus bottleneck crossing SLR boundaries inside the VU19P. We used manual HSTDM partitioning and iterative constraint generation to spread read and write channels across the best FPGA locations, and SerDes serialization cut inter-SLR SSI traffic by 16x. The design closed place and route at 80% SSI utilization despite 98.8% LUT consumption.</description></item><item><title>[Design] Taming 1.55B Gates: Atom Chip FPGA Emulation (1)</title><link>https://hj4.it/posts/atom-chip-emulation-fpga/</link><pubDate>Fri, 21 Feb 2025 00:00:00 +0000</pubDate><guid isPermaLink="true">https://hj4.it/posts/atom-chip-emulation-fpga/</guid><dc:creator>Hyunje Jo</dc:creator><category>Design</category><category domain="tag">FPGA</category><category domain="tag">Atom Chip</category><category domain="tag">Emulation</category><category domain="tag">Engineering</category><description>This post walks through the system we built to emulate the 1.55 billion gate Atom chip: a HAPS-100 carrying four VU19P FPGAs plus nine Xilinx U250 boards linked by Aurora IP. The Neural Engine was offloaded to the U250s while shared memory kept dedicated resources. We pushed back severe routing congestion on the HAPS-100 with AXI bus partitioning, SerDes logic, and manual LVDS bundle mapping.</description></item><item><title>[Design] Shrinking NPU Shared Memory: 48% Fewer Muxes on a Shared Bus</title><link>https://hj4.it/posts/rbln-chip-bist/</link><pubDate>Mon, 01 Jul 2024 21:20:00 +0900</pubDate><guid isPermaLink="true">https://hj4.it/posts/rbln-chip-bist/</guid><dc:creator>Hyunje Jo</dc:creator><category>Design</category><category domain="tag">ASIC</category><category domain="tag">Memory Architecture</category><category domain="tag">DFT</category><category domain="tag">Atom Chip</category><description>This post details how we optimized NPU shared memory by pulling BIST logic away from the SRAM macros and onto a shared bus, cutting mux count 48% from 27K down to 14K. Applying CPU hardening techniques to the ASIC flow evened out cell density and removed IR drop hotspots, giving wider timing margins and denser SRAM placement than the previous scratchpad-style layout.</description></item></channel></rss>