BEST POST 🕺
-
[UPMEM PIM] 졸업논문 최종 발표
2025. 12. 10. Wednesday
2025.12.10 15:10 -
[Paper Review] Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Servinghttps://arxiv.org/abs/2407.00079 Mooncake: A KVCache-centric Disaggregated Architecture for LLM ServingMooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI. It features a KVCache-centric disaggregated architecture that separates the prefill and decoding clusters. It also leverages the underu..
2026.05.28 16:20 -
[PIM] UPMEM Simulator Example
UPMEM Hello World! Examplehttps://sdk.upmem.com/stable/02_HelloWorld.html Hello World! Example — UPMEM DPU SDK 2025.1.0 Documentation© Copyright 2015-2024, UPMEM SAS - All rights reserved.sdk.upmem.com 0. UPMEM SDK 설치하기 https://sdk.upmem.com UPMEM DPU SDKUPMEM SDK The Software Development Kit for programming and using the DPU provided by the UPMEM Acceleration platform.sdk.upmem.com tar -..
2025.09.21 15:35
NEW POST ✨
-
[Paper Review] ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs
ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs (ISCA 2026)https://arxiv.org/abs/2604.03298 ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUsThe rapid scaling of Large Language Models presents significant challenges for their deployment and inference, particularly on resource-constrained specialized AI hardware accelerators su..
2026.08.07 02:18 -
[Paper Review] NetZIP: Algorithm/Hardware Co-design of In-network Lossless Compression for Distributed Large Model Training (MICRO'25)
NetZIP: Algorithm/Hardware Co-design of In-network Lossless Compression for Distributed Large Model Traininghttps://dl.acm.org/doi/full/10.1145/3725843.3756079 NetZIP: Algorithm/Hardware Co-design of In-network Lossless Compression for Distributed Large Model Training | Proceedings of thPublication History Published: 17 October 2025dl.acm.org Abstract In distributed large model training, t..
2026.06.10 22:06 -
[Paper Review] PreSto: An In-Storage Data Preprocessing System for Training Recommendation Models (ISCA'24)
PreSto: An In-Storage Data Preprocessing System for Training Recommendation Models (ISCA'24) https://arxiv.org/abs/2406.14571 PreSto: An In-Storage Data Preprocessing System for Training Recommendation ModelsTraining recommendation systems (RecSys) faces several challenges as it requires the "data preprocessing" stage to preprocess an ample amount of raw data and feed them to the GPU for trainin..
2026.06.08 17:23 -
[Presentation] Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
2026.06.01 17:53 -
[Paper Review] Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Servinghttps://arxiv.org/abs/2407.00079 Mooncake: A KVCache-centric Disaggregated Architecture for LLM ServingMooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI. It features a KVCache-centric disaggregated architecture that separates the prefill and decoding clusters. It also leverages the underu..
2026.05.28 16:20 -
[Paper Review] Reducing Solid-State Drive Read Latency by Optimizing Read-Retry (ASPLOS'21)
Reducing Solid-State Drive Read Latency by Optimizing Read-Retry (ASPLOS'21)https://dl.acm.org/doi/10.1145/3445814.3446719 Reducing solid-state drive read latency by optimizing read-retry | Proceedings of the 26th ACM International Conference on ArchiAs data privacy and security rapidly become key requirements, securely erasing data from a storage system becomes as important as reliably storing ..
2026.05.27 21:26 -
[Paper Review] DPES: Lifetime Improvement of NAND Flash-based Storage SystemsUsing Dynamic Program and Erase Scaling
DPES: Lifetime Improvement of NAND Flash-based Storage SystemsUsing Dynamic Program and Erase Scalinghttps://www.usenix.org/conference/fast14/technical-sessions/presentation/jeong Lifetime Improvement of NAND Flash-based Storage Systems Using Dynamic Program and Erase Scaling | USENIX www.usenix.org Abstract The cost-per-bit of NAND flash memory has been continuously improved by semiconductor ..
2026.05.25 21:46 -
[Paper Review] AERO: Adaptive Erase Operation for Improving Lifetime and Performance of Modern NAND Flash-Based SSDs
AERO: Adaptive Erase Operation for Improving Lifetime and Performance of Modern NAND Flash-Based SSDshttps://arxiv.org/abs/2404.10355 20 V) to flash cells for a long time (e.g.," data-og-host="arxiv.org" data-og-source-url="https://arxiv.org/abs/2404.10355" data-og-url="https://arxiv.org/abs/2404.10355v1" data-og-image="https://blog.kakaocdn.net/dna/uZrr6/dJMb81fXwvb/AAAAAAAAAAAAAAAAAAAAAKyjoNtCZhSarbZnEulkxa1SqmFvpTDEzaKyEcK_nsVG/img.p..?credential=yqXZFxpELC7KVnFOS48ylbz2pIh7yKj8&expires=1790780399&allow_ip=&allow_referer=&signature=mU7c1HbXHFUWBdHxdKn8ab7uRpU%3D
2026.05.14 16:36 -
[Paper Review] ZipLLM: Efficient LLM Storage via Model-AwareSynergistic Data Deduplication and Compression
ZipLLM: Efficient LLM Storage via Model-AwareSynergistic Data Deduplication and Compression (NSDI'26)https://www.usenix.org/conference/nsdi26/presentation/wang-zirui ZipLLM: Efficient LLM Storage via Model-Aware Synergistic Data Deduplication and Compression | USENIXOpen Access Media USENIX is committed to Open Access to the research presented at our events. Papers and proceedings are freely ava..
2026.05.07 19:34 -
[Paper Review] Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric Perspective
Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric Perspective https://arxiv.org/abs/2511.00739 Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric PerspectiveAgentic AI serving converts monolithic LLM-based inference to autonomous problem-solvers that can plan, call tools, perform reasoning, and adapt on the fly. Due to diverse ..
2026.05.06 19:39 -
[Paper Review] Processing in Memory: The Terasys Massively Parallel PlM Array
Processing in Memory: The Terasys Massively Parallel PlM Array (PACT'95) The notion of computing in memory has been with us for several decades. For example, Stone’ proposed a logic-in-memory computer consisting of an enhanced cache memory array that serves as a high-speed buffer between CPU and conventional memory. More recently, a group at the University of Toronto has designed a computati..
2026.04.22 21:21 -
[Paper Review] DFTL: A Flash Translation Layer Employing Demand-basedSelective Caching of Page-level Address Mappings
DFTL: A Flash Translation Layer Employing Demand-based Selective Caching of Page-level Address Mappings (ASPLOS'09) Abstract Recent technological advances in the development of flash memory based devices have consolidated their leadership position as the preferred storage media in the embedded systems market and opened new vistas for deployment in enterprise-scale storage systems. Unlike hard ..
2026.04.22 17:46 -
[Paper Review] Hitting the Memory Wall: Implications of the Obvious
Hitting the Memory Wall: Implications of the Obvious (1994) This brief note points out something obvious—something the authors “knew” without really understanding. With apologies to those who did understand, we offer it to those others who, like us, missed the point.We all know that the rate of improvement in microprocessor speed exceeds the rate of improvement in DRAM memory speed; each is imp..
2026.04.17 16:09 -
[Paper Review] Near-Memory Computing: Past, Present, and Future
Near-Memory Computing: Past, Present, and Future The conventional approach of moving data to the CPU for computation has become a significant performance bottleneck for emerging scale-out data-intensive applications due to their limited data reuse. At the same time, the advancement in 3D integration technologies has made the decade-old concept of coupling compute units close to the memory — call..
2026.04.17 14:29 -
[Paper Review] QuCo: Efficient and Flexible Hardware-Driven Automatic Configure
QuCo: Efficient and Flexible Hardware-Driven Automatic Configuration of Tile Transfers in GPUs (HPCA'26) Abstract The growing complexity and parallelism demands of modern GPU workloads have driven architectural innovations toward asynchronous tile transfers (ATTs) to overlap computation and data movement. While ATT units such as the NVIDIA’s Tensor Memory Accelerator (TMA) introduce high-thro..
2026.04.03 00:41 -
[Interconnection Networks] Chap24. Simulation
Interconnection Networks, Simulation simulation is a double-edged sword — while it can provide excellent models of complex network designs, simulators and simulations are equally complex. To that end, the quality of simulation results is only as good as the methodology used to generate and measure these results. In this chapter, we address the basics of simulation input, measurement, and design...
2026.03.25 21:03 -
[CUDA] Chap11. Prefix sum (scan)
Chap11. Prefix sum (scan) Our next parallel pattern is prefix sum, which is also commonly known as scan. Parallel scan is frequently used to parallelize seemingly sequential operations, such as resource allocation, work assignment, and polynomial evaluation. In general, if a computation is naturally described as a mathematical recursion in which each item in a series is defined in terms of the p..
2026.03.24 20:48 -
[CUDA] Chap10. Reduction
Chap10. Reduction A reduction derives a single value from an array of values. The single value could be the sum, the maximum value, the minimal value, and so on among all elements. The value can also be of various types: integer, single-precision floating-point, double-precision floating-point, half-precision floating-point, characters, and so on. All these types of reductions have the same comp..
2026.03.24 17:06 -
[CUDA] Chap9. Parallel histogram
Chap9. Parallel histogram In practice, whenever there is a large volume of data that needs to be analyzed to distill interesting events, histograms are likely used as a foundational computation. Note that multiple threads need to update the same counter (m-p), which is a conflict that is referred to as output interference. Programmers must understand the concepts of race conditions and atomic o..
2026.03.24 16:54 -
[CUDA] Chap8. Stencil
Chap8. Stencil The data that is processed by stencil-based algorithms consists of discretized quantities of physical significance, such as mass, velocity, force, acceleration, temperature, electrical field, and energy, whose relationships with each other are governed by differential equations. A common use of stencils is to approximate the derivative values of a function based on the function va..
2026.03.24 16:40