Sitemap
A list of all the posts and pages found on the site. For you robots out there, there is an XML version available for digesting as well.
Pages
Posts
[gem5 Q&A] Why there is miss prediction of non-control instructions
Published:
Hello, In function checkSignalsAndUpdate(ThreadID tid) in src/cpu/o3/fetch.cc file, it seems miss prediction can still happen from commit and decode even if mispredictInst->isControl() is false.
[gem5 Q&A] Page Walker: Where the PTE hits in the memory hierarchy
Published:
Hi, I am working on the x86 page walker in gem5. I understand that the page walker accesses the page walker cache (PWC) first and, in case of a miss, it accesses the memory hierarchy (L1, then L2, then L3 caches and lastly the memory). This happens through the packetpointer read, which reads the physical address of the entry at each level (PML4, PDP.. etc.).
[gem5 tech mark] Why are stores in the SQ assumed to have valid addresses?
Published:
Hi, I’m doing some gem5 hacking for research and have been confused over the timing of when loads search the store queue (SQ) and when stores have valid addresses that can be compared against. Gem5 includes an assert in the read() method in the LSQ unit that the addresses of all stores before the executing loads are valid, but I don’t understand how this can be guaranteed in OoO execution.
[gem5 Q&A] Fixed I/O Address Range in x86
Published:
Hi all, I’m trying to model the SPEC HPC benchmark suite in gem5 with an x86 ISA using KVM. As a result, I am trying to link the “_addr” version of the m5ops against the binaries in order to model the region of interest. Unfortunately, I get the following error when trying to build the sample hello world example:
[gem5 Q&A] Microcode_ROM Instruction and fetchRomMicroop() Function
Published:
Hello, I am looking at the AtomicSimpleCPU code in src/cpu/simple for x86 ISA. I am trying to understand the following code snippet. Whenever this condition is true for a given PC, it does NOT follow the regular fetch from the instruction cache and then decode. This results in a macroop called Microcode_ROM, which is not an x86 macroop that has a sequence of uops (can be seen in the O3 CPU). Example: Instruction is: Microcode_ROM : ldst t0, HS:[t0 + t6 + 0x20] (This is taken from the O3 logs running the same workload by checking the same PC in the Debug logs).
[gem5 Q&A] Squashing Instructions after Page Table Fault
Published:
Hello, I am currently trying to locate the code that is used to squash instructions if a Page Table Fault is triggered in the O3 CPU. After using the PageTableWalker Debug Flags, my current guess would be gem5/src/arch/x86/pagetable_walker.cc in line 199. Furthermore I inspected the files in the src/cpu/o3 directory, but couldn’t find anything specific to squashing instructions after a fault.
Is my assumption correct, that the O3 CPU implementation does not handle these things on its own, but the architectural part of the implementation does it? I am missing something, feel free to point it out.
portfolio
Portfolio item number 1
Short description of portfolio item number 1
Portfolio item number 2
Short description of portfolio item number 2 
publications
Fuzzy Flow Regulation for Network-on-Chip based Chip Multiprocessors Systems
In 2014 19th Asia and South Pacific Design Automation Conference (ASP-DAC), 2014
Delays;Regulators;Fuzzy logic;Network-on-chip;Throughput;Pragmatics;Multiprocessor interconnection;Network-on-Chip;Chip Multiprocessor;Flow Regulation;Fuzzy Logic
Towards Stochastic Delay Bound Analysis for Network-on-Chip
In 2014 Eighth IEEE/ACM International Symposium on Networks-on-Chip (NoCS), 2014
Stochastic processes;Delays;Interference;Calculus;Analytical models;Servers;System-on-chip[<35;31;32M
DVFS for NoCs in CMPs: A Thread Voting Approach
In 2016 IEEE International Symposium on High Performance Computer Architecture (HPCA), 2016
DVFS, Multi-core
Memory-Access Aware DVFS for Network-on-Chip in CMPs
In 2016 Design, Automation & Test in Europe Conference & Exhibition (DATE), 2016
Switches;Resource management;Delays;Load modeling;Nickel;Tuning;Benchmark testing
Opportunistic Competition Overhead Reduction for Expediting Critical Section in NoC Based CMPs
In 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA), 2016
Critical Section; CMP; NoC; OS
Aggregate Flow-Based Performance Fairness in CMPs
In ACM Transactions on Architecture and Code Optimization (TACO), Volume 13, Issue 4 Article No.: 53, Pages 1 - 27, 2016
computer architecture, performance fairness, quality of service
Dynamic Traffic Regulation in NoC-Based Systems
In IEEE Transactions on Very Large Scale Integration (VLSI) Systems (Volume: 25, Issue: 2, February 2017) , 2017
Delays;System performance;IP networks;Nickel;Regulators;Calculus;Network-on-chip;Chip multi/many-core processor (CMP);fuzzy control;multi/many-processor systems-on-chip (MPSoC);network-on-chip (NoC);traffic engineering
Marginal Performance: Formalizing and Quantifying Power Over/Under Provisioning in NoC DVFS
In IEEE Transactions on Computers (Volume: 66, Issue: 11, 01 November 2017) , 2017
Energy efficiency;Power demand;Measurement;Benchmark testing;Energy efficiency;Network-on-chip;Program processors;Performance evaluation;power efficiency;DVFS;network-on-chip (NoC);CMP
iNPG: Accelerating Critical Section Access with In-network Packet Generation for NoC Based Many-Cores
In 2018 IEEE International Symposium on High Performance Computer Architecture (HPCA), 2018
Instruction sets;Spinning;Liquid crystal on silicon;Coherence;Acceleration;Routing protocols;In Network Packet Generation;Critical Section;Synchronisation Primitive;Cache Coherency;Network on Chip;CMP
Best-paper candidate
Thread Voting DVFS for Manycore NoCs
In IEEE Transactions on Computers ( Volume: 67, Issue: 10, 01 October 2018) , 2018
Measurement;Message systems;System-on-chip;Instruction sets;Voltage control;Load modeling;Power system management;Chip manycore processor (CMP);DVFS;network on chip (NoC);power/energy efficiency
Pursuing Extreme Power Efficiency With PPCC Guided NoC DVFS
In IEEE Transactions on Computers (Volume: 69, Issue: 3, 01 March 2020) , 2020
Power demand;Message systems;Tuning;Thermal management;Monitoring;Energy consumption;Power system management;Manycore processor;DVFS;NoC;power efficiency;CMP
TSOPER: Efficient Coherence-Based Strict Persistency
In 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA), 2021
Protocols;Program processors;Nonvolatile memory;Computational modeling;Semantics;Coherence;Computer architecture;non-volatile memory;persistent memory;persistency;total store order;coherence
I am co-first author
Game-of-Life Temperature-Aware DVFS Strategy for Tile-based Chip Many-Core Processor
In IEEE Journal on Emerging and Selected Topics in Circuits and Systems (JETCAS, Volume: 13, Issue: 1, March 2023), 2023
Dynamic voltage scaling, multiprocessor interconnection, automata.
SE-CNN: Convolution Neural Network Acceleration via Symbolic Value Prediction
In IEEE Journal on Emerging and Selected Topics in Circuits and Systems (JETCAS, Volume: 13, Issue: 1, March 2023), 2023
Artificial intelligence, artificial neural networks, AI accelerators.
Silent Stores in the Battery-less Internet of Things: A Good Idea?
In Proceedings of the 2023 International Conference on Embedded Wireless Systems and Networks (EWSN), 2023
sensor network;store buffer;silent store;low power devices
TaDA: Task Decoupling Architecture for the Battery-less Internet of Things
In Proceedings of the 22nd ACM Conference on Embedded Networked Sensor Systems (SenSys), 2024
Task decoupling, Internet of Things (IoT), energy harvesting, intermittent computing
TangramFP: Energy-Efficient, Bit-Parallel, Multiply-Accumulate for Deep Neural Networks
In 2024 IEEE 36th International Symposium on Computer Architecture and High Performance Computing (SBAC-PAD), 2024
Bit-Parallel, Energy-Efficient Multiply Accumulate, Deep Neural Networks
Best computer architecture track paper
RXT: RefleXive address Translation for Pointer-Chasing Workloads
In 39th IEEE International Parallel & Distributed Processing Symposium (IPDPS), 2025
Early acceptance paper
Anatomy of the gem5 Simulator: AtomicSimpleCPU, TimingSimpleCPU, O3CPU, and Their Interaction with the Ruby Memory System
In arXiv preprint, 2025
An architectural guide to gem5 CPU models and the Ruby memory system.
The Fake-Busy and True-Idle Problems of Running Graph Applications on Chiplet-Based Multi-Cores
In 2025 IEEE International Symposium on Workload Characterization (IISWC), 2025
Chiplet-specific execution phenomena in graph workloads.
Prefill vs. Decode Bottlenecks: SRAM-Frequency Tradeoffs and the Memory-Bandwidth Ceiling
In arXiv preprint, 2025
SRAM and frequency tradeoffs for the compute-bound prefill and memory-bound decode phases of LLM inference.
Understanding Simulated Architecture via gem5 Call-Stack Profiling
In 2026 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS), 2026
gem5; call-stack profiling; architecture simulation; observability
ConvReflex: Efficient Ultra-Low-Power CNN Inference via Clamping Prediction
In ACM/IEEE International Conference on Embedded Artificial Intelligence and Sensing Systems (SenSys), 2026
Embedded AI; CNN inference; ultra-low-power computing; microcontrollers
NoCWalk: In-Network Page Walks for Concurrent Data Structure Workloads on Multicore
In 2026 ACM International Conference on Computing Frontiers (CF), 2026
Virtual memory; page-table walks; networks-on-chip; concurrent data structures
DICE: Detailed Inter-Chiplet End-to-End PHY Modeling for Accurate Chiplet Simulation
In 2026 ACM/IEEE 53rd Annual International Symposium on Computer Architecture (ISCA), 2026
Chiplets; PHY modeling; inter-chiplet links; architecture simulation
[To appear]
Links: arXiv, GitHub repository
DICE: Detailed Inter-Chiplet End-to-End PHY Modeling for Accurate Chiplet Simulation
In arXiv preprint, 2026
Detailed inter-chiplet end-to-end PHY modeling for accurate chiplet simulation.
Links: arXiv, GitHub repository
talks
Anatomy of the gem5 Simulator
Published:
teaching
Computer Architecture I
Undergraduate course, Uppsala University, Department of IT, 2024
I have been the course responsible since 2021
Course’s webpage at Uppsala University
Accelerating Systems with Programmable Logic Components
Graduate course, Uppsala University, Department of IT, 2024
I have been the course responsible since 2020.
Course’s webpage at Uppsala University
tools
Statically-linked PARSEC-3.0 benchmarks
Published:
In this project I modified PARSEC-3.0 benchmarks using static linking with gem5 hooks for the x86_64 architecture.
Statically-linked WHISPER benchmarks
Published:
WHISPER benchmark suite with static linkage.
My config files for some gnu tools such as emacs, vim, etc
Published:
Some of my personal configurations for some of my personal favoriate GNU Tools
TangramFP
Published:
TangramFP: Energy-Efficient, Bit-Parallel Multiply-Accumulate for Deep Neural Networks
DICE Simulator
Published:
DICE is an open architecture-level simulation framework for detailed inter-chiplet end-to-end PHY-link modeling.
