cv

Curriculum vitae in HTML, with a downloadable PDF version.

Basics

Name Piyush Sao
Label Staff Scientist
Email saopk@ornl.gov
Phone +1(404) 405 9940
Summary Staff scientist at Oak Ridge National Laboratory (ORNL), focusing on improving scientific computing and machine learning algorithms using high-performance computing platforms.

Work

  • 2019.08 - Present
    Staff Scientist
    Oak Ridge National Laboratory
    Staff scientist in the Computational Data Analytics Group, focusing on enhancing scientific computing and machine learning algorithms through high-performance computing platforms.
  • 2018.08 - 2019.07
    Postdoctoral Research Associate
    Oak Ridge National Laboratory
    Postdoctoral research associate in the Computer Science Research Group.
  • 2014.05 - 2014.08
    Summer Intern
    Intel Corporation
    Summer intern at the Parallel Computing Lab.
  • 2013.05 - 2013.08
    Summer Intern
    Lawrence Berkeley National Laboratory
    Summer intern at the Computational Research Division.
  • 2011.01 - 2018.01
    Graduate Research Assistant
    Georgia Institute of Technology
    Graduate research assistant at Georgia Institute of Technology.

Education

  • 2018

    Atlanta, GA, USA

    PhD
    Georgia Institute of Technology
    Computational Science and Engineering
    • Thesis: Scalable and Resilient Sparse Linear Solvers
    • Advisor: Dr. Richard Vuduc
  • 2016

    Atlanta, GA, USA

    Master of Science
    Georgia Institute of Technology
    Computational Science and Engineering

    GPA: 4.0/4.0

  • 2011

    Madras, India

    Bachelor of Technology
    Indian Institute of Technology
    Electrical Engineering, Minor: Theoretical Computer Science

    GPA: 8.5/10.0

  • 2011

    Madras, India

    Master of Technology
    Indian Institute of Technology
    Electrical Engineering, Specialization: Microelectronics and VLSI design

    GPA: 8.5/10.0

Awards

  • 2022.01.01
    ORNL Special Performance Award
    For outstanding research contributions in the Computer Science and Mathematics Division
  • 2022.01.01
    SC22 Gordon Bell Finalist
    (Media), Finalist for submission "Exaflops biomedical knowledge graph analytics"
  • 2022.01.01
    SIAM PP22 Best Paper Prize
    (Link), Winner of the SIAM Activity Group on Supercomputing Best Paper Prize; (SIAM News), (ORNL News)
  • 2021.01.01
    R&D 100 Award Finalist
    (Link)
  • 2020.01.01
    SC20 Gordon Bell Finalist
    (Link), Finalist for submission "Scalable Knowledge Graph Analytics at 136 PetaFlop/s"
  • 2019.01.01
    ORNL Outstanding Postdoctoral Research Associate
    For outstanding research contributions in the Computer Science and Mathematics Division
  • 2019.01.01
    Graph500
    (Link), Member of the technical team that placed the Summit Supercomputer at ORNL 4th in the prestigious Graph500 List

Publications

2026

  1. A Second-Moment Theory for Floating-Point Reduction Trees
    Piyush Sao, Narasinga Miniskar, Pedro Valero-Lara, and 2 more authors
    arXiv preprint arXiv:2607.18758, Jul 2026
  2. Contraction-Gauge Preconditioning for Quantized Matrix Multiplication
    Piyush Sao, Narasinga Miniskar, Pedro Valero-Lara, and 2 more authors
    arXiv preprint arXiv:2607.18745, Jul 2026
  3. Ghosts of Softmax: Complex Singularities That Limit Safe Step Sizes in Cross-Entropy
    Piyush Sao
    arXiv preprint arXiv:2603.13552, Mar 2026
  4. Fast Evaluation of Truncated Neumann Series by Low-Product Radix Kernels
    Piyush Sao
    arXiv preprint arXiv:2602.11843, Feb 2026
  5. What Trace Powers Reveal About Log-Determinants: Closed-Form Estimators, Certificates, and Failure Modes
    Piyush Sao
    arXiv preprint arXiv:2601.12612, Jan 2026

2025

  1. Fast Active-Set Thresholding Method for Nonnegative Least Squares
    Benjamin Cobb, Ramakrishnan Kannan, Konstantin Pieper, and 5 more authors
    In 2025 IEEE International Conference on Big Data (BigData), 2025
  2. Knowledge graph analytics kernels in high performance computing
    Ramakrishnan Kannan, Piyush K Sao, Hao Lu, and 5 more authors
    2025
    US Patent 12,417,246
  3. Forward Error Bounds and Efficient Algorithms for Computing a Tensor Times Matrix Chain in Low Precision on GPUs
    Julian Bellavita, Piyush Sao, and Ramakrishnan Kannan
    2025
    SC25 poster

2024

  1. PANDORA: A Parallel Dendrogram Construction Algorithm for Single Linkage Clustering on GPU
    Piyush Sao, Andrey Prokopenko, and Damien Lebrun-Grandié
    In Proceedings of the 53rd International Conference on Parallel Processing, 2024
  2. Interface for sparse linear algebra operations
    Ahmad Abdelfattah, Willow Ahrens, Hartwig Anzt, and 32 more authors
    arXiv preprint arXiv:2411.13259, 2024
  3. Accelerated Constrained Sparse Tensor Factorization on Massively Parallel Architectures
    Yongseok Soh, Ramakrishnan Kannan, Piyush Sao, and 1 more author
    In Proceedings of the 53rd International Conference on Parallel Processing, 2024

2023

  1. Newly Released Capabilities in Distributed-memory SuperLU Sparse Direct Solver
    Xiaoye S Li, Paul Lin, Yang Liu, and 1 more author
    ACM Transactions on Mathematical Software, 2023
  2. Unified Communication Optimization Strategies for Sparse Triangular Solver on CPU and GPU Clusters
    Yang Liu, Nan Ding, Piyush Sao, and 2 more authors
    In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, 2023
  3. Optimizing Communication in 2D Grid-Based MPI Applications at Exascale
    Hao Lu, Piyush Sao, Michael Matheson, and 3 more authors
    In Proceedings of the 30th European MPI Users’ Group Meeting, 2023
  4. Brief Announcement: Communication Optimal Sparse LU Factorization for Planar Matrices
    Piyush Sao, and Xiaoye Sherry Li
    In Proceedings of the 35th ACM Symposium on Parallelism in Algorithms and Architectures, 2023

2022

  1. A single-tree algorithm to compute the Euclidean minimum spanning tree on GPUs
    Andrey Prokopenko, Piyush Sao, and Damien Lebrun-Grandie
    In Proceedings of the 51st International Conference on Parallel Processing, 2022
  2. Exaflops biomedical knowledge graph analytics
    Ramakrishnan Kannan, Piyush Sao, Hao Lu, and 8 more authors
    In 2022 SC22: International Conference for High Performance Computing, Networking, Storage and Analysis (SC), 2022
  3. FUNNL: Fast Nonlinear Nonnegative Unmixing for Alternate Energy Systems
    Jeffrey A Graves, Thomas F Blum, Piyush Sao, and 2 more authors
    In Knowledge-Guided Machine Learning, 2022
  4. Resilience Design Patterns: A Structured Approach to Resilience at Extreme Scale (Version 2.0)
    Christian Engelmann, Rizwan Ashraf, Saurabh Hukerikar, and 2 more authors
    2022

2021

  1. Scalable All-pairs Shortest Paths for Huge Graphs on Multi-GPU Clusters
    Piyush Sao, Hao Lu, Ramakrishnan Kannan, and 3 more authors
    In Proceedings of the 30th International Symposium on High-Performance Parallel and Distributed Computing, 2021
  2. Sparse Binary Matrix-Vector Multiplication on Neuromorphic Computers
    Catherine D Schuman, Bill Kay, Prasanna Date, and 3 more authors
    In 2021 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW), 2021
  3. Dense semiring linear algebra on modern cuda hardware
    Vijay Thakkar, Ramakrishnan Kannan, Piyush Sao, and 5 more authors
    2021
    SIAM Computational Sciences and Engineering. SIAM

2020

  1. Scalable knowledge graph analytics at 136 petaflop/s
    Ramakrishnan Kannan, Piyush Sao, Hao Lu, and 5 more authors
    In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis, 2020
  2. A supernodal all-pairs shortest path algorithm
    Piyush Sao, Ramakrishnan Kannan, Prasun Gera, and 1 more author
    In Proceedings of the 25th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming, 2020
  3. Traversing large graphs on GPUs with unified memory
    Prasun Gera, Hyojong Kim, Piyush Sao, and 2 more authors
    Proceedings of the VLDB Endowment, 2020

2019

  1. Multifrontal Non-negative Matrix Factorization
    Piyush Sao, and Ramakrishnan Kannan
    In International Conference on Parallel Processing and Applied Mathematics, 2019
  2. Self-stabilizing Connected Components
    Piyush Sao, Christian Engelmann, Srinivas Eswar, and 2 more authors
    In 2019 IEEE/ACM 9th Workshop on Fault Tolerance for HPC at eXtreme Scale (FTXS), 2019
  3. A Communication-avoiding 3D Sparse Triangular Solve Algorithm
    Piyush Sao, Ramakrishnan Kannan, Xiaoye Li, and 1 more author
    In International Conference on Supercomputing, Jun 2019
  4. A communication-avoiding 3D algorithm for sparse LU factorization on heterogeneous systems
    Piyush Sao, Xiaoye S Li, and Richard Vuduc
    Journal of Parallel and Distributed Computing, 2019

2018

  1. A communication-avoiding 3D LU factorization algorithm for sparse matrices
    Piyush Sao, Xiaoye S. Li, and Richard Vuduc
    In Proceedings of the IEEE International Parallel and Distributed Processing Symposium (IPDPS), May 2018
  2. Scalable and Resilient Sparse Linear Solvers
    Piyush Sao
    Georgia Institute of Technology, Aug 2018

2016

  1. A Self-Correcting Connected Components Algorithm
    Piyush Sao, Oded Green, Chirag Jain, and 1 more author
    In Proceedings of the ACM Workshop on Fault-Tolerance for HPC at Extreme Scale, 2016

2015

  1. A Sparse Direct Solver for Distributed Memory Xeon Phi-accelerated Systems
    Piyush Sao, Xing Liu, Richard Vuduc, and 1 more author
    In Parallel and Distributed Processing Symposium (IPDPS), 2015 IEEE International, 2015

2014

  1. A distributed CPU-GPU sparse direct solver
    Piyush Sao, Richard Vuduc, and Xiaoye Sherry Li
    In European Conference on Parallel Processing, 2014
  2. A distributed kernel summation framework for general-dimension machine learning
    Dongryeol Lee, Piyush Sao, Richard Vuduc, and 1 more author
    Statistical Analysis and Data Mining: The ASA Data Science Journal, 2014

2013

  1. Self-stabilizing iterative solvers
    Piyush Sao, and Richard Vuduc
    In Proceedings of the Workshop on Latest Advances in Scalable Algorithms for Large-Scale Systems, 2013

2011

  1. Model Order Reduction Techniques for VLSI Circuit Simulation
    Piyush Sao
    IIT Madras, May 2011

1999

  1. SuperLU Users’ Guide
    Xiaoye S Li, James W Demmel, John R Gilbert, and 4 more authors
    1999

Skills

C
Python
C++
MPI
Parallel Programming
OpenMP
Parallel Programming
CUDA
Parallel Programming
Matlab
Scientific Packages
Visualization
NumPy
Scientific Packages
Pandas
Scientific Packages
LAPACK
Scientific Packages
Matplotlib
Visualization
d3.js
Visualization
GraphViz
Visualization
Inkscape
Visualization
SQL
General Purpose
Latex
General Purpose
Git
General Purpose