Research

Single-pass learning on data streams

Data streams arrive continuously and fast, often too fast to store. My doctoral research built classifiers that learn in a single pass under a discard-after-learn constraint: each sample updates a set of hyperellipsoidal regions and is then thrown away, so memory stays bounded however long the stream runs.

Illustration of hyperellipsoid regions for three classes and the movement of their centres

Illustration: hyperellipsoid regions for three classes. Dashed shapes are earlier states of each region; arrows show how its centre moved as new samples were absorbed and discarded.
  • TRACED (Information Sciences, 2026) is a trend-adaptive classifier built on hyper-ellipsoidal regions. It resolves two kinds of ambiguity: points that fall outside every learned region (exterior regions) and points that fall inside several regions with different labels (coincident regions). For exterior points it uses how each region’s centre and width have been changing, its trend, to decide which region the point belongs to.
  • D4 (Expert Systems with Applications, 2025) keeps a single hyperellipsoid per class. Where the regions of two classes meet, it pairs their principal axes, the directions of each class’s data distribution, and decides in that reduced subspace.
  • Compared with state-of-the-art streaming learners, both methods run substantially faster and use less memory, which makes them suitable for TinyML on edge devices.
  • Both methods, along with baseline hyperellipsoid classifiers, are in the open-source Python package spdal (pip install spdal).
  • Try it: an interactive demo runs the real TRACED implementation chunk by chunk on a two-dimensional stream, standing still or moving from A to B. A second demo reruns the paper’s ablation on the 20-dimensional HDFR-B stream.

Applied NLP and large language models

I build NLP systems that work at national scale. Current work includes:

  • hierarchical classification of research proposals using LLMs
  • an embedding pipeline (Qwen3-Embedding) over more than 46,000 proposals for semantic search
  • multi-stage semantic deduplication that detects overlapping projects across funding repositories

Research directions: retrieval-augmented generation, semantic information retrieval, and LLM evaluation and alignment.

Graph kernels and bioinformatics

In my M.S. work I represented biochemical compounds as graphs of primitive structures and designed graph kernels for support vector machine classification of molecular target characteristics (JCSSE 2020). I also contributed to network-based target identification for compounds against dengue virus (Molecules, 2020).

The method is available as the Python package molprim (pip install molprim). Try it: an interactive demo shows how molprim turns a molecule into its primitive graph, with the published benchmark results.

Other projects

Year Project Methods
2025 Voter behavior and patterns Hierarchical clustering with Jaccard similarity, cluster stability analysis
2022 Risk assessment for digital asset portfolios CAPM, regression, time-series volatility
2019 Risk factors for heat stroke in military training Clustering, statistical inference
2018 Discovering governing equations from noisy data Sparse regression (LASSO)

Collaboration and supervision

I am open to collaborating on streaming and continual learning, efficient ML, and applied NLP. Students interested in these topics are welcome to get in touch.