Samyamoy Rakshit

Samyamoy Rakshit

Data Scientist

I rebuild published models from the paper — the transformer, BERT, ViT — tensor by tensor in PyTorch, all three on Bengali data, all three trained end to end on one M1 laptop with 16 GB of RAM. Reading a paper and reimplementing it turn out to be very different kinds of understanding. By day I build forecasting and agentic systems at Sigmoid.

Samyamoy Rakshit
Bengaluru, India

Experience

  1. Jan 2026 – present

    Data Scientist · Sigmoid

    Re-engineered daily demand forecasting across 9,200 machine × product series — full run from ~3 hours to ~3 minutes. Plus an MCP proof of concept driven from three model backends.

  2. Dec 2024 – Dec 2025

    Associate Data Scientist · Sigmoid

    Built the feature-engineering pipeline behind a four-agent RAG recommendation engine, and worked on one of those agents — live since June 2025, ~80% accuracy in its core categories. Earlier, price elasticity and promo baselines for a global CPG manufacturer.

  3. Aug 2024 – Dec 2024

    Data Science Intern · Sigmoid

    Measured which levers — price, promo, media, retail execution — actually move sell-through for a global spirits manufacturer, so spend could be aimed where it pays.

Full record

Projects

papers-from-scratchSOTA architectures rebuilt in PyTorch

Three papers replicated tensor by tensor on Bengali data — no nn.Transformer, no nn.MultiheadAttention, no pre-built HuggingFace models.

  1. BERT

    BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingDevlin et al., 2019

    A 7.5M-parameter Bengali BERT, pre-trained from nothing on a laptop, that scores above published mBERT and IndicBERT on Bengali news classification.

    Params
    ~7.5M
    Corpus
    ~114MB
    Pre-trained
    ~28h
    Accuracy
    86.5%
  2. Transformer

    Attention Is All You NeedVaswani et al., 2017

    The original transformer rebuilt component by component, then trained to translate English into Bengali on one laptop.

    Params
    ~11M
    Paper
    ~65M
    Hardware
    M1, 16GB
    Data
    Samanantar
  3. Vision Transformer

    An Image Is Worth 16x16 WordsDosovitskiy et al., 2021

    A ViT trained on 1,935 photographs of Bengali terracotta temples — a dataset I had to build first — to test the paper's claim about pre-training. It failed, exactly as predicted.

    Images
    1,935
    Classes
    10
    From scratch
    14.9%
    Pre-trained ViT-B/16
    88.9%
All three in full

Education

  • Nov 2020 – Jul 2024

    B.Tech, Computer Science and Engineering (Data Science)

    Haldia Institute of Technology

    CGPA 9.05. Affiliated to Maulana Abul Kalam Azad University of Technology, West Bengal.

  • 2020

    Higher Secondary (WBCHSE), Science — 94.4%

    Jhantipahari High School

  • 2018

    Secondary (WBBSE) — 90.4%

    Jhantipahari High School

Writing

BERT, from scratch

A 7.5M-parameter Bengali BERT, pre-trained from nothing on a laptop, that scores above published mBERT and IndicBERT on Bengali news classification.

1 min read
All posts

Skills

Everything here is something I have shipped with or trained with.

Languages

  • Python
  • SQL

Deep learning

  • PyTorch
  • Transformers
  • BERT
  • Vision Transformers

Generative and agentic

  • Azure OpenAI
  • Azure ML Prompt Flow
  • RAG
  • Multi-agent systems
  • Model Context Protocol
  • Ollama

Modelling and statistics

  • scikit-learn
  • LightGBM
  • Prophet
  • Mixed-effects models
  • SARIMAX
  • Ridge and Lasso
  • Experiment design

Data and cloud

  • Snowflake
  • Microsoft Fabric
  • Azure ML