Robotics software engineer focused on perception-to-action pipelines, motion planning, and real robot deployment.
I’m a mechatronics engineer and robotics student, currently pursuing a Master’s in Robotics. I like building systems that connect perception to action, from computer vision pipelines to motion planning and real robot execution.
I’m especially interested in robot manipulation, motion planning, and multi-robot systems (swarm-style coordination and distributed autonomy). I enjoy turning messy real-world problems into robust, testable software that runs on hardware.
I’m actively looking for robotics software roles and internships involving ROS 2, perception, planning, simulation, and embedded or real-time systems. I’m happiest on teams shipping tools that make robots more reliable in the real world.
A LoRA fine-tune of Llama 3.1 8B that runs a tabletop RPG, with retrieval memory
A LoRA fine-tune of Llama-3.1-8B that acts as the game master of a tabletop fantasy RPG. It narrates in second person, plays NPCs in character, applies consequences without deciding the player's choices, and uses a retrieval-augmented memory to keep the world consistent across a long session.
RPG Game Master is a LoRA fine-tune of Llama-3.1-8B that runs a tabletop fantasy RPG. Each turn it narrates the scene in second person, voices the NPCs in character, and reacts to whatever the player tries, then hands the turn back at a point where the player actually has something to do. The important design rule is that it applies consequences but never resolves the player’s intent: it will not decide whether your attack lands, speak for you, or make your choices.
The base Llama model can roleplay, but it tends to let every NPC react at once, plays along with anything, and ends every turn with a chipper “What will you do?”. The fine-tune was built to fix exactly those habits, and a retrieval memory runs alongside it so the world stays consistent even in a long session. The screenshot above is the running app: the fine-tuned GM narrating a turn on the merchant ship, with the world bible (RAG memory) filling up on the right.
The clearest way to see the difference is to give the same scene and the same player action to the fine-tuned GM and to stock Llama-3.1-8B-Instruct. The scene is a merchant ship, the Marigold, with Captain John, first mate Mara, deckhand Tom, and the cook Bess.
Player: I draw my cutlass and lunge at Captain John, shouting "The Marigold is mine now!"
Fine-tuned GM
Captain John doesn't flinch. He keeps steering with one hand and raps the back of your head with the flat of his palm, hard enough that stars flash, not enough to break bone. "Put the blade away before I put you over the rail," he says, calm, glancing once at Mara, who's already moving toward the rigging. "We're three days from shore. Three days, and no place on this boat where you can hide." The deckhands go quiet; Tom looks like he might vomit.
As you draw your cutlass and shout, the entire crew stops what they're doing to look at you. Captain John's bushy eyebrows shoot up in surprise. "By the gods, lad! What's gotten into ye?" he exclaims, taking a step back, his hand hovering near his own sword but not drawing it. Mara rushes over, her eyes wide with concern. "Easy there, boyo!" she asks. Tom drops his brush and gasps. "Oh no! Has he gone mad from the sun?" Bess pokes her head out of the galley. "What's all this commotion?" she calls out. The captain holds up a calming hand. "Now, now, there's no need for violence among friends! Why don't ye put that blade away and we'll talk this over like civilized folk, eh?" What will you do?
Every NPC on the ship reacts in the base version, the threat evaporates into friendly chatter, and it closes with the tell-tale “What will you do?”. The fine-tune keeps John in command, moves only the people who would actually move, and leaves a real threat hanging.
Player: I leap off the deck, fly up into the storm clouds, and call down a bolt of lightning to split the mast.
Fine-tuned GM
You vault off the rail and your feet meet nothing but air; there's no flying here, and no lightning answers a man who asks for it. Captain John hauls you back by the collar before you can drop far. "Watch the rail," he growls. "Mara keeps her eyes open." Mara doesn't look away from the grey line where horizon meets cloud. "If we lose that mast in a squall, I'd rather know why than watch him flapping about the sky." Tom gapes after you like you're mad, which maybe you are.
You leap dramatically off the deck, reaching your arms toward the swirling storm clouds above, willing the lightning to obey your command! For a moment, the wind howls and the clouds churn ominously... but alas, no bolt of lightning answers your call. You land back on the deck with a thud, slightly winded. Mara gasps. Captain John strokes his chin thoughtfully. "Well now, that was quite the spectacle!" Tom stares, jaw hanging open. "Whoa! That was AMAZING! Can you teach me how to do that?" What will you do?
Here the fine-tune simply denies the impossible action and grounds the player back on the deck, while the base model stages it as a dramatic spectacle the crew admires. Refusing overpowered actions without breaking immersion was one of the behaviors I trained for directly.
Player: I try to convince Mara to take over the ship, whispering, "The captain's leading us to our deaths. Join me."
Fine-tuned GM
Mara glances at your mouth, then out past the rail to where Captain John stands by the wheel. She doesn't answer right away; her hand drifts once to her sword-hilt before letting go. "He hasn't steered wrong yet," she says finally. "Not like some I could name who've lost ships off these same shoals." Her gaze comes back to you. "And even if he was blind, the crew'd know better than to follow a stranger's word against theirs. You want my blade? Earn one of mine first."
The persuasion attempt gets a reasoned, in-character refusal that reflects loyalty and reputation, rather than a nervous “we’ll talk about this later.” The GM plays Mara’s judgment instead of caving to the player.
Every turn is a chat exchange with a fixed shape. A system message states the GM’s rules. A user message carries five labeled slots. The assistant message is the GM’s narration, which is the target during training and the output at inference.
The same five-slot format is used for both training and inference, so the retrieval memory can drop facts straight into the World state slot the model already knows how to read.
Fine-tuning uses QLoRA on top of Llama-3.1-8B-Instruct. The base model is loaded in 4-bit (NF4 with double quantization and a bfloat16 compute type) so it fits on a single GPU, and its weights stay frozen. Only small LoRA adapters are trained, with rank 16, alpha 32, and a light dropout, applied to all of the attention and MLP projection layers. That keeps the number of trainable parameters tiny compared to the full 8B model.
Training runs for 3 epochs with an effective batch size of 16 (batch of 4 with 4 gradient-accumulation steps), a 2e-4 learning rate on a cosine schedule with a short warmup, a 4096-token context, bfloat16, and gradient checkpointing. It is built with Hugging Face TRL and PEFT. The finished adapter is published to the Hugging Face hub and pulled down on first run.
The quality of a fine-tune like this lives or dies on the data, so most of the work went into building a clean training set rather than tuning hyperparameters. Examples come from four sources: transcripts from real play like Critical Role (the CRD3 dataset) and FIREBALL, fantasy character dialogue from the LIGHT dataset, and a batch of hand-written examples. There is also a dedicated bucket of proactive turns that teach the model to advance the scene on its own when the player does nothing.
Raw transcripts are noisy, full of dice talk, table chatter, and donation shout-outs, none of which belong in GM narration. To filter them, every candidate example is scored by an LLM judge against a strict rubric that rewards concrete sensory detail, second-person present-tense narration, and turns that stop where the player can react, while penalizing mechanics talk, out-of-character chatter, and incoherence. The judge scores each example from 1 to 5, and only 4s and 5s survive into the final mix. Judging runs many calls concurrently and caches every score to disk so the dataset can be rebuilt without paying to re-judge. The surviving examples are shuffled together, split into training and validation, and written out in the five-slot chat format.
Four noisy sources are filtered by a strict LLM judge, mixed, split, and used to train LoRA adapters over a frozen 4-bit Llama base.
A model only sees what fits in its context window, so over a long session the GM starts to slip: it renames an NPC, rearranges a room it already described, or forgets who died a few scenes ago. The retrieval memory fixes this. After each GM turn, a small extractor model pulls the durable, concrete facts out of what just happened, named NPCs, places, items, promises, and relationships, formatted as one “Name, short fact” per line, and drops anything generic like mood or tension. Those facts are embedded and stored in a Chroma vector database, keyed one record per entity so the newest fact about a character overwrites the old one.
On the next turn, the system embeds the current scene and player action, retrieves the most relevant stored facts, and drops them into the World state slot in the same format the model was trained on. One nice touch is a contradiction guard on relationships: the first time the memory learns that two characters are, say, siblings, it pins that kinship, and later facts that contradict it are rejected rather than silently overwriting canon. In the GUI you can watch this memory fill up live in a world bible panel as you play.
The retrieval loop that runs on every turn: pull relevant facts in, narrate, then extract and store what just happened.
At inference the fine-tuned adapter is loaded back onto the 4-bit base model and sampled with a temperature of 0.6, a repetition penalty of 1.15, and a 300-token cap per turn, which keeps the narration focused and stops it from rambling past the point where the player should take over.
Everything comes together in a Gradio interface, shown at the top of this page. You set the scene and the cast (a merchant-ship scenario is preloaded), type your action, and the GM narrates the result. There are buttons to regenerate the last turn, reset the history, and clear the memory, and a world bible panel on the side that shows the retrieval memory filling up in real time as facts about the world accumulate.
There is also a plain terminal client for quick testing. It runs the same fine-tuned model and the same retrieval memory, just as a REPL: you type at the Player> prompt, the GM narrates back, and slash commands like /facts print the current world bible, /reset clears the history, and /forget wipes the memory.
An example session in the terminal client (chat.py): the same GM and retrieval memory without the GUI, with /facts printing the world bible.
An image-to-image diffusion model that turns Pokémon artwork into pixel-art sprites
A conditional image-to-image diffusion model that turns official Pokémon front-view artwork into 96×96 pixel-art sprites. It uses a UNet with cross-attention onto an encoded copy of the artwork, trained on regular and shiny artwork/sprite pairs, so it can sprite Pokémon it has never seen.
PixelArtGen turns the front-view of official Pokémon artwork into 96×96 3DS-style pixel-art sprites. It is a conditional image-to-image diffusion model: instead of denoising from text, it denoises a 96×96 sprite while continuously looking at an encoded copy of the source artwork. Because the artwork is the condition rather than the starting image, the model can produce a proper low-resolution sprite with hard pixel edges rather than just downscaling the original.
I trained it on paired official artwork and game sprites, using both regular and shiny forms, so it learns the mapping from a clean illustration to the blocky in-game look and can generate sprites for Pokémon that have no official sprite yet.
The real test is generalization to Pokémon that were never in the training set. I fed the model the freshly released official artwork for the three Gen 10 starters. Nothing about these Pokémon was seen during training, so every sprite below is produced purely from the artwork on the left.
Browt (grass), with the official artwork on the left and three sprites generated from it.
Pombon (fire), with the official artwork on the left and three sprites generated from it.
Gecqua (water), with the official artwork on the left and three sprites generated from it.
The core is a conditional UNet that works at the 96×96 sprite resolution and predicts the noise added to a sprite at a given timestep. It has three resolution levels with channel widths of 64, 128, and 256, two residual blocks per level, and a bottleneck at the smallest spatial size.
A few choices inside the residual blocks are deliberate. I use GroupNorm instead of BatchNorm because in diffusion every image in a batch carries a different amount of noise, so batch statistics are misleading, whereas group statistics stay stable per sample. Activations are SiLU rather than ReLU because it is smooth everywhere, which suits the regression-style noise prediction, and each block carries a small dropout to fight overfitting on a fairly small dataset. The timestep is turned into a sinusoidal embedding, passed through a small MLP, and injected into every residual block as a per-channel bias so the network always knows how noisy the current sprite is.
Conditioning on the artwork. A separate CNN encoder compresses the 256×256 artwork down to a 16×16 feature map with 256 channels. That encoded artwork is then fed into the UNet through cross-attention, which is the piece that made the project actually work. In each cross-attention block the queries come from the sprite being denoised, while the keys and values come from the artwork features. In other words, every location in the sprite gets to ask the artwork “what belongs here,” and pull back the matching shape and color. I place cross-attention at the deepest downsampling level, the bottleneck, and the first upsampling level, where the spatial size is small enough that attention is cheap. Multi-head self-attention runs alongside it at the same levels.
Keeping edges crisp. On the decoder side I upsample with nearest-neighbor followed by a convolution instead of bilinear upsampling or a transposed conv. Bilinear smooths everything, which is exactly wrong for pixel art, whereas nearest-neighbor preserves the hard blocky edges and lets the following conv clean up the result.
The artwork is encoded once, then read by the sprite through cross-attention at the deepest blocks and the bottleneck.
Data is built by pairing each official artwork with its game sprite by Pokédex ID, and I include shiny forms as additional pairs, which roughly doubles the data and forces the model to key off shape rather than memorizing a single color per Pokémon. Artwork is resized to 256×256 with bicubic interpolation so it stays clean, while the target sprite is resized to 96×96 with nearest-neighbor so its pixels stay intact. Transparent backgrounds are composited onto solid black, and everything is normalized to the range -1 to 1. Data is split 90/10 into train and validation by Pokémon, so validation Pokémon are never seen during training.
Augmentation is paired carefully. Horizontal flips are applied to the artwork and its sprite together so they stay aligned, but color jitter (brightness, contrast, and saturation) is applied only to the input artwork and never the target. That teaches the model to be robust to lighting and color shifts in the input without ever corrupting the sprite it is supposed to reproduce.
The diffusion process uses a continuous-time cosine schedule, where the signal and noise rates are the cosine and sine of the timestep. Timesteps are sampled uniformly in the range 0 to 1, noise is added to the sprite, and the model is trained to predict that noise with an L1 loss, which is more forgiving of the occasional hard outlier pixel than an L2 loss. Optimization is AdamW at a 2e-4 learning rate with weight decay, a linear warmup over the first 1000 steps followed by a cosine decay down to a tenth of the base rate, and gradient clipping for stability. I keep an exponential moving average of the weights with a decay of 0.999, and it is that EMA copy, not the raw training weights, that gets used for all sampling.
Each step noises a sprite to a random level, asks the UNet to predict that noise from the noisy sprite plus the artwork, and trains on the L1 gap. A slow EMA copy of the weights is what actually gets used to generate.
At inference the model starts from pure Gaussian noise at 96×96 and runs a deterministic DDIM sampler. At each step it predicts the noise, reconstructs an estimate of the clean sprite, clamps it to the valid range, and takes a deterministic step toward the next, lower noise level. The whole trajectory is conditioned on the same encoded artwork the entire way down. Sampling defaults to 50 steps but works anywhere from 10 to 100, trading speed for quality.
To produce variety, I generate several samples at once from the same artwork by starting each one from a different random noise seed. The artwork condition is shared, so every sample is a valid sprite of the same Pokémon, but small differences in shading and posture appear across them.
A worked example: Browt's artwork conditions a chain of denoising steps that turns pure noise into its sprite.
I wrapped the model in a Gradio interface called PixelForge so it is usable without touching the code. You upload official front-view artwork, set the number of diffusion steps and how many samples you want, and it returns a strip of generated sprites. It works best with artwork at 256×256 or larger, since smaller inputs get upscaled and lose detail. The interface also reproduces the training preprocessing at upload time: it keys out the near-white background to transparency, composites onto black, and resizes to 256×256, so what the model sees at inference matches what it saw during training. Outputs are shown upscaled with nearest-neighbor so the pixel grid stays sharp on screen.
PixelForge: upload artwork, tune the diffusion settings, and generate a strip of sprites.
Two problems dominated the project. The first was spatial awareness. Early versions produced sprites with roughly the right colors but no sense of where the ears, legs, arms, and head actually belonged, because the model was leaning on a single compressed summary of the artwork. Adding cross-attention, so the sprite could reference the full artwork feature map throughout denoising, fixed most of the misplaced-anatomy problems. The second was sharpness: the first outputs were soft and blurry, the opposite of pixel art. Switching the decoder to nearest-neighbor upsampling followed by a convolution, and resizing the sprite targets with nearest-neighbor rather than a smoothing filter, brought back the crisp blocky edges that make it read as a real sprite.
A generalizable bimanual policy for the LeHome Challenge 2026
A bimanual robot policy that folds clothes from arbitrary crumpled states. Trained on simulation for the LeHome Challenge 2026 and handmade demos with benchmarking across Diffusion Policy, ACT, and SmolVLA.
This is our entry for the LeHome Challenge 2026, a competition for robotic garment manipulation. The goal is a single policy that takes a piece of clothing from an arbitrary initial state, crumpled, twisted, or flat, and manipulates it into a correctly folded state using a pair of bimanual arms. Hardware is the LeRobot SO-ARM101, and submissions are scored by a mesh-keypoint pairwise-proximity metric against a target fold.
This was a team project: Conor Hayes, Andnet DeBoer, and myself, run as a Northwestern CS396 final project and robotics research project.
Our best policy, SmolVLA, folding a garment in the LeHome Isaac Sim environment
Our pipeline has three stages:
Collecting real teleoperation episodes by hand is slow, so we used NVIDIA Cosmos to grow our dataset. Cosmos takes one of our clips and re-renders it in a completely different setting while keeping the robot’s motion exactly the same. From a single demo we can spin up dozens of variants with new backgrounds, lighting, colors, and textures. All that variety is what teaches the policy to handle scenes it never actually saw while training.
An example of our NVIDIA Cosmos augmentation. The same manipulation, re-rendered with a different background and lighting.
We trained and tested three imitation-learning policies: SmolVLA, ACT, and Diffusion Policy. SmolVLA, shown in the overview above, and ACT both folded the garment cleanly and did it the same way every time. Diffusion Policy struggled. It usually bunched or dragged the cloth instead of folding it, and once the fabric ended up in an odd shape it almost never recovered.
ACT folds cleanly and decisively.
Diffusion Policy often leaves the garment bunched instead of folded.
So why the gap? ACT plans a short burst of motion at a time, which fits the smooth, repetitive rhythm of folding really well. SmolVLA leans on a big pre-trained vision-language model, so it handles new starting positions with ease. Diffusion Policy builds each move through many small refinement steps, and that process gets fragile when the cloth is partly hidden or hard to read. A small mistake early on snowballs into the messy folds you see here.
Submissions are scored with a mesh-keypoint metric. We place a set of check points on the garment, and a fold passes a check when each point lands close enough to where it should be. The more points that match their targets, the higher the score.
How the scoring works. Each check point on the garment has to land close to its target position.
Across our evaluation episodes, the three policies landed right where the folds above suggested:
| Policy | Success rate | Fold quality |
|---|---|---|
| SmolVLA | 69.44% | 8 / 10 |
| ACT | 61.11% | 7 / 10 |
| Diffusion | 41.67% | 2 / 10 |
SmolVLA came out on top, ACT was close behind, and Diffusion Policy trailed well back. That lines up with what the clips show.
Learning from direct human demonstrations
Teaching robots to grasp delicate objects by learning from human hand demonstrations with tactile feedback.
Hand2Rob teaches a Franka robot to grasp delicate objects by watching human hand demonstrations captured with a stereo camera pair. MediaPipe and CoTracker extract hand and object keypoints that get triangulated into 3D trajectories the robot can follow directly. The catch is that spatial imitation alone isn’t enough for fragile things, so I integrated a ResKin tactile sensor into custom gripper fingertips to give the robot a sense of force.
This project builds on Point Policy and Feel The Force by Siddhant Haldar and Lerrel Pinto. Big thanks to them for the original work. I adapted and extended both systems with my own changes to get them running on the Franka Panda, including modifications to the trajectory execution, force control integration, data collection, and the custom gripper hardware.
Without force feedback
With ResKin force control
Without force feedback, the robot has no way to modulate its grip strength, it simply closes the gripper until the binary close command is fully executed. For fragile objects like eggs, this can lead to it breaking. Here the robot successfully reaches and grasps the egg using the learned trajectory, but the uncontrolled gripper force crushes it. This failure motivates the need for closed-loop force control during the grasp phase.
With the ResKin sensor in the loop, the robot can feel how much force it is applying and stop closing once a target threshold is reached. The force controller reads the live tactile signal and adjusts the gripper incrementally, holding the egg securely without exceeding the pressure that would crack it. This demonstrates that combining learned spatial policies with real-time tactile feedback enables safe manipulation of objects that would otherwise be damaged.
Above are two evaluation runs showing the full pipeline end-to-end, the robot approaches the egg, descends to the grasp position, and closes with force-controlled grip. Each camera view is shown separately, with the live force reading overlaid in blue. The consistency across runs demonstrates that the learned policy generalizes reliably from the human demonstrations.
Figure 1. Overview of the Hand2Rob pipeline. Stereo camera footage and MediaPipe hand tracking are processed through CoTracker and stereo triangulation to build the dataset. A training policy learns both the end-effector trajectory and force grasping, which are deployed on the Franka robot with ResKin tactile feedback for manipulating fragile objects.
Gripper without sensor
Gripper with ResKin tactile sensor
Special thanks to Miguel Pegues for helping me design custom gripper fingertips to house the ResKin magnetometer-based tactile sensor. The left model shows the standard Franka gripper fingers, while the right model integrates a recessed pocket that secures the ResKin sensing pad flush against the contact surface.
A robot arm that detects a pen using a camera and autonomously grasps it with accurate positioning.
This project implements a vision-guided grasping system for a robotic arm that autonomously detects and grasps a pen using an RGB-D camera. The objective was to build a complete perception-to-action pipeline that converts raw camera data into executable robot motion, enabling the robot to locate and grasp a pen without manual alignment or human intervention.
The task was intentionally constrained to a known object type and workspace, allowing the system to emphasize robustness, accuracy, and correct geometric reasoning rather than general-purpose object recognition.
Perception was implemented using classical computer vision techniques. Since the target object was a purple pen, the RGB image was converted to the HSV color space and color thresholding was applied to segment purple regions from the background. Depth data from the RealSense camera was used to remove background pixels beyond a fixed range, improving robustness under clutter and lighting variation.
Contours were extracted from the segmented mask, and the most relevant contour was selected based on geometric properties. From this contour, the pen’s image-space centroid and orientation were estimated. The centroid pixel was aligned with the depth image, and the corresponding depth value was used to deproject the pixel into a 3D point in the camera coordinate frame using the camera intrinsics. To reduce sensor noise, multiple measurements were collected over a short time window and averaged.
The 3D pen position estimated in the camera frame was transformed into the robot base frame using a precomputed camera-to-robot extrinsic calibration. This calibration was represented as a rigid-body transform consisting of a rotation matrix (R) and translation vector (t), applied directly as:
Probot = R · Pcamera + t
A small tool offset was then added to account for the physical geometry of the gripper.
Once the target position was expressed in the robot frame, the PincherX 100 arm was controlled using direct API commands. The robot moved to a hover pose above the pen, descended to the grasp location, closed the gripper, lifted the object to verify a successful grasp, and returned to a safe pose. This demonstrated a complete vision-driven manipulation pipeline operating under real sensor noise and hardware constraints.
Camera-guided grasping with a Franka robot arm
A robot arm that sees objects with a camera and automatically picks them up and places them in the correct location and orientation without human control.
OmniPlace is an autonomous pick and place system built on ROS 2 for a Franka Panda with an Intel RealSense camera. The robot scans a tabletop, detects both objects (squares, rectangles, cylinders) and targets (their flat cross sections), matches each object to the correct target, and executes pick and place until no targets remain.
Hardware setup: objects, camera, and robot arm
Software pipeline from sensing to planning and execution
Perception is handled by a YOLO-based detector that outputs oriented bounding boxes, providing both object position and in-plane rotation directly from the image. Rather than relying on single-frame detections, the system performs a structured scan of the workspace and aggregates detections across multiple frames.

Detections are associated over time by comparing center location, shape class, bounding box dimensions, and orientation. Candidates that do not appear consistently across the scan window are discarded as noise. This temporal filtering produces a stable set of object and target poses that can be safely used for motion planning.
To train the YOLO model, a Python-based synthetic data generation pipeline was developed. The script programmatically generated large datasets of labeled images containing geometric objects with randomized pose, scale, orientation, lighting, and background variation. This approach made it possible to rapidly scale the training dataset without manual labeling and ensured strong coverage of object orientations required for reliable OBB prediction.
Objects and targets are distinguished using a height-based heuristic. Targets are flat cross-sections placed on the table, while objects have nonzero height. This separation allows the system to independently build object and target sets and perform shape-based matching between them.
Motion planning is handled through a custom Python interface built on top of MoveIt 2, designed to simplify interaction with the robot while still exposing fine control when needed. Instead of calling MoveIt APIs directly throughout the codebase, the system is structured around a small set of modular wrappers that separate state queries, planning logic, and environment management.
The motion planning interface provides utilities for querying the robot’s current state, generating collision-aware trajectories, and executing both Cartesian and joint-space motions. A dedicated planning scene manager dynamically adds and removes collision objects corresponding to detected items and targets, ensuring that planned motions remain valid as the workspace changes. All motion commands are funneled through a single high-level interface that exposes actions such as moving to poses, executing grasps, and controlling the gripper, keeping task logic clean and readable.
Task execution is coordinated by a central control script that drives the full pick-and-place pipeline. When triggered, the robot first moves to a known home configuration and performs camera calibration using an ArUco marker to establish consistent transforms between the camera, marker, and robot base. The perception pipeline then scans the workspace, producing a stable set of object and target poses that are visualized in RViz and added to the planning scene.
Objects are matched to targets based on shape and size, and the robot executes pick-and-place operations one pair at a time. After each attempt, the system re-scans the workspace rather than assuming a static scene. This design allows the robot to recover from failed grasps, handle objects being moved during execution, and remain robust to detection ordering changes. The task continues until no valid targets remain, at which point the robot safely returns to its home position.

A physics-based simulation of a jack-in-the-box that realistically models bouncing and collisions using real-world motion principles.
This project simulates a jack-in-the-box mechanism to study the interaction between rigid-body motion and internal impacts. The system’s motion is computed from analytical dynamics, with collisions between the jack and the box resolved using impulse-based methods.
The result is a compact simulation that highlights the role of coordinate frames, equations of motion, and contact dynamics in mechanical behavior.
The system is represented using multiple coordinate frames. A fixed world frame defines the inertial reference. A box frame moves and rotates with the enclosure. A jack frame rotates inside the box and defines the positions of the four tip masses. Points are mapped between frames using rigid-body transformations of the form
pworld = R · pbody + t
Expressing the jack in the box frame simplifies contact detection, since the box walls are fixed in that frame.
The motion of the box and jack is derived using the Euler–Lagrange formulation,
d/dt(∂L/∂q̇) − ∂L/∂q = Q
which governs smooth translational and rotational motion between collisions. These equations are numerically integrated forward in time.
When a jack tip contacts a wall, the collision is treated as an instantaneous event. Instead of integrating through contact, velocities are updated using an impulse-based formulation. The post-impact generalized velocities satisfy
∂L/∂q̇⁺ = ∂L/∂q̇⁻ + Jᵀλ
where (J) is the contact Jacobian and (λ) is the impulse magnitude. The impulse is solved by enforcing the contact constraint together with an energy consistency condition, resulting in realistic momentum transfer between the internal jack and the box. These repeated impacts produce rotation, bouncing, and jitter of the enclosure.

Building a full SLAM stack from scratch
A TurtleBot3 that maps its environment and localizes itself in real time using EKF SLAM, built from scratch in ROS 2.
Red = ground truth · Blue = odometry · Green = SLAM estimate and landmark map
This project builds feature-based EKF SLAM on a TurtleBot3 from scratch. Five ROS 2 packages make up the system:
diff_params.yaml for wheel radius, track width, and collision geometry.The odometry node integrates wheel encoder deltas using the constant-curvature arc model from the DiffDrive class. It publishes nav_msgs/Odometry and broadcasts the odom → base_footprint TF, with an initial_pose service to reset the origin.
Physical TurtleBot3 driving in a circle
Odometry estimate visualized in RViz
Landmarks are cylinders detected from 2D lidar scans in three stages:
The EKF maintains a joint state vector [θ, x, y, mx₁, my₁, …, mxN, myN] and a single covariance matrix over robot pose and all landmark positions simultaneously.