An image-to-image diffusion model that turns Pokémon artwork into pixel-art sprites
A conditional image-to-image diffusion model that turns official Pokémon front-view artwork into 96×96 pixel-art sprites. It uses a UNet with cross-attention onto an encoded copy of the artwork, trained on regular and shiny artwork/sprite pairs, so it can sprite Pokémon it has never seen.
PixelArtGen turns the front-view of official Pokémon artwork into 96×96 3DS-style pixel-art sprites. It is a conditional image-to-image diffusion model: instead of denoising from text, it denoises a 96×96 sprite while continuously looking at an encoded copy of the source artwork. Because the artwork is the condition rather than the starting image, the model can produce a proper low-resolution sprite with hard pixel edges rather than just downscaling the original.
I trained it on paired official artwork and game sprites, using both regular and shiny forms, so it learns the mapping from a clean illustration to the blocky in-game look and can generate sprites for Pokémon that have no official sprite yet.
The real test is generalization to Pokémon that were never in the training set. I fed the model the freshly released official artwork for the three Gen 10 starters. Nothing about these Pokémon was seen during training, so every sprite below is produced purely from the artwork on the left.
Browt (grass), with the official artwork on the left and three sprites generated from it.
Pombon (fire), with the official artwork on the left and three sprites generated from it.
Gecqua (water), with the official artwork on the left and three sprites generated from it.
The core is a conditional UNet that works at the 96×96 sprite resolution and predicts the noise added to a sprite at a given timestep. It has three resolution levels with channel widths of 64, 128, and 256, two residual blocks per level, and a bottleneck at the smallest spatial size.
A few choices inside the residual blocks are deliberate. I use GroupNorm instead of BatchNorm because in diffusion every image in a batch carries a different amount of noise, so batch statistics are misleading, whereas group statistics stay stable per sample. Activations are SiLU rather than ReLU because it is smooth everywhere, which suits the regression-style noise prediction, and each block carries a small dropout to fight overfitting on a fairly small dataset. The timestep is turned into a sinusoidal embedding, passed through a small MLP, and injected into every residual block as a per-channel bias so the network always knows how noisy the current sprite is.
Conditioning on the artwork. A separate CNN encoder compresses the 256×256 artwork down to a 16×16 feature map with 256 channels. That encoded artwork is then fed into the UNet through cross-attention, which is the piece that made the project actually work. In each cross-attention block the queries come from the sprite being denoised, while the keys and values come from the artwork features. In other words, every location in the sprite gets to ask the artwork “what belongs here,” and pull back the matching shape and color. I place cross-attention at the deepest downsampling level, the bottleneck, and the first upsampling level, where the spatial size is small enough that attention is cheap. Multi-head self-attention runs alongside it at the same levels.
Keeping edges crisp. On the decoder side I upsample with nearest-neighbor followed by a convolution instead of bilinear upsampling or a transposed conv. Bilinear smooths everything, which is exactly wrong for pixel art, whereas nearest-neighbor preserves the hard blocky edges and lets the following conv clean up the result.
The artwork is encoded once, then read by the sprite through cross-attention at the deepest blocks and the bottleneck.
Data is built by pairing each official artwork with its game sprite by Pokédex ID, and I include shiny forms as additional pairs, which roughly doubles the data and forces the model to key off shape rather than memorizing a single color per Pokémon. Artwork is resized to 256×256 with bicubic interpolation so it stays clean, while the target sprite is resized to 96×96 with nearest-neighbor so its pixels stay intact. Transparent backgrounds are composited onto solid black, and everything is normalized to the range -1 to 1. Data is split 90/10 into train and validation by Pokémon, so validation Pokémon are never seen during training.
Augmentation is paired carefully. Horizontal flips are applied to the artwork and its sprite together so they stay aligned, but color jitter (brightness, contrast, and saturation) is applied only to the input artwork and never the target. That teaches the model to be robust to lighting and color shifts in the input without ever corrupting the sprite it is supposed to reproduce.
The diffusion process uses a continuous-time cosine schedule, where the signal and noise rates are the cosine and sine of the timestep. Timesteps are sampled uniformly in the range 0 to 1, noise is added to the sprite, and the model is trained to predict that noise with an L1 loss, which is more forgiving of the occasional hard outlier pixel than an L2 loss. Optimization is AdamW at a 2e-4 learning rate with weight decay, a linear warmup over the first 1000 steps followed by a cosine decay down to a tenth of the base rate, and gradient clipping for stability. I keep an exponential moving average of the weights with a decay of 0.999, and it is that EMA copy, not the raw training weights, that gets used for all sampling.
Each step noises a sprite to a random level, asks the UNet to predict that noise from the noisy sprite plus the artwork, and trains on the L1 gap. A slow EMA copy of the weights is what actually gets used to generate.
At inference the model starts from pure Gaussian noise at 96×96 and runs a deterministic DDIM sampler. At each step it predicts the noise, reconstructs an estimate of the clean sprite, clamps it to the valid range, and takes a deterministic step toward the next, lower noise level. The whole trajectory is conditioned on the same encoded artwork the entire way down. Sampling defaults to 50 steps but works anywhere from 10 to 100, trading speed for quality.
To produce variety, I generate several samples at once from the same artwork by starting each one from a different random noise seed. The artwork condition is shared, so every sample is a valid sprite of the same Pokémon, but small differences in shading and posture appear across them.
A worked example: Browt's artwork conditions a chain of denoising steps that turns pure noise into its sprite.
I wrapped the model in a Gradio interface called PixelForge so it is usable without touching the code. You upload official front-view artwork, set the number of diffusion steps and how many samples you want, and it returns a strip of generated sprites. It works best with artwork at 256×256 or larger, since smaller inputs get upscaled and lose detail. The interface also reproduces the training preprocessing at upload time: it keys out the near-white background to transparency, composites onto black, and resizes to 256×256, so what the model sees at inference matches what it saw during training. Outputs are shown upscaled with nearest-neighbor so the pixel grid stays sharp on screen.
PixelForge: upload artwork, tune the diffusion settings, and generate a strip of sprites.
Two problems dominated the project. The first was spatial awareness. Early versions produced sprites with roughly the right colors but no sense of where the ears, legs, arms, and head actually belonged, because the model was leaning on a single compressed summary of the artwork. Adding cross-attention, so the sprite could reference the full artwork feature map throughout denoising, fixed most of the misplaced-anatomy problems. The second was sharpness: the first outputs were soft and blurry, the opposite of pixel art. Switching the decoder to nearest-neighbor upsampling followed by a convolution, and resizing the sprite targets with nearest-neighbor rather than a smoothing filter, brought back the crisp blocky edges that make it read as a real sprite.