DDPM Sampling, Applications & Training from Scratch
CS180 Project 5
The original project spec at UC Berkeley CS180 is here.
Mirror of CS180 Project 5 Showcase Webpage.
Part A
0. Setup
2 Stages
We first use 3 prompts to let the model generate output images. Here are images and captions displayed below, with different inference steps:
- 5 steps (i.e.
num_inference_steps=5):- Size: 64px * 64px (Stage 1)
an oil painting of
a snowy mountain village
a man wearing a hat
a rocket ship - Size: 256px * 256px (Stage 2)
an oil painting of a snowy mountain village
a man wearing a hat
a rocket ship
- Size: 64px * 64px (Stage 1)
- 20 steps:
- Stage 1:
an oil painting of a
snowy mountain village
a man wearing a hat
a rocket ship - Stage 2:
an oil painting of a
snowy mountain village
a man wearing a hat
a rocket ship
- Stage 1:
- 100 steps:
- Stage 1:
an oil painting of
a snowy mountain village
a man wearing a hat
a rocket ship - Stage 2:
an oil painting of a
snowy mountain village
a man wearing a hat
a rocket ship
- Stage 1:
Reflection on the generation
We find that for 5 steps, the outputs are not so clear, specifically, the noise added are not removed so completely. We can observe lots of noisy dots in the generated images. The generated feature is also not so clear.
For 20 steps, the noise is removed, and the generated image starts to be decent. The generated images are quite close to the text prompts.
For 100 steps, the generated images are quite clear and the features are well generated, also closer to the text prompts.
Seed
We use the seed SEED=42 in this project part.
1. Sampling Loops
1.1 Implementing the Forward Process
A key part of diffusion is the forward process, which takes a clean image and adds noise to it.
\[ q(x_t | x_0) = N(x_t ; \sqrt{\bar\alpha_t} x_0, (1 - \bar\alpha_t)I)\tag{1} \]which is equivalent to an equation giving \(x_t\):
\[ x_t = \sqrt{\bar\alpha_t} x_0 + \sqrt{1 - \bar\alpha_t} \epsilon \quad \text{where}~ \epsilon \sim N(0, I) \tag{2} \]That is, given a clean image \(x_0\), we get a noisy image \( x_t \) at timestep \(t\) by sampling from a Gaussian with mean \( \sqrt{\bar\alpha_t} x_0 \) and variance \( (1 - \bar\alpha_t) \). Note that the forward process is not just adding noise – we also scale the image by \(\sqrt{\bar\alpha_t}\) and scale the noise by \(\sqrt{1-\bar\alpha_t}\). The alpha’s cumulated product is actually an equivalent from an iterative noise adding with scheduled \(\alpha_t\)’s, which is expressed as \(\bar{\alpha}_t = \prod_{s=1}^t \alpha_s\). \(\bar\alpha_t\) is close to 1 for small \(t\), and close to 0 for large \(t\).
We run the forward process on the test image with \( t \in [250, 500, 750] \). Here is the noisy images in different time steps as the results:
1.2 Classical Denoising
From the noisy images above in different time steps, we try using Gaussian blurring filters to denoise them. Respectively, the kernel size for \(t\in[250,500,750]\) is 3,5,7, and the sigma is 1.5,2.5,3.5. Here we show the trials of denoising using this classical way.
We can see the Gaussian filters denoise the images so poorly: the original noises are not eliminated, while main features and shapes of the Campanile is blurred.
1.3 One-step Denoising
Now, we try to recover \(x_0\) using UNet from \(x_t\), where \(t\in[250,500,750]\). The usage of UNet is not to directly predict \(x_0\), but to predict the added noise \(\epsilon\). We denote the noise-predicting model (UNet) as \(\epsilon_\theta(x_t,t)\), which is also conditioned on the time step \(t\) as in the expression.
Here, the expression of \(x_0\) can be directly given by the forward equation (2) above, which is the one-step denoising:
\[ x_0 = \frac{1}{\sqrt{\bar{\alpha}_t}} x_t - \frac{\sqrt{1 - \bar{\alpha}_t}}{\sqrt{\bar{\alpha}_t}} \epsilon_\theta(x_t, t)\tag{2.1} \]where \(\epsilon_\theta\) is the UNet as the noise predictor.
The one-step denoising results (the original image, the noisy image, and the estimate of the original image) are shown below.
We have seen a much better denoising performance in 1.3 i.e. one-step denoising. But when \(t\) is larger, it still goes worse: the denoised image is blurred.
1.4 Iterative Denoising
The diffusion model by an iterative denoising can solve the problem in 1.3 that for larger \(t\), the denoised image starts blurring. Though in the math, the one-step equation is somehow equivalent to the iterative scheme if the models are all (both) perfect, but in real, for a model with limited capability, the latter would be better because it tears the task apart into smaller and easier procedures.
The formula for iterative denoising to estimate the previous step of forwarding (i.e. the next iterated step in denoising) is
\[ x_{t'} = \frac{\sqrt{\bar{\alpha}_{t'} }\beta_t}{1 - \bar{\alpha}_t} x_0 + \frac{\sqrt{\alpha_t }(1 - \bar{\alpha}_{t'})}{1 - \bar{\alpha}_t} x_t + v_\sigma \]where
- \(t'\) is the previous forward step i.e. the step we are reducing to in strided timesteps (the model can skip over an amount of steps and still give decent outputs), \(t'<t\);
- \(\alpha_t={\bar\alpha_{t'}\over \bar\alpha_t}\);
- \(\beta_t = 1-\alpha_t\);
- \(v_\sigma\) is a variance term also predicted by the model in our case.
Given \(x_t\) from the last step, and \(x_0\) in this step predicted from the formula (2.1) in 1.3, we can compute \(x_{t'}\) from the formula (3). In this project, we set the start of \(t\) as 990, and the stride as 30, so that the model skips 30 steps each time and finally arrives at \(t=0\) i.e. the original image. The results of denoising are shown below.
1.5 Diffusion Model Sampling
In this part, we use another important use of diffusion models other than denoising: sampling from the real-image manifold. We feed the iterative denoising function with randomly (drawn from Gaussian) generated noises, using the prompt "a high quality photo" as a “null” prompt as a way to let the model simply do unconditional generation.
Here are 5 images from sampling from the “null” prompt:
1.6 Classifier-free Guidance
For a noise or generally, input image, we have the generation conditioned on some prompts. For the same input without conditioning, the model can estimate an unconditional noise denoted as \(\epsilon_u\), and another estimated noise conditioned on the prompt as usual denoted as \(\epsilon_c\). Note that we use a truly empty prompt for generating \(\epsilon_u\), not the “null” prompt mentioned above. Actually, the “null” prompt can be the conditioning in this case, for an unconditional generation in the outer context.
The estimate of the noise, from above, is expressed as
\[ \epsilon=\epsilon_u+\gamma(\epsilon_c-\epsilon_u)=\gamma\epsilon_c+(1-\gamma)\epsilon_u \tag 4 \]where \(\gamma\) is the scale factor, which we set as \(\gamma=7\) in this project.
Basically, this can be seen as a guidance, i.e. a push (\(\epsilon_c-\epsilon_u\)) from the unconditional point in the manifold to the conditional point, that pushes the image to have more “conditional-ness”. For example, for a dog as the conditioning, pushing this can make the image resemble a dog more, i.e. have more dog-ness.
If we set \(\gamma=1\), the push will be equivalent as that in the above section, which is shown not so effective. If \(\gamma>1\), the push will be enhanced, which is what we are doing in CFG.
Here are 5 images from sampling from the “null” prompt, with CFG at scale \(\gamma=7\):
1.7 Image-to-Image Translation
In this part, we follow SDEdit algorithm to transform one image (our inputs) to another with some conditioning. This can be done with inputting the iterative denoising pipeline with our input images, and set a t (or an equivalent index of the strided time steps i.e. i_start a.k.a. noise level), which is the forward step. t is seen as a claimed level of the noises added to the input, i.e. how much “noise” should the model “reduce” into “the original image”. The smaller the noise level is, the more t is, and the more the image is denoised (edited).
We use given noise levels [1, 3, 5,7, 10, 20] and the “null” prompt i.e. "a high quality photo" as the conditioning. Results are shown below:
Result 1: Berkeley Campanile
Result 2: Self-selected image 1: kusa.png
Result 3: Self-selected image 2: pien.png
1.7.1 Editing Hand-Drawn and Web Images
Same as above, we pick some images from the web & hand-drawn and feed them into the translation.
Result 1: Web image
Result 2: Hand-drawn image 1: A Cruise
Result 3: Hand-drawn image 2: A Lemon
1.7.2 Inpainting
Now, we implement a hole-filling (inpainting) algorithm. We use the same iterative denoising pipeline, but with a mask \(\bold m\) on the input image. Mask values i.e. values in \(\bold m\) for pixels to be inpainted are set to 1, and those for the rest (the known pixels) are set to 0. The mask is fed into the model as an additional input. Initially, we also produce a Gaussian noise as above, and we also hold the original image as \(x_{orig}\). Then, every time we iteratively denoise from \(t\) to \(t'\), we follow this formula according to this paper:
\[ x_{t'} = \bold m \odot x_{t'} + (1-\bold m) \odot \text{forward}(x_{orig},t')\tag 5 \]where \(\odot\) is the element-wise multiplication. This formula is to fill the holes in the image with the inpainted pixels from the iterative denoising. The results are shown below.
Result 1: Berkeley Campanile
1.7.3 Text-Conditional Image-to-image Translation
In this part, we do the same as in 1.7 and 1.7.1. But we use a text prompt as the conditioning. The text prompt is "a rocket ship". The results are shown below.
Result 1: Berkeley Campanile
Result 2: Self-selected image 1: kusa.png
Result 3: Self-selected image 2: pien.png
1.8 Visual Anagrams
In this part, we use the iterative denoising pipeline to generate visual anagrams (according to this research), which is basically a image that shows a feature when watched ordinarily without being transformed, and another feature when watched upside down.
We can implement this by modifying the noise estimate. One estimate from the noised image now i.e. \(x_t\) is based on \(p_1\) which is the first prompt, and another estimate from the flipped image \(\text{flip}(x_t)\) is based on \(p_2\) i.e. the second prompt. Then the estimate for the flipped image is flipped again, aligned with the direction of the ordinary view. Finally, these two estimates are averaged, and the desired estimate is outputted. The process can be expressed as:
\[ \epsilon_1 = \text{UNet}(x_t, t, p_1) \\ \epsilon_2 = \text{flip}(\text{UNet}(\text{flip}(x_t), t, p_2)) \\ \epsilon = \frac{\epsilon_1 + \epsilon_2}{2}\tag 6. \]The results are shown below.
Result 1
Result 2
Result 3
1.9 Hybrid Images
In this section, we perform the hybrid image generation, which is to generate an image that shows one feature in low frequency (far away / blurredly) and another feature in high frequency (closely / clearly), based on this paper (Factorized Diffusion). We estimate the noise by these formulas:
\[ \epsilon_1 = \text{UNet}(x_t, t, p_1) \\ \epsilon_2 = \text{UNet}(x_t, t, p_2) \\ \epsilon = f_\text{lowpass}(\epsilon_1) + f_\text{highpass}(\epsilon_2) \tag 7 \]where \(f_\text{lowpass}\) and \(f_\text{highpass}\) are the low-pass and high-pass filters, respectively.
We use a kernel size of 33 and sigma of 2 as is recommended in the project spec for the LP filter as a Gaussian filter, and the HP filter is to find the difference between the original image and the LP-filtered image, i.e. the difference between identity and the LP filter. The results are shown below.
I used the text encoder instead of predetermined .pth embeddings to get the embeddings for my DIY prompts as in Result 2 and 3.
Result 1 Low pass: a lithograph of a skull High pass: a lithograph of waterfalls
Result 2 Low pass: a salmon sushi nigiri High pass: a sitting orange cat with a white belly
Result 3 Low pass: a photo of the Ayers rock High pass: a photo of a dog lying on stomach
2. Bells & Whistles
- I used the text encoder instead of predetermined
.pthembeddings to get the embeddings for my DIY prompts as above.
2.1 A logo for the CS180 course
I designed a logo for this course, CS180, using the model stage 1 above, and also upsampled it to a higher resolution using stage 2 of the model.
The logo is a pixelated bear holding a camera, ready to taking a photo.
The logo is shown below:
Part B
1. Training a Single-Step Denoising UNet
Given a noisy image \(z\), we want to train a denoiser \(D_\theta\) with UNet so as to map \(z\) to a clean image \(x\). L2 loss is used in this training process (as well as in the whole Part B)
\[ L=E_{z,x}||D_\theta(z)-x||^2\tag8 \]1.1 Implementing the UNet

We implement an unconditional UNet as shown in the graph above, where operation blocks mentioned above are:

1.2 Using the UNet to Train a Denoiser
To train the unconditional UNet denoiser, we dynamically (not with pre-computed noises) generate \((z,x)\) pairs from clean images from the training data. The clean image drawn from the training data is \(x\), and
\[ z=x+\sigma \epsilon,\quad\epsilon\sim N(0,I)\tag9 \]We show varying levels of noise on MNIST digits, with \(\sigma=[0.0, 0.2, 0.4, 0.5, 0.6, 0.8, 1.0]\).
Now, we train the denoiser with \(\sigma=0.5\), batch size 256, 128 hidden channels (\(D=128\) where \(D\) is mentioned in the above computation graph), and an Adam optimizer with a learning rate of 1e-4 on 5 epochs.
The training loss curve is shown below.
We visualize denoised results on the test set at the end of training.
1.2.2 OOD Testing
Though the denoiser is trained where \(\sigma=0.5\), we can also perform out-of-distribution testing with a range of \(\sigma\) which is \(\sigma=[0.0, 0.2, 0.4, 0.5, 0.6, 0.8, 1.0]\).
2. Training a Diffusion Model
Now, we are to implement DDPM. We now want the UNet to predict the noise instead of the clean image, i.e. the model is \(\epsilon_\theta\) and the loss is
\[ L=E_{\epsilon,z}||\epsilon_\theta(z)-\epsilon||^2\tag{10} \]From (2) we know
\[ x_t = \sqrt{\bar\alpha_t} x_0 + \sqrt{1 - \bar\alpha_t} \epsilon \quad \text{where}~ \epsilon \sim N(0, I) \tag{2} \]for a certain time step \(t\in\{0,\cdots,T\}\) as the noise-adding process (the forward process). Because we have a varying noise level now, so we should condition the model on \(t\) to let the model work. The time-conditional diffusion model has a computation graph as follows:

where the FCBlock is

In DDPM, we also have a noise schedule which is a list of \(\beta_t,\alpha_t,\bar\alpha_t\). The relationship is:
- \(\beta_0=1e-4,\beta_T=0.02\) and \(\beta_t\)’s in between are uniformly spaced;
- \(\alpha_t=1-\beta_t\)
- \(\bar\alpha_t=\prod_{s=1}^t \alpha_s\) is a cumulative product.
2.1 Adding Time Conditioning to UNet
We add an encoded time conditioning using broadcasting to the results of an UpBlock and an Unflatten layer as shown in the graph above.
Now, the objective with time conditioning is
\[ L=E_{\epsilon,x_0,t}||\epsilon(x_t,t)-\epsilon||^2\tag{11} \]where \(x_t\) is produced in (2).
2.2 Training the Time-Conditional DDPM
The training algorithm is as follows:

In the implementation, we train the DDPM on MNIST (same in parts below) with batch size 128, 20 epochs, \(D=64\) and an Adak optimizer with an initial learning rate of 1e-3. An exponential LR decay scheduler with a gamma of \(0.1^{1/\text{n\_epochs}}\) is also used. Also, \(t\) is always normalized.
The training loss curve is shown below.
2.3 Sampling from the Time-Conditional DDPM
Following the sampling algorithm of DDPM as follows:

we can now sample from the model. We show sampling results after the 5th and 20th epoch.
2.4 Adding Class-Conditioning to UNet
We want the DDPM generate images given a specific class. To modify the UNet architecture, we can now add 2 more FCBlocks and feed them both with the one-hot class vectors which are masked to 0 with a probability \(p_{\rm uncond}=0.1\) because we still want the model to preserve the ability of unconditional generation.
When we are adding time conditioning, we now multiply the pre-addition hiddens elementwisely with the outputs of the FCBlocks passing the class signals.
We use a same set of hyperparameters as in 2.2. The class-conditional training algorithm is as follows:

The training loss curve is shown below.
2.5 Sampling from the Class-Conditional DDPM
With class conditioning, we should also use classifier-free guidance mentioned in Part A. We use CFG with a guidance scale \(\gamma=5.0\) for this part, and the sampling algorithm is as follows, where \(\epsilon_u\) is the unconditioned predicted noise and \(\epsilon_c\) is the conditioned one.

The sampling results are shown below. We can see the class signals are received very well.
3. Bells & Whistles: Improving Time-conditional UNet Architecture
For ease of explanation and implementation, our UNet architecture above is pretty simple.
I added skip connections (shortcuts) in ConvBlock, DownBlock and Upblock, which is to add a plainly convoluted input (working as the residual a.k.a. the “identity”) to the output of the block. We train with the same set of hyperparameters as in 2.2.
The improved UNet can achieve a better test loss (0.02820390514746497) than the original (0.02956294636183147).
The training loss curve is shown below.
The sampling results are shown below.
4. Bells & Whistles: Rectified Flow
Instead of DDPM, we now implement a novel SOTA framework named Rectified Flow.
Please see here.