Quicktour
Last updated
Last updated
Diffusion models are trained to denoise random Gaussian noise step-by-step to generate a sample of interest, such as an image or audio. This has sparked a tremendous amount of interest in generative AI, and you have probably seen examples of diffusion generated images on the internet. ๐งจ Diffusers is a library aimed at making diffusion models widely accessible to everyone.
Whether youโre a developer or an everyday user, this quicktour will introduce you to ๐งจ Diffusers and help you get up and generating quickly! There are three main components of the library to know about:
The is a high-level end-to-end class designed to rapidly generate samples from pretrained diffusion models for inference.
Popular pretrained architectures and modules that can be used as building blocks for creating diffusion systems.
Many different - algorithms that control how noise is added for training, and how to generate denoised images during inference.
The quicktour will show you how to use the for inference, and then walk you through how to combine a model and scheduler to replicate whatโs happening inside the .
The quicktour is a simplified version of the introductory ๐งจ Diffusers to help you get started quickly. If you want to learn more about ๐งจ Diffusers goal, design philosophy, and additional details about itโs core API, check out the notebook!
Before you begin, make sure you have all the necessary libraries installed:
Copied
๐ speeds up model loading for inference and training.
๐ is required to run the most popular diffusion models, such as .
The is the easiest way to use a pretrained diffusion system for inference. It is an end-to-end system containing the model and the scheduler. You can use the out-of-the-box for many tasks. Take a look at the table below for some supported tasks, and for a complete list of supported tasks, check out the table.
Task
Description
Pipeline
Unconditional Image Generation
generate an image from Gaussian noise
Text-Guided Image Generation
generate an image given a text prompt
Text-Guided Image-to-Image Translation
adapt an image guided by a text prompt
Text-Guided Image-Inpainting
fill the masked part of an image given the image, the mask and a text prompt
Text-Guided Depth-to-Image Translation
adapt parts of an image guided by a text prompt while preserving structure via depth estimation
Copied
Copied
We strongly recommend running the pipeline on a GPU because the model consists of roughly 1.4 billion parameters. You can move the generator object to a GPU, just like you would in PyTorch:
Copied
Copied
Save the image by calling save
:
Copied
You can also use the pipeline locally. The only difference is you need to download the weights first:
Copied
Then load the saved weights into the pipeline:
Copied
Now you can run the pipeline as you would in the section above.
Copied
Try generating an image with the new scheduler and see if you notice a difference!
Copied
To access the model parameters, call model.config
:
Copied
The model configuration is a ๐ง frozen ๐ง dictionary, which means those parameters canโt be changed after the model is created. This is intentional and ensures that the parameters used to define the model architecture at the start remain the same, while other parameters can still be adjusted during inference.
Some of the most important parameters are:
sample_size
: the height and width dimension of the input sample.
in_channels
: the number of input channels of the input sample.
down_block_types
and up_block_types
: the type of down- and upsampling blocks used to create the UNet architecture.
block_out_channels
: the number of output channels of the downsampling blocks; also used in reverse order for the number of input channels of the upsampling blocks.
layers_per_block
: the number of ResNet blocks present in each UNet block.
To use the model for inference, create the image shape with random Gaussian noise. It should have a batch
axis because the model can receive multiple random noises, a channel
axis corresponding to the number of input channels, and a sample_size
axis for the height and width of the image:
Copied
For inference, pass the noisy image to the model and a timestep
. The timestep
indicates how noisy the input image is, with more noise at the beginning and less at the end. This helps the model determine its position in the diffusion process, whether it is closer to the start or the end. Use the sample
method to get the model output:
Copied
To generate actual examples though, youโll need a scheduler to guide the denoising process. In the next section, youโll learn how to couple a model with a scheduler.
Schedulers manage going from a noisy sample to a less noisy sample given the model output - in this case, it is the noisy_residual
.
Copied
๐ก Notice how the scheduler is instantiated from a configuration. Unlike a model, a scheduler does not have trainable weights and is parameter-free!
Some of the most important parameters are:
num_train_timesteps
: the length of the denoising process or in other words, the number of timesteps required to process random Gaussian noise into a data sample.
beta_schedule
: the type of noise schedule to use for inference and training.
beta_start
and beta_end
: the start and end noise values for the noise schedule.
Copied
The less_noisy_sample
can be passed to the next timestep
where itโll get even less noisier! Letโs bring it all together now and visualize the entire denoising process.
First, create a function that postprocesses and displays the denoised image as a PIL.Image
:
Copied
To speed up the denoising process, move the input and model to a GPU:
Copied
Now create a denoising loop that predicts the residual of the less noisy sample, and computes the less noisy sample with the scheduler:
Copied
Sit back and watch as a cat is generated from nothing but noise! ๐ป
Hopefully you generated some cool images with ๐งจ Diffusers in this quicktour! For your next steps, you can:
Start by creating an instance of a and specify which pipeline checkpoint you would like to download. You can use the for any stored on the Hugging Face Hub. In this quicktour, youโll load the checkpoint for text-to-image generation.
For models, please carefully read the first before running the model. ๐งจ Diffusers implements a to prevent offensive or harmful content, but the modelโs improved image generation capabilities can still produce potentially harmful content.
Load the model with the method:
The downloads and caches all modeling, tokenization, and scheduling components. Youโll see that the Stable Diffusion pipeline is composed of the and among other things:
Now you can pass a text prompt to the pipeline
to generate an image, and then access the denoised image. By default, the image output is wrapped in a object.
Different schedulers come with different denoising speeds and quality trade-offs. The best way to find out which one works best for you is to try them out! One of the main features of ๐งจ Diffusers is to allow you to easily switch between schedulers. For example, to replace the default with the , load it with the method:
In the next section, youโll take a closer look at the components - the model and scheduler - that make up the and learn how to use these components to generate an image of a cat.
Most models take a noisy sample, and at each timestep it predicts the noise residual (other models learn to predict the previous sample directly or the velocity or ), the difference between a less noisy image and the input image. You can mix and match models to create other diffusion systems.
Models are initiated with the method which also locally caches the model weights so it is faster the next time you load the model. For the quicktour, youโll load the , a basic unconditional image generation model with a checkpoint trained on cat images:
๐งจ Diffusers is a toolbox for building diffusion systems. While the is a convenient way to get started with a pre-built diffusion system, you can also choose your own model and scheduler components separately to build a custom diffusion system.
For the quicktour, youโll instantiate the with itโs method:
To predict a slightly less noisy image, pass the following to the schedulerโs method: model output, timestep
, and current sample
.
Train or finetune a model to generate your own images in the tutorial.
See example official and community for a variety of use cases.
Learn more about loading, accessing, changing and comparing schedulers in the guide.
Explore prompt engineering, speed and memory optimizations, and tips and tricks for generating higher quality images with the guide.
Dive deeper into speeding up ๐งจ Diffusers with guides on , and inference guides for running and .