> For the complete documentation index, see [llms.txt](https://boinc-ai.gitbook.io/accelerate/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://boinc-ai.gitbook.io/accelerate/reference/torch-wrapper-classes.md).

# Torch wrapper classes

## Wrapper classes for torch Dataloaders, Optimizers, and Schedulers

The internal classes Accelerate uses to prepare objects for distributed training when calling [prepare()](https://huggingface.co/docs/accelerate/v0.24.0/en/package_reference/accelerator#accelerate.Accelerator.prepare).

### Datasets and DataLoaders

**accelerate.data\_loader.prepare\_data\_loader**

[\<source>](https://github.com/huggingface/accelerate/blob/v0.24.0/src/accelerate/data_loader.py#L732)

( dataloader: DataLoaderdevice: typing.Optional\[torch.device] = Nonenum\_processes: typing.Optional\[int] = Noneprocess\_index: typing.Optional\[int] = Nonesplit\_batches: bool = Falseput\_on\_device: bool = Falserng\_types: typing.Union\[typing.List\[typing.Union\[str, accelerate.utils.dataclasses.RNGType]], NoneType] = Nonedispatch\_batches: typing.Optional\[bool] = Noneeven\_batches: bool = Trueslice\_fn\_for\_dispatch: typing.Optional\[typing.Callable] = None ) → `torch.utils.data.dataloader.DataLoader`

Parameters

* **dataloader** (`torch.utils.data.dataloader.DataLoader`) — The data loader to split across several devices.
* **device** (`torch.device`) — The target device for the returned `DataLoader`.
* **num\_processes** (`int`, *optional*) — The number of processes running concurrently. Will default to the value given by [AcceleratorState](https://huggingface.co/docs/accelerate/v0.24.0/en/package_reference/state#accelerate.state.AcceleratorState).
* **process\_index** (`int`, *optional*) — The index of the current process. Will default to the value given by [AcceleratorState](https://huggingface.co/docs/accelerate/v0.24.0/en/package_reference/state#accelerate.state.AcceleratorState).
* **split\_batches** (`bool`, *optional*, defaults to `False`) — Whether the resulting `DataLoader` should split the batches of the original data loader across devices or yield full batches (in which case it will yield batches starting at the `process_index`-th and advancing of `num_processes` batches at each iteration).

  Another way to see this is that the observed batch size will be the same as the initial `dataloader` if this option is set to `True`, the batch size of the initial `dataloader` multiplied by `num_processes` otherwise.

  Setting this option to `True` requires that the batch size of the `dataloader` is a round multiple of `batch_size`.
* **put\_on\_device** (`bool`, *optional*, defaults to `False`) — Whether or not to put the batches on `device` (only works if the batches are nested list, tuples or dictionaries of tensors).
* **rng\_types** (list of `str` or `RNGType`) — The list of random number generators to synchronize at the beginning of each iteration. Should be one or several of:
  * `"torch"`: the base torch random number generator
  * `"cuda"`: the CUDA random number generator (GPU only)
  * `"xla"`: the XLA random number generator (TPU only)
  * `"generator"`: the `torch.Generator` of the sampler (or batch sampler if there is no sampler in your dataloader) or of the iterable dataset (if it exists) if the underlying dataset is of that type.
* **dispatch\_batches** (`bool`, *optional*) — If set to `True`, the datalaoder prepared is only iterated through on the main process and then the batches are split and broadcast to each process. Will default to `True` when the underlying dataset is an `IterableDataset`, `False` otherwise.
* **even\_batches** (`bool`, *optional*, defaults to `True`) — If set to `True`, in cases where the total batch size across all processes does not exactly divide the dataset, samples at the start of the dataset will be duplicated so the batch can be divided equally among all workers.
* **slice\_fn\_for\_dispatch** (`Callable`, *optional*`) -- If passed, this function will be used to slice tensors across` num\_processes`. Will default to` slice\_tensors()`. This argument is used only when` dispatch\_batches`is set to`True\` and will be ignored otherwise.

Returns

`torch.utils.data.dataloader.DataLoader`

A new data loader that will yield the portion of the batches

Wraps a PyTorch `DataLoader` to generate batches for one of the processes only.

Depending on the value of the `drop_last` attribute of the `dataloader` passed, it will either stop the iteration at the first batch that would be too small / not present on all processes or loop with indices from the beginning.

`BatchSampler`s with varying batch sizes are not enabled by default. To enable this behaviour, set `even_batches` equal to `False`

**accelerate.skip\_first\_batches**

[\<source>](https://github.com/huggingface/accelerate/blob/v0.24.0/src/accelerate/data_loader.py#L985)

( dataloadernum\_batches = 0 )

Creates a `torch.utils.data.DataLoader` that will efficiently skip the first `num_batches`.

#### class accelerate.data\_loader.BatchSamplerShard

[\<source>](https://github.com/huggingface/accelerate/blob/v0.24.0/src/accelerate/data_loader.py#L102)

( \*args\*\*kwds )

Parameters

* **batch\_sampler** (`torch.utils.data.sampler.BatchSampler`) — The batch sampler to split in several shards.
* **num\_processes** (`int`, *optional*, defaults to 1) — The number of processes running concurrently.
* **process\_index** (`int`, *optional*, defaults to 0) — The index of the current process.
* **split\_batches** (`bool`, *optional*, defaults to `False`) — Whether the shards should be created by splitting a batch to give a piece of it on each process, or by yielding different full batches on each process.

  On two processes with a sampler of `[[0, 1, 2, 3], [4, 5, 6, 7]]`, this will result in:

  * the sampler on process 0 to yield `[0, 1, 2, 3]` and the sampler on process 1 to yield `[4, 5, 6, 7]` if this argument is set to `False`.
  * the sampler on process 0 to yield `[0, 1]` then `[4, 5]` and the sampler on process 1 to yield `[2, 3]` then `[6, 7]` if this argument is set to `True`.
* **even\_batches** (`bool`, *optional*, defaults to `True`) — Whether or not to loop back at the beginning of the sampler when the number of samples is not a round multiple of (original batch size / number of processes).

Wraps a PyTorch `BatchSampler` to generate batches for one of the processes only. Instances of this class will always yield a number of batches that is a round multiple of `num_processes` and that all have the same size. Depending on the value of the `drop_last` attribute of the batch sampler passed, it will either stop the iteration at the first batch that would be too small / not present on all processes or loop with indices from the beginning.

`BatchSampler`s with varying batch sizes are not enabled by default. To enable this behaviour, set `even_batches` equal to `False`

#### class accelerate.data\_loader.IterableDatasetShard

[\<source>](https://github.com/huggingface/accelerate/blob/v0.24.0/src/accelerate/data_loader.py#L255)

( \*args\*\*kwds )

Parameters

* **dataset** (`torch.utils.data.dataset.IterableDataset`) — The batch sampler to split in several shards.
* **batch\_size** (`int`, *optional*, defaults to 1) — The size of the batches per shard (if `split_batches=False`) or the size of the batches (if `split_batches=True`).
* **drop\_last** (`bool`, *optional*, defaults to `False`) — Whether or not to drop the last incomplete batch or complete the last batches by using the samples from the beginning.
* **num\_processes** (`int`, *optional*, defaults to 1) — The number of processes running concurrently.
* **process\_index** (`int`, *optional*, defaults to 0) — The index of the current process.
* **split\_batches** (`bool`, *optional*, defaults to `False`) — Whether the shards should be created by splitting a batch to give a piece of it on each process, or by yielding different full batches on each process.

  On two processes with an iterable dataset yielding of `[0, 1, 2, 3, 4, 5, 6, 7]`, this will result in:

  * the shard on process 0 to yield `[0, 1, 2, 3]` and the shard on process 1 to yield `[4, 5, 6, 7]` if this argument is set to `False`.
  * the shard on process 0 to yield `[0, 1, 4, 5]` and the sampler on process 1 to yield `[2, 3, 6, 7]` if this argument is set to `True`.

Wraps a PyTorch `IterableDataset` to generate samples for one of the processes only. Instances of this class will always yield a number of samples that is a round multiple of the actual batch size (depending of the value of `split_batches`, this is either `batch_size` or `batch_size x num_processes`). Depending on the value of the `drop_last` attribute of the batch sampler passed, it will either stop the iteration at the first batch that would be too small or loop with indices from the beginning.

#### class accelerate.data\_loader.DataLoaderShard

[\<source>](https://github.com/huggingface/accelerate/blob/v0.24.0/src/accelerate/data_loader.py#L390)

( \*args\*\*kwds )

Parameters

* **dataset** (`torch.utils.data.dataset.Dataset`) — The dataset to use to build this datalaoder.
* **device** (`torch.device`, *optional*) — If passed, the device to put all batches on.
* **rng\_types** (list of `str` or `RNGType`) — The list of random number generators to synchronize at the beginning of each iteration. Should be one or several of:
  * `"torch"`: the base torch random number generator
  * `"cuda"`: the CUDA random number generator (GPU only)
  * `"xla"`: the XLA random number generator (TPU only)
  * `"generator"`: an optional `torch.Generator`
* **synchronized\_generator** (`torch.Generator`, *optional*) — A random number generator to keep synchronized across processes.
* **skip\_batches** (`int`, *optional*, defaults to 0) — The number of batches to skip at the beginning. kwargs — All other keyword arguments to pass to the regular `DataLoader` initialization.

Subclass of a PyTorch `DataLoader` that will deal with device placement and current distributed setup.

**Available attributes:**

* **total\_batch\_size** (`int`) — Total batch size of the dataloader across all processes. Equal to the original batch size when `split_batches=True`; otherwise the original batch size \* the total number of processes
* **total\_dataset\_length** (`int`) — Total length of the inner dataset across all processes.

#### class accelerate.data\_loader.DataLoaderDispatcher

[\<source>](https://github.com/huggingface/accelerate/blob/v0.24.0/src/accelerate/data_loader.py#L543)

( \*args\*\*kwds )

Parameters

* **split\_batches** (`bool`, *optional*, defaults to `False`) — Whether the resulting `DataLoader` should split the batches of the original data loader across devices or yield full batches (in which case it will yield batches starting at the `process_index`-th and advancing of `num_processes` batches at each iteration). Another way to see this is that the observed batch size will be the same as the initial `dataloader` if this option is set to `True`, the batch size of the initial `dataloader` multiplied by `num_processes` otherwise. Setting this option to `True` requires that the batch size of the `dataloader` is a round multiple of `batch_size`.
* **skip\_batches** (`int`, *optional*, defaults to 0) — The number of batches to skip at the beginning of an iteration.

Subclass of a PyTorch `DataLoader` that will iterate and preprocess on process 0 only, then dispatch on each process their part of the batch.

**Available attributes:**

* **total\_batch\_size** (`int`) — Total batch size of the dataloader across all processes. Equal to the original batch size when `split_batches=True`; otherwise the original batch size \* the total number of processes
* **total\_dataset\_length** (`int`) — Total length of the inner dataset across all processes.

### Optimizers

#### class accelerate.optimizer.AcceleratedOptimizer

[\<source>](https://github.com/huggingface/accelerate/blob/v0.24.0/src/accelerate/optimizer.py#L38)

( optimizerdevice\_placement = Truescaler = None )

Parameters

* **optimizer** (`torch.optim.optimizer.Optimizer`) — The optimizer to wrap.
* **device\_placement** (`bool`, *optional*, defaults to `True`) — Whether or not the optimizer should handle device placement. If so, it will place the state dictionary of `optimizer` on the right device.
* **scaler** (`torch.cuda.amp.grad_scaler.GradScaler`, *optional*) — The scaler to use in the step function if training with mixed precision.

Internal wrapper around a torch optimizer.

Conditionally will perform `step` and `zero_grad` if gradients should be synchronized when performing gradient accumulation.

### Schedulers

#### class accelerate.scheduler.AcceleratedScheduler

[\<source>](https://github.com/huggingface/accelerate/blob/v0.24.0/src/accelerate/scheduler.py#L25)

( scheduleroptimizersstep\_with\_optimizer: bool = Truesplit\_batches: bool = False )

Parameters

* **scheduler** (`torch.optim.lr_scheduler._LRScheduler`) — The scheduler to wrap.
* **optimizers** (one or a list of `torch.optim.Optimizer`) — The optimizers used.
* **step\_with\_optimizer** (`bool`, *optional*, defaults to `True`) — Whether or not the scheduler should be stepped at each optimizer step.
* **split\_batches** (`bool`, *optional*, defaults to `False`) — Whether or not the dataloaders split one batch across the different processes (so batch size is the same regardless of the number of processes) or create batches on each process (so batch size is the original batch size multiplied by the number of processes).

A wrapper around a learning rate scheduler that will only step when the optimizer(s) have a training step. Useful to avoid making a scheduler step too fast when gradients went overflow and there was no training step (in mixed precision training)

When performing gradient accumulation scheduler lengths should not be changed accordingly, Accelerate will always step the scheduler to account for it.
