Metadata-Version: 2.4
Name: asr-attacks
Version: 0.3.0
Summary: White-box adversarial attacks for CTC automatic speech recognition models.
Author-email: Hammad Ali Khan <114487807+hammaad2002@users.noreply.github.com>
License-Expression: Apache-2.0
Project-URL: Homepage, https://github.com/hammaad2002/ASRAdversarialAttacks
Project-URL: Documentation, https://hammaad2002.github.io/ASRAdversarialAttacks/
Project-URL: Changelog, https://github.com/hammaad2002/ASRAdversarialAttacks/blob/main/CHANGELOG.md
Project-URL: Issues, https://github.com/hammaad2002/ASRAdversarialAttacks/issues
Project-URL: Source, https://github.com/hammaad2002/ASRAdversarialAttacks
Keywords: asr,adversarial-attacks,speech,robustness,wav2vec2,ctc
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Typing :: Typed
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
License-File: NOTICE
Requires-Dist: numpy>=1.22
Requires-Dist: torch>=2.0
Requires-Dist: tqdm>=4.64
Requires-Dist: jiwer>=3.0
Requires-Dist: scipy>=1.10
Provides-Extra: hf
Requires-Dist: transformers>=4.30; extra == "hf"
Provides-Extra: wav2vec2
Requires-Dist: torchaudio>=2.0; extra == "wav2vec2"
Provides-Extra: rooms
Requires-Dist: pyroomacoustics>=0.7; extra == "rooms"
Provides-Extra: benchmark
Requires-Dist: torchaudio>=2.0; extra == "benchmark"
Requires-Dist: huggingface_hub>=0.20; extra == "benchmark"
Requires-Dist: pyarrow>=14; extra == "benchmark"
Requires-Dist: soundfile>=0.12; extra == "benchmark"
Provides-Extra: dev
Requires-Dist: pytest>=7.4; extra == "dev"
Requires-Dist: pytest-cov>=5; extra == "dev"
Requires-Dist: mypy>=1.10; extra == "dev"
Requires-Dist: ruff<0.17,>=0.16.7; extra == "dev"
Requires-Dist: pre-commit>=3.5; extra == "dev"
Provides-Extra: docs
Requires-Dist: mkdocs<2,>=1.6; extra == "docs"
Requires-Dist: mkdocs-material>=9.5; extra == "docs"
Requires-Dist: mkdocstrings[python]>=0.26; extra == "docs"
Provides-Extra: build
Requires-Dist: build>=1.2; extra == "build"
Requires-Dist: twine>=5.1; extra == "build"
Dynamic: license-file

# ASR Adversarial Attacks

[![CI](https://github.com/hammaad2002/ASRAdversarialAttacks/actions/workflows/ci.yml/badge.svg)](https://github.com/hammaad2002/ASRAdversarialAttacks/actions/workflows/ci.yml)
[![codecov](https://codecov.io/gh/hammaad2002/ASRAdversarialAttacks/graph/badge.svg)](https://codecov.io/gh/hammaad2002/ASRAdversarialAttacks)
[![PyPI](https://img.shields.io/pypi/v/asr-attacks)](https://pypi.org/project/asr-attacks/)
[![Python](https://img.shields.io/pypi/pyversions/asr-attacks)](https://pypi.org/project/asr-attacks/)
[![License](https://img.shields.io/pypi/l/asr-attacks)](LICENSE)
[![Docs](https://github.com/hammaad2002/ASRAdversarialAttacks/actions/workflows/docs.yml/badge.svg)](https://hammaad2002.github.io/ASRAdversarialAttacks/)
[![pre-commit](https://img.shields.io/badge/pre--commit-enabled-brightgreen?logo=pre-commit)](https://github.com/pre-commit/pre-commit)
[![Ruff](https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/ruff/main/assets/badge/v2.json)](https://github.com/astral-sh/ruff)

This package tests an ASR (speech-to-text) model against well-known white-box adversarial attacks.

**Docs:** [hammaad2002.github.io/ASRAdversarialAttacks](https://hammaad2002.github.io/ASRAdversarialAttacks/)

## Install

```bash
pip install asr-attacks[wav2vec2]
```

The development version:

```bash
pip install "asr-attacks[wav2vec2] @ git+https://github.com/hammaad2002/ASRAdversarialAttacks.git"
```

| Extra | For |
| --- | --- |
| `wav2vec2` | torchaudio wav2vec2 pipelines |
| `hf` | Hugging Face CTC models |
| `rooms` | pyroomacoustics, to generate the rooms of `RoomSimulator` (`mode="robust"`) |
| `benchmark` | `scripts/benchmark_wav2vec2.py` |
| `dev` | pytest, pytest-cov, mypy, ruff, pre-commit |
| `docs` | MkDocs |
| `build` | `python -m build` and twine |

Python 3.10+ and PyTorch 2.x are required. Install a CUDA/CPU torch wheel from [pytorch.org](https://pytorch.org/get-started/locally/) first if you need a specific build.

## Supported attacks

Every attack supports targeted and untargeted modes (untargeted: model prediction
or `label=`). Untargeted Imperceptible is a package extension.

| Attack | Paper | Notes |
| --- | --- | --- |
| FGSM | [Goodfellow et al., 2015](https://arxiv.org/abs/1412.6572) | Single-step; `norm` in `{inf, 2, 1}` |
| BIM | [Kurakin et al., 2017](https://arxiv.org/abs/1607.02533) | Iterative; `norm` in `{inf, 2, 1}` |
| PGD | [Madry et al., 2018](https://arxiv.org/abs/1706.06083) | Random start + optional `restarts`; `norm` in `{inf, 2, 1}` |
| CW | [Carlini & Wagner, 2018](https://arxiv.org/abs/1801.01944) | Audio C&W Sec. III-B/C/F (`\|\|δ\|\|_2² + c·CTC`) |
| Imperceptible | [Qin et al., 2019](https://arxiv.org/abs/1903.10346) | Offline by default; `mode="robust"` needs a `RoomSimulator` |

How they behave on a real model (wav2vec2 on LibriSpeech, success rate, distortion and
run time) is in the [benchmark](https://hammaad2002.github.io/ASRAdversarialAttacks/guide/benchmark/).

## Quickstart

```python
import torch
import torchaudio
from asr_attacks import ASRAttacker, CTCModuleBackend

bundle = torchaudio.pipelines.WAV2VEC2_ASR_BASE_960H
model = bundle.get_model()
device = "cuda" if torch.cuda.is_available() else "cpu"

backend = CTCModuleBackend(model, labels=list(bundle.get_labels()), device=device)
attacker = ASRAttacker(backend)

waveform, sample_rate = torchaudio.load("example.wav")
assert sample_rate == 16000

print("clean:", attacker.decode(waveform))
adv = attacker.fgsm(waveform, epsilon=0.01, targeted=False)
adv = attacker.fgsm(waveform, epsilon=0.01, label="THE CAT", norm=2, targeted=False)
print("adversarial:", attacker.decode(adv))
```

Targeted BIM:

```python
adv = attacker.bim(
    waveform,
    target="THE CAT SAT",
    epsilon=0.03,
    alpha=0.005,
    num_iter=50,
    targeted=True,
    early_stop=True,
    nested=False,
)
mean_wer, counts = attacker.wer(["THE CAT SAT"], [adv])
```


Hugging Face CTC:

```python
from asr_attacks import ASRAttacker, HuggingFaceCTCBackend

backend = HuggingFaceCTCBackend("facebook/wav2vec2-base-960h", device="cpu")
attacker = ASRAttacker(backend)
```

Alternatively, wrap a CTC module with `ASRAttacks`:

```python
from asr_attacks import ASRAttacks

attacks = ASRAttacks(model, "cpu", list(bundle.get_labels()))
adv = attacks.FGSM_ATTACK(waveform, epsilon=0.01, targeted=False)
```

## Coverage

Tests run on Python 3.10–3.12 (Linux) and 3.12 (macOS, Windows) for every pull request,
with branch coverage reported to [Codecov](https://codecov.io/gh/hammaad2002/ASRAdversarialAttacks).

[![Coverage sunburst](https://codecov.io/gh/hammaad2002/ASRAdversarialAttacks/graphs/sunburst.svg)](https://codecov.io/gh/hammaad2002/ASRAdversarialAttacks)

## Responsible use

These methods exist to measure and defend ASR systems. Use them only on models and data you are authorized to evaluate. Do not use this project to interfere with production speech systems you do not own.

## Cite

See [`CITATION.cff`](CITATION.cff). If you use the imperceptible attack, also cite Qin et al. (2019) and IBM ART.

## License

Apache License 2.0. Psychoacoustic helpers are derived from the IBM Adversarial Robustness Toolbox; see [`NOTICE`](NOTICE).
