Metadata-Version: 2.4
Name: rapidds
Version: 0.2.1
Summary: An explainable data analysis and cleaning toolkit for Python
Author: Joel John Jobinse
License-Expression: MIT
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE.txt
Requires-Dist: numpy>=1.23
Requires-Dist: pandas>=1.5
Requires-Dist: scikit-learn>=1.2
Requires-Dist: matplotlib>=3.6
Dynamic: license-file

# 🚀 rapidds

**rapidds makes data science simpler.**

rapidds is a guided dataset companion built around a simple workflow:

> **Detect → Suggest → Execute**

It sits on top of pandas and scientific Python libraries and helps you understand what is happening in an unfamiliar dataset before you start making transformations.

## ✨ What rapidds Does

- 📊 Profile datasets automatically
- 🧠 Detect common data-quality and modeling issues
- 💡 Suggest practical next steps with severity and confidence
- 🧹 Clean data explicitly or with conservative `auto_clean()`
- 📈 Explore statistics and relationships
- 🛠 Prepare data for machine learning
- 📏 Evaluate classification, regression, and clustering models
- 📝 Produce human-readable inspection reports

## ⚡ Quick Start

```python
from rapidds import Dataset

ds = Dataset("students.csv")

print(ds.profile())
ds.explain()
```

### Guided cleaning

```python
# See what rapidds would change without modifying the data
preview, plan = ds.auto_clean(dry_run=True)
print(plan)

# Apply conservative, high-confidence fixes
cleaned = ds.auto_clean()
```

By default, `auto_clean()` can remove exact duplicates, fill straightforward missing values, convert unambiguously numeric strings, and trim whitespace. It does **not** automatically drop columns, remove outliers, or rewrite semantic labels unless you explicitly opt in.

### Data science utilities

```python
profile = ds.profile()
quality = ds.quality()
corr = ds.correlation()
outliers = ds.outliers("income")

X, y, warnings = ds.prepare_for_modeling("target")
X_train, X_test, y_train, y_test, warnings = ds.split("target", stratify=True)
```

## 🧩 Architecture

```text
Dataset
  │
  ├── analysis
  │   ├── profile
  │   ├── quality
  │   ├── statistics
  │   └── outliers
  │
  ├── cleaning
  │   ├── missing values
  │   ├── duplicates
  │   ├── type fixes
  │   └── text standardization
  │
  ├── transform
  │   ├── scaling
  │   ├── encoding
  │   ├── datetime features
  │   └── numeric transforms
  │
  ├── modeling
  │   └── evaluation
  │
  └── reporting
```

## 🎯 Philosophy

Most data science tools assume you already know what you are looking for. rapidds is designed for the earlier part of the workflow: **what is in this dataset, what looks unusual, why might it matter, and what could I do next?**

It does not try to hide decisions behind a black box. Suggestions are structured and automatic actions are auditable.

## 🛠 Installation

```bash
pip install rapidds
```

For development:

```bash
pip install -e .
pytest
```

## 📦 Version

`v0.2.0` expands the original foundation with structured profiling, data-quality checks, outlier detection, statistics, transformations, model evaluation, and conservative automated cleaning.

## 🤝 Who Is rapidds For?

- Students learning data science
- Developers prototyping quickly
- Analysts exploring unfamiliar datasets
- Data scientists who want a lightweight first-pass diagnostic layer

## 📜 License

MIT
