Initial project setup: NF Hotel Data Analysis API

- Add project structure with domain-driven design organization
- Implement FastAPI endpoints for hotel booking reports
- Add data cleaning, statistics, and LLM report generation services
- Include configuration management and security utilities
- Add repository layer for bookings and metadata access
- Setup testing framework with conftest.py
- Include example files and documentation
This commit is contained in:
2026-09-21 13:17:49 +08:00
commit 225d28feb7
39 changed files with 41693 additions and 0 deletions
+11
View File
@@ -0,0 +1,11 @@
{
"version": "0.0.1",
"configurations": [
{
"name": "nf-hotel-api",
"runtimeExecutable": ".venv/Scripts/python.exe",
"runtimeArgs": ["-m", "uvicorn", "nf_hotel_api.main:app", "--app-dir", "src", "--port", "8000"],
"port": 8000
}
]
}
+22
View File
@@ -0,0 +1,22 @@
# Copy to .env and fill in real values. Never commit the real .env file.
# Required: clients must send this value in the X-API-Key header.
API_KEY=change-me-to-a-long-random-secret
# Comma-separated list is not supported by pydantic-settings list parsing via
# env vars directly; set as a JSON array string if overriding, e.g.:
# ALLOWED_ORIGINS=["http://localhost:3000"]
# Local dataset fallback used by POST /api/v1/report/from-file
DEFAULT_DATA_PATH=data/nf_hotel_bookings.csv
CSV_SEPARATOR=;
# Room types, standard prices and room counts
METADATA_PATH=data/hotel_metadata.json
# OpenAI-compatible local inference server hosting the analysis LLM
LLM_BASE_URL=http://localhost:11434/v1
LLM_API_KEY=not-needed
LLM_MODEL=qwen3.8-whittle-moe-27b-a17.8b
# Reasoning models can take several minutes on a full report - keep generous.
LLM_TIMEOUT_SECONDS=900
+8
View File
@@ -0,0 +1,8 @@
.venv/
__pycache__/
*.pyc
.env
*.egg-info/
.pytest_cache/
.coverage
docs/htmlcov/
+179
View File
@@ -0,0 +1,179 @@
# NF Hotel Analytics API
Secure FastAPI service that cleans NF Hotel booking data, computes descriptive
statistics, and asks a local LLM to turn those statistics into a marketing /
operations report in markdown.
## Architecture
Clean-architecture layering, each concern isolated behind classes/interfaces:
```
src/nf_hotel_api/
domain/ # Pydantic models (Booking, ReportResponse, HotelMetadata) - no framework/IO deps
repositories/ # Where data comes from: CsvBookingRepository, JsonBookingRepository, JsonHotelMetadataRepository
services/ # Business logic
cleaning.py # DataCleaningService - wrong format / empty cells / wrong data / duplicates / pricing
statistics.py # DescriptiveStatsService - df.describe() (booking_id dropped)
llm_report.py # LLMReportService - calls the local LLM, strips <think> blocks
report.py # ReportService - orchestrates the above use case
api/ # FastAPI routers, DI wiring, HTTP-only concerns
core/ # Settings (env config) and API-key security
```
Dependencies point inward: `api` -> `services` -> `domain`, and `repositories`
implement an abstract interface consumed by `services`, so the analysis logic
never depends on FastAPI or pandas I/O directly.
## Setup
```bash
python -m venv .venv
.venv/Scripts/pip install -e ".[dev]"
cp .env.example .env
# edit .env: set API_KEY, and point LLM_BASE_URL at your local
# OpenAI-compatible inference server serving qwen3.8-whittle-moe-27b-a17.8b
```
## Run
```bash
.venv/Scripts/python -m uvicorn nf_hotel_api.main:app --app-dir src --reload
```
Docs: http://localhost:8000/docs
## Security
Every `/api/v1/*` route requires an `X-API-Key` header matching `API_KEY`
from the environment, checked with a constant-time comparison
(`secrets.compare_digest`). CORS is locked to `ALLOWED_ORIGINS`. Never commit
the real `.env`.
## Endpoints
- `GET /health` - unauthenticated liveness check.
- `POST /api/v1/report/from-file` - runs the report over the bundled
`data/nf_hotel_bookings.csv`.
- `POST /api/v1/report/from-json` - runs the report over booking records
supplied in the request body, same shape as the CSV columns.
### JSON format for `POST /api/v1/report/from-json`
Request body is `{"records": [<Booking>, ...]}`, at least one record. Each
`Booking` has these fields (mirrors [domain/schemas.py](src/nf_hotel_api/domain/schemas.py)):
| Field | Type | Example |
|---|---|---|
| `booking_id` | int | `1` |
| `hotel` | string | `"NF Hotel"` |
| `is_canceled` | int (0/1) | `0` |
| `lead_time` | int | `342` |
| `arrival_date_week_number` | int | `27` |
| `booking_date` | string (`YYYY-MM-DD` or `DD-MM-YYYY`) | `"2017-07-24"` |
| `arrival_date` | string (`YYYY-MM-DD` or `DD-MM-YYYY`) | `"2018-07-01"` |
| `arrival_date_day_of_month` | int | `1` |
| `stays_in_weekend_nights` | int | `0` |
| `stays_in_week_nights` | int | `0` |
| `adults` | int | `2` |
| `children` | int | `0` |
| `babies` | int | `0` |
| `meal` | string | `"BB"` |
| `country` | string | `"Portugal"` |
| `market_segment` | string | `"Direct"` |
| `is_repeated_guest` | int (0/1) | `0` |
| `previous_cancellations` | int | `0` |
| `assigned_room_type` | string (`"A"` small, `"B"` large) | `"A"` |
| `booking_changes` | int | `3` |
| `deposit_type` | string | `"No Deposit"` |
| `agent` | int | `0` |
| `customer_type` | string | `"No Contract (Single)"` |
| `required_car_parking_spaces` | int | `0` |
| `total_of_special_requests` | int | `0` |
| `prize_per_nigth` | number, optional (price paid per night; standard price if omitted) | `20` |
Fields are intentionally accepted as raw/untrusted input - values with the
wrong format, blanks, or implausible numbers are fine; `DataCleaningService`
fixes them before statistics are computed.
A ready-to-use example payload lives in
[examples/sample_booking_request.json](examples/sample_booking_request.json):
```bash
curl -X POST http://localhost:8000/api/v1/report/from-json \
-H "X-API-Key: $API_KEY" \
-H "Content-Type: application/json" \
--data @examples/sample_booking_request.json
```
Both report endpoints return:
```json
{
"descriptive_stats": { "<column>": { "mean": ..., "top": ..., "...": ... } },
"llm_report": "## Markdown report from the LLM, <think> blocks stripped"
}
```
## Data cleaning pipeline (`DataCleaningService`)
Applied in order, before any statistics are computed:
1. **Wrong format** - dates parsed to `datetime`, numeric columns coerced to
numeric, categorical columns trimmed of whitespace.
2. **Empty cells** - blank/placeholder values normalized to `"Unknown"` for
categoricals and `0` for numerics; rows with unparseable dates are dropped.
3. **Wrong data** - implausible guest counts (e.g. 55 adults) are clipped to
a sane maximum; bookings with zero total occupants are corrected to one
adult instead of being discarded.
4. **Duplicates** - exact duplicate bookings (ignoring `booking_id`, which is
just a row identifier) are removed, keeping the first occurrence.
5. **Pricing** - room size and revenue are derived, and missing prices are
filled from the hotel metadata (see below).
Dates are accepted as `YYYY-MM-DD` (JSON) and `DD-MM-YYYY` (the CSV export).
Each value is tried against both layouts, so a mixed column does not lose rows.
## Hotel metadata and pricing
Reference data that is not part of the bookings lives in
[data/hotel_metadata.json](data/hotel_metadata.json) (path set by
`METADATA_PATH`, loaded by `JsonHotelMetadataRepository`):
```json
{
"hotel": "NF Hotel",
"currency": "USD",
"room_types": {
"A": { "size": "Small", "standard_price_per_night": 20, "room_count": null },
"B": { "size": "Large", "standard_price_per_night": 25, "room_count": null }
}
}
```
| `assigned_room_type` | `room_size` | Standard price | Rooms in hotel |
|---|---|---|---|
| `A` | Small | $20 | *to be filled in* |
| `B` | Large | $25 | *to be filled in* |
- **`prize_per_nigth` in the CSV is the price the customer actually paid** and
is never overwritten. Only a missing or negative value is replaced with the
room type's standard price.
- Room types that are not in the metadata get `room_size = "Unknown"` and no
invented price.
- `revenue` = (`stays_in_weekend_nights` + `stays_in_week_nights`) x
`prize_per_nigth` for non-cancelled bookings, `0` for cancelled ones.
- `room_size`, `prize_per_nigth` and `revenue` are part of the descriptive
statistics sent to the LLM, and the prompt includes the room catalogue
(standard prices and room counts) from the metadata file.
- `room_count` is `null` until the real number of rooms is known. Occupancy
rate needs it, so it is not calculated yet. Events are also still missing.
To change a standard price or set a room count, edit `data/hotel_metadata.json`
and restart the service (it is read once at startup).
## Tests
```bash
.venv/Scripts/python -m pytest
```
+8
View File
@@ -0,0 +1,8 @@
{
"hotel": "NF Hotel",
"currency": "USD",
"room_types": {
"A": { "size": "Small", "standard_price_per_night": 20, "room_count": null },
"B": { "size": "Large", "standard_price_per_night": 25, "room_count": null }
}
}
File diff suppressed because it is too large Load Diff
+160
View File
@@ -0,0 +1,160 @@
# Classes vs. functions: why each file is shaped the way it is
The brief asked for a class-based project. That rule is applied wherever a
file actually has **state** (constructor arguments it reuses) or is one of
**several interchangeable implementations of the same contract**. Where a
file is one-shot, stateless glue code that a framework (FastAPI) expects in
a specific shape, it is left as a plain function - wrapping it in a class
would just be a single-method class, i.e. a function wearing a costume.
Rule of thumb used throughout:
- **Class** - has constructor state it reuses across methods, groups several
related private steps behind one public operation, or is one of multiple
interchangeable implementations of an abstract contract (repositories).
- **Function** - a single, stateless piece of wiring/glue, especially where
the framework itself expects a plain callable (FastAPI route handlers,
dependency providers).
This mirrors Clean Architecture's layering: the inner layers (domain,
services, repositories) hold the real behavior and are classes; the
outermost layer (API/framework glue) is intentionally thin and written the
way FastAPI wants it - plain functions.
## Domain layer
### `src/nf_hotel_api/domain/schemas.py` - classes: `Booking`, `BookingBatch`, `ReportResponse`
Pydantic `BaseModel` subclasses. This is not a style choice: Pydantic
requires a class to generate field validation, JSON (de)serialization, and
the OpenAPI schema FastAPI exposes at `/docs`. Each class is a **data
contract**, not behavior - no methods, just typed fields.
### `src/nf_hotel_api/domain/metadata.py` - classes: `RoomType`, `HotelMetadata`
Pydantic models for the reference data in `data/hotel_metadata.json`: per room
type its size, standard price per night and (optional) number of rooms.
`HotelMetadata.describe()` renders the catalogue as text for the LLM prompt.
## Core (config & security)
### `src/nf_hotel_api/core/config.py` - class `Settings` + function `get_settings()`
`Settings` is a `pydantic-settings` `BaseSettings` subclass - again required
by the library to get typed, validated configuration loaded from the
environment / `.env`. `get_settings()` is a two-line factory wrapped in
`@lru_cache` so the `Settings` object is built once and reused; a function
is all that's needed to memoize a constructor call.
### `src/nf_hotel_api/core/security.py` - class `ApiKeyAuthenticator`
Implemented as a class with `__call__` so a single instance
(`require_api_key = ApiKeyAuthenticator()`) can be reused as a FastAPI
dependency across every protected route. Being a class also makes it
trivial to unit test in isolation and to extend later (e.g. multiple valid
keys, per-key rate limiting) without touching every route that depends on
it.
## Repository layer
### `src/nf_hotel_api/repositories/metadata_repository.py` - class `JsonHotelMetadataRepository`
Loads and validates `hotel_metadata.json` into a `HotelMetadata`.
### `src/nf_hotel_api/repositories/booking_repository.py` - classes: `BookingRepository` (ABC), `CsvBookingRepository`, `JsonBookingRepository`
The textbook case for classes: an abstract base class defines one contract
(`load() -> DataFrame`), and two concrete classes implement it against
different sources while holding their own state (a file path + separator,
or a list of JSON records). Every service depends only on the abstract
`BookingRepository` type, so the CSV source can be swapped for the JSON
source (or a future database-backed repository) without touching any
business logic. Plain functions can't express "two interchangeable
implementations of one contract" this cleanly - that's exactly what
classes + inheritance are for.
## Services (business logic)
### `src/nf_hotel_api/services/cleaning.py` - class `DataCleaningService`
Groups five related private steps (`_fix_wrong_format`,
`_clean_empty_cells`, `_fix_wrong_data`, `_remove_duplicates`, `_add_pricing`)
behind one public `clean()` method. It is constructed with a `HotelMetadata`;
`_add_pricing` keeps the price paid in the CSV, fills missing prices from the
standard price, and derives `room_size` and `revenue`. The class keeps these steps cohesive, individually
testable (see `tests/test_cleaning.py`), and lets the whole pipeline be
swapped out in `ReportService`.
### `src/nf_hotel_api/services/statistics.py` - class `DescriptiveStatsService`
A single-purpose class computing and JSON-serializing `df.describe()`. It
holds no state today, but is a class for consistency with its sibling
services and so it can be constructor-injected into `ReportService` the
same way they are - adding config later (e.g. which percentiles to include)
won't change any call sites.
### `src/nf_hotel_api/services/llm_report.py` - classes `LLMReportService`, `LLMServiceError`
`LLMReportService` holds real constructor state (`base_url`, `api_key`,
`model`, `timeout_seconds`) used across its methods, and encapsulates the
network call plus the thinking-block-stripping post-processing behind one
`generate_report()` method - a natural fit for a class.
`LLMServiceError` is a small custom exception type so callers can catch
"the LLM failed" distinctly from a generic `httpx` error.
### `src/nf_hotel_api/services/report.py` - class `ReportService`
The orchestrator. Takes the other three services as constructor
dependencies and wires the raw-bookings -> clean -> stats -> LLM pipeline in
one `generate()` method. This is dependency injection in practice: each
collaborator is swappable, which is exactly how the tests substitute a fake
LLM service without touching the real cleaning/statistics logic.
## API layer (FastAPI wiring) - deliberately function-based
### `src/nf_hotel_api/api/deps.py` - functions: `get_cleaning_service`, `get_stats_service`, `get_llm_service`, `get_report_service`
FastAPI dependency-provider functions. FastAPI's `Depends()` system is
built around plain callables: it inspects a function's parameters and
return type to build the dependency graph and the OpenAPI schema. Each
function here does exactly one thing - construct and return a service
instance - and holds no state of its own, so wrapping it in a class would
add a layer of indirection with no benefit.
### `src/nf_hotel_api/api/v1/endpoints/report.py` - functions: `generate_report_from_file`, `generate_report_from_json`
Route handlers. FastAPI's `@router.post(...)` decorators are applied to
functions; this is the idiomatic (and for FastAPI, effectively required)
shape for an endpoint. Each handler is stateless - it receives its
dependencies via `Depends(...)` and immediately delegates to the
class-based service layer to do the actual work.
### `src/nf_hotel_api/api/v1/router.py` - no functions or classes, just module-level wiring
Purely declarative: creates one `APIRouter` and registers the report
router on it. There is no behavior here to encapsulate in either a
function or a class.
### `src/nf_hotel_api/main.py` - module-level app + function `health_check`
Creates the FastAPI app instance and registers middleware/routers at import
time - this file is the composition root. `health_check` is a trivial,
stateless liveness probe; a single-method class here would add nothing
over a function.
## Summary table
| File | Shape | Why |
|---|---|---|
| `domain/schemas.py` | Classes (Pydantic models) | Required by Pydantic/FastAPI for validation + OpenAPI schema |
| `core/config.py` | Class (`Settings`) + factory function | `BaseSettings` requires a class; caching a constructor call only needs a function |
| `core/security.py` | Class | Reusable, testable, extensible FastAPI dependency object |
| `repositories/booking_repository.py` | Classes (ABC + 2 impls) | Interchangeable implementations of one contract - inheritance/polymorphism |
| `services/cleaning.py` | Class | Groups related private steps behind one public operation |
| `services/statistics.py` | Class | Consistency with sibling services; injectable into `ReportService` |
| `services/llm_report.py` | Classes | Holds constructor state (URL, key, model, timeout) used across methods |
| `services/report.py` | Class | Orchestrator with injected, swappable collaborators |
| `api/deps.py` | Functions | FastAPI `Depends()` expects plain callables; no state to hold |
| `api/v1/endpoints/report.py` | Functions | FastAPI route handlers must be functions; stateless delegation |
| `api/v1/router.py` | Module-level wiring | No behavior to encapsulate |
| `main.py` | Module-level app + 1 function | Composition root; trivial health check |
+44
View File
@@ -0,0 +1,44 @@
"""Example: call POST /api/v1/report/from-file.
This endpoint takes no request body - it always analyzes the bundled
data/nf_hotel_bookings.csv on the server. Only the API key is required.
Usage:
API_KEY=your-key python examples/call_from_file_report.py
API_KEY=your-key python examples/call_from_file_report.py --base-url http://localhost:8000
"""
import argparse
import json
import os
import sys
import httpx
def main() -> None:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--base-url", default="http://localhost:8000")
args = parser.parse_args()
api_key = os.environ.get("API_KEY")
if not api_key:
print("Set the API_KEY environment variable first.", file=sys.stderr)
raise SystemExit(1)
response = httpx.post(
f"{args.base_url}/api/v1/report/from-file",
headers={"X-API-Key": api_key},
timeout=180.0,
)
response.raise_for_status()
report = response.json()
print("=== descriptive_stats ===")
print(json.dumps(report["descriptive_stats"], indent=2))
print("\n=== llm_report ===")
print(report["llm_report"])
if __name__ == "__main__":
main()
+58
View File
@@ -0,0 +1,58 @@
{
"records": [
{
"booking_id": 1,
"hotel": "NF Hotel",
"is_canceled": 0,
"lead_time": 342,
"arrival_date_week_number": 27,
"booking_date": "2017-07-24",
"arrival_date": "2018-07-01",
"arrival_date_day_of_month": 1,
"stays_in_weekend_nights": 0,
"stays_in_week_nights": 0,
"adults": 2,
"children": 0,
"babies": 0,
"meal": "BB",
"country": "Portugal",
"market_segment": "Direct",
"is_repeated_guest": 0,
"previous_cancellations": 0,
"assigned_room_type": "A",
"booking_changes": 3,
"deposit_type": "No Deposit",
"agent": 0,
"customer_type": "No Contract (Single)",
"required_car_parking_spaces": 0,
"total_of_special_requests": 0
},
{
"booking_id": 2,
"hotel": "NF Hotel",
"is_canceled": 0,
"lead_time": 737,
"arrival_date_week_number": 27,
"booking_date": "2016-06-24",
"arrival_date": "2018-07-01",
"arrival_date_day_of_month": 1,
"stays_in_weekend_nights": 0,
"stays_in_week_nights": 0,
"adults": 2,
"children": 0,
"babies": 0,
"meal": "BB",
"country": "Portugal",
"market_segment": "Direct",
"is_repeated_guest": 0,
"previous_cancellations": 0,
"assigned_room_type": "B",
"booking_changes": 4,
"deposit_type": "No Deposit",
"agent": 0,
"customer_type": "No Contract (Single)",
"required_car_parking_spaces": 0,
"total_of_special_requests": 0
}
]
}
+50
View File
@@ -0,0 +1,50 @@
[project]
name = "nf-hotel-api"
version = "0.1.0"
description = "Secure FastAPI service for NF Hotel booking analytics: descriptive statistics + LLM-generated markdown reports"
requires-python = ">=3.11"
dependencies = [
"fastapi>=0.115",
"uvicorn[standard]>=0.34",
"pandas>=2.2",
"pydantic>=2.9",
"pydantic-settings>=2.6",
"httpx>=0.27",
]
[project.optional-dependencies]
dev = [
"pytest>=8.3",
"pytest-asyncio>=0.24",
"pytest-cov>=5.0",
"httpx2>=2.0",
]
[tool.pytest.ini_options]
pythonpath = ["src"]
asyncio_mode = "auto"
addopts = "--cov=nf_hotel_api --cov-report=term-missing"
filterwarnings = [
# Third-party noise: starlette's testclient.py still references anyio's
# deprecated BlockingPortal alias; nothing our code touches directly.
"ignore:The anyio.abc.BlockingPortal alias is deprecated:DeprecationWarning",
]
[tool.coverage.run]
source = ["nf_hotel_api"]
[tool.coverage.report]
exclude_lines = [
"if TYPE_CHECKING:",
"raise NotImplementedError",
]
[tool.coverage.html]
directory = "docs/htmlcov"
[build-system]
requires = ["setuptools>=68"]
build-backend = "setuptools.build_meta"
[tool.setuptools.packages.find]
where = ["src"]
View File
View File
+48
View File
@@ -0,0 +1,48 @@
from functools import lru_cache
from fastapi import Depends
from nf_hotel_api.core.config import Settings, get_settings
from nf_hotel_api.domain.metadata import HotelMetadata
from nf_hotel_api.repositories.metadata_repository import JsonHotelMetadataRepository
from nf_hotel_api.services.cleaning import DataCleaningService
from nf_hotel_api.services.llm_report import LLMReportService
from nf_hotel_api.services.report import ReportService
from nf_hotel_api.services.statistics import DescriptiveStatsService
@lru_cache
def get_hotel_metadata() -> HotelMetadata:
return JsonHotelMetadataRepository(get_settings().metadata_path).load()
def get_cleaning_service(
metadata: HotelMetadata = Depends(get_hotel_metadata),
) -> DataCleaningService:
return DataCleaningService(metadata)
@lru_cache
def get_stats_service() -> DescriptiveStatsService:
return DescriptiveStatsService()
def get_llm_service(
settings: Settings = Depends(get_settings),
metadata: HotelMetadata = Depends(get_hotel_metadata),
) -> LLMReportService:
return LLMReportService(
metadata=metadata,
base_url=settings.llm_base_url,
api_key=settings.llm_api_key,
model=settings.llm_model,
timeout_seconds=settings.llm_timeout_seconds,
)
def get_report_service(
cleaning_service: DataCleaningService = Depends(get_cleaning_service),
stats_service: DescriptiveStatsService = Depends(get_stats_service),
llm_service: LLMReportService = Depends(get_llm_service),
) -> ReportService:
return ReportService(cleaning_service, stats_service, llm_service)
View File
@@ -0,0 +1,47 @@
from fastapi import APIRouter, Depends, HTTPException, status
from nf_hotel_api.api.deps import get_report_service
from nf_hotel_api.core.config import Settings, get_settings
from nf_hotel_api.core.security import require_api_key
from nf_hotel_api.domain.schemas import BookingBatch, ReportResponse
from nf_hotel_api.repositories.booking_repository import (
CsvBookingRepository,
JsonBookingRepository,
)
from nf_hotel_api.services.llm_report import LLMServiceError
from nf_hotel_api.services.report import ReportService
router = APIRouter(
prefix="/report",
tags=["report"],
dependencies=[Depends(require_api_key)],
)
@router.post("/from-file", response_model=ReportResponse)
async def generate_report_from_file(
report_service: ReportService = Depends(get_report_service),
settings: Settings = Depends(get_settings),
) -> ReportResponse:
"""Generate the analytics report from the bundled nf_hotel_bookings.csv file."""
repository = CsvBookingRepository(settings.default_data_path, settings.csv_separator)
try:
return await report_service.generate(repository)
except FileNotFoundError as exc:
raise HTTPException(status_code=status.HTTP_404_NOT_FOUND, detail=str(exc)) from exc
except LLMServiceError as exc:
raise HTTPException(status_code=status.HTTP_502_BAD_GATEWAY, detail=str(exc)) from exc
@router.post("/from-json", response_model=ReportResponse)
async def generate_report_from_json(
batch: BookingBatch,
report_service: ReportService = Depends(get_report_service),
) -> ReportResponse:
"""Generate the analytics report from booking records supplied as JSON."""
records = [record.model_dump() for record in batch.records]
repository = JsonBookingRepository(records)
try:
return await report_service.generate(repository)
except LLMServiceError as exc:
raise HTTPException(status_code=status.HTTP_502_BAD_GATEWAY, detail=str(exc)) from exc
+6
View File
@@ -0,0 +1,6 @@
from fastapi import APIRouter
from nf_hotel_api.api.v1.endpoints import report
api_router = APIRouter(prefix="/api/v1")
api_router.include_router(report.router)
View File
+48
View File
@@ -0,0 +1,48 @@
from functools import lru_cache
from pathlib import Path
from pydantic import field_validator
from pydantic_settings import BaseSettings, SettingsConfigDict
# src/nf_hotel_api/core/config.py -> project root is three levels up.
_PROJECT_ROOT = Path(__file__).resolve().parents[3]
class Settings(BaseSettings):
"""Central application configuration, loaded from environment / .env."""
model_config = SettingsConfigDict(env_file=".env", env_file_encoding="utf-8", extra="ignore")
# --- API security ---
api_key: str
allowed_origins: list[str] = ["http://localhost:3000"]
# --- Local data source (fallback / default dataset) ---
default_data_path: Path = Path("data/nf_hotel_bookings.csv")
csv_separator: str = ";"
# Reference data: room types, standard prices and how many rooms exist.
metadata_path: Path = Path("data/hotel_metadata.json")
@field_validator("default_data_path", "metadata_path")
@classmethod
def _resolve_relative_to_project_root(cls, value: Path) -> Path:
# Anchored to the project root, not the process's current working
# directory - otherwise this 404s whenever uvicorn/pytest is
# launched from anywhere other than the repo root.
return value if value.is_absolute() else _PROJECT_ROOT / value
# --- LLM (OpenAI-compatible local inference server, e.g. vLLM / LM Studio) ---
llm_base_url: str = "http://localhost:1234/v1"
llm_api_key: str = "not-needed"
llm_model: str = "qwen3.8-whittle-moe-27b-a17.8b"
# Reasoning models can spend minutes generating <think>/reasoning_content
# tokens before the final answer - keep this generous.
llm_timeout_seconds: float = 900.0
@lru_cache
def get_settings() -> Settings:
# api_key (and other required fields) come from the environment / .env
# at runtime via BaseSettings, not from constructor arguments - static
# type checkers can't see that, hence the ignore.
return Settings() # type: ignore[call-arg]
+30
View File
@@ -0,0 +1,30 @@
import secrets
from fastapi import Depends, HTTPException, Security, status
from fastapi.security import APIKeyHeader
from nf_hotel_api.core.config import Settings, get_settings
_api_key_header = APIKeyHeader(name="X-API-Key", auto_error=False)
class ApiKeyAuthenticator:
"""Validates inbound requests against the configured API key.
Uses a constant-time comparison to avoid leaking key length/content via
response-time side channels.
"""
def __call__(
self,
api_key: str | None = Security(_api_key_header),
settings: Settings = Depends(get_settings),
) -> None:
if not api_key or not secrets.compare_digest(api_key, settings.api_key):
raise HTTPException(
status_code=status.HTTP_401_UNAUTHORIZED,
detail="Invalid or missing API key",
)
require_api_key = ApiKeyAuthenticator()
View File
+34
View File
@@ -0,0 +1,34 @@
from pydantic import BaseModel, Field
class RoomType(BaseModel):
"""One kind of room the hotel offers."""
size: str
standard_price_per_night: float = Field(..., ge=0)
# None = not known yet; anything that needs the count (e.g. occupancy)
# must be skipped rather than guess.
room_count: int | None = Field(default=None, ge=0)
class HotelMetadata(BaseModel):
"""Reference data about the hotel that is not part of the booking records.
``room_types`` is keyed by the ``assigned_room_type`` code used in the CSV.
"""
hotel: str
currency: str = "USD"
room_types: dict[str, RoomType]
def describe(self) -> str:
"""Plain-text summary of the room catalogue, for the LLM prompt."""
lines = []
for code, room in self.room_types.items():
count = "unknown" if room.room_count is None else str(room.room_count)
lines.append(
f"- Room type {code}: {room.size}, standard price "
f"{room.standard_price_per_night:g} {self.currency} per night, "
f"number of rooms in hotel: {count}"
)
return "\n".join(lines)
+54
View File
@@ -0,0 +1,54 @@
from typing import Any
from pydantic import BaseModel, Field
class Booking(BaseModel):
"""A single hotel booking record, mirroring nf_hotel_bookings.csv.
Fields are intentionally loosely typed (str for dates/free-form values)
because incoming JSON is treated as *raw* data that still needs to pass
through the cleaning pipeline before analysis.
"""
booking_id: int
hotel: str
is_canceled: int
lead_time: int
arrival_date_week_number: int
booking_date: str
arrival_date: str
arrival_date_day_of_month: int
stays_in_weekend_nights: int
stays_in_week_nights: int
adults: int
children: int
babies: int
meal: str
country: str
market_segment: str
is_repeated_guest: int
previous_cancellations: int
assigned_room_type: str
booking_changes: int
deposit_type: str
agent: int
customer_type: str
required_car_parking_spaces: int
total_of_special_requests: int
# Optional: when omitted (or wrong) it is derived from assigned_room_type.
# Spelling mirrors the CSV column header.
prize_per_nigth: float | None = None
class BookingBatch(BaseModel):
"""Payload for submitting raw booking records as JSON."""
records: list[Booking] = Field(..., min_length=1)
class ReportResponse(BaseModel):
"""Result of a full analytics report: stats + LLM narrative."""
descriptive_stats: dict[str, dict[str, Any]]
llm_report: str
+28
View File
@@ -0,0 +1,28 @@
from fastapi import FastAPI
from fastapi.middleware.cors import CORSMiddleware
from nf_hotel_api.api.v1.router import api_router
from nf_hotel_api.core.config import get_settings
settings = get_settings()
app = FastAPI(
title="NF Hotel Analytics API",
description="Secure API for descriptive statistics and LLM-generated reports over NF Hotel bookings.",
version="0.1.0",
)
app.add_middleware(
CORSMiddleware,
allow_origins=settings.allowed_origins,
allow_credentials=True,
allow_methods=["GET", "POST"],
allow_headers=["X-API-Key", "Content-Type"],
)
app.include_router(api_router)
@app.get("/health", tags=["health"])
def health_check() -> dict[str, str]:
return {"status": "ok"}
@@ -0,0 +1,36 @@
from abc import ABC, abstractmethod
from pathlib import Path
from typing import Any
import pandas as pd
class BookingRepository(ABC):
"""Abstraction over "where booking records come from"."""
@abstractmethod
def load(self) -> pd.DataFrame:
"""Return the raw (unclean) bookings as a DataFrame."""
class CsvBookingRepository(BookingRepository):
"""Loads bookings from the on-disk NF Hotel CSV export."""
def __init__(self, path: Path, separator: str = ";") -> None:
self._path = path
self._separator = separator
def load(self) -> pd.DataFrame:
if not self._path.exists():
raise FileNotFoundError(f"Booking data file not found: {self._path}")
return pd.read_csv(self._path, sep=self._separator)
class JsonBookingRepository(BookingRepository):
"""Loads bookings from JSON records supplied directly by the caller."""
def __init__(self, records: list[dict[str, Any]]) -> None:
self._records = records
def load(self) -> pd.DataFrame:
return pd.DataFrame.from_records(self._records)
@@ -0,0 +1,15 @@
from pathlib import Path
from nf_hotel_api.domain.metadata import HotelMetadata
class JsonHotelMetadataRepository:
"""Loads hotel reference data (room types, standard prices, room counts)."""
def __init__(self, path: Path) -> None:
self._path = path
def load(self) -> HotelMetadata:
if not self._path.exists():
raise FileNotFoundError(f"Hotel metadata file not found: {self._path}")
return HotelMetadata.model_validate_json(self._path.read_text(encoding="utf-8"))
+176
View File
@@ -0,0 +1,176 @@
import pandas as pd
from nf_hotel_api.domain.metadata import HotelMetadata
_DATE_COLUMNS = ("booking_date", "arrival_date")
_NUMERIC_COLUMNS = (
"booking_id",
"is_canceled",
"lead_time",
"arrival_date_week_number",
"arrival_date_day_of_month",
"stays_in_weekend_nights",
"stays_in_week_nights",
"adults",
"children",
"babies",
"is_repeated_guest",
"previous_cancellations",
"booking_changes",
"agent",
"required_car_parking_spaces",
"total_of_special_requests",
)
_CATEGORICAL_COLUMNS = (
"hotel",
"meal",
"country",
"market_segment",
"assigned_room_type",
"deposit_type",
"customer_type",
)
# Accepted date layouts: ISO (JSON payloads) and day-first (the CSV export).
_DATE_FORMATS = ("%Y-%m-%d", "%d-%m-%Y")
_PRICE_COLUMN = "prize_per_nigth"
_EMPTY_MARKERS = ("", "nan", "null", "none", "n/a", "na")
# Bookings with more occupants than this in a single field are treated as
# data-entry errors (e.g. "55 adults" in one room) rather than real guests.
_MAX_PLAUSIBLE_ADULTS = 10
_MAX_PLAUSIBLE_CHILDREN = 5
class DataCleaningService:
"""Cleans raw booking data before it is analyzed.
Each private method covers one classic pandas cleaning concern, run in
a fixed order via :meth:`clean`. Once the data is clean, :meth:`_add_pricing`
derives room size and revenue using the hotel metadata.
"""
def __init__(self, metadata: HotelMetadata) -> None:
self._metadata = metadata
def clean(self, df: pd.DataFrame) -> pd.DataFrame:
df = df.copy()
df = self._fix_wrong_format(df)
df = self._clean_empty_cells(df)
df = self._fix_wrong_data(df)
df = self._remove_duplicates(df)
df = self._add_pricing(df)
return df.reset_index(drop=True)
def _fix_wrong_format(self, df: pd.DataFrame) -> pd.DataFrame:
"""Coerce columns into their expected dtype (dates, numbers)."""
for column in _DATE_COLUMNS:
if column in df.columns:
df[column] = self._parse_dates(df[column])
for column in _NUMERIC_COLUMNS:
if column in df.columns:
df[column] = pd.to_numeric(df[column], errors="coerce")
if _PRICE_COLUMN in df.columns:
df[_PRICE_COLUMN] = pd.to_numeric(df[_PRICE_COLUMN], errors="coerce")
for column in _CATEGORICAL_COLUMNS:
if column in df.columns:
df[column] = df[column].astype("string").str.strip()
return df
@staticmethod
def _parse_dates(series: pd.Series) -> pd.Series:
"""Parse each value with the first matching layout in ``_DATE_FORMATS``.
A single ``pd.to_datetime`` call would infer one layout from the first
value and turn every date in the other layout into NaT, which would
then silently drop those rows.
"""
parsed = pd.Series(pd.NaT, index=series.index, dtype="datetime64[ns]")
for date_format in _DATE_FORMATS:
candidate = pd.to_datetime(series, format=date_format, errors="coerce")
parsed = parsed.fillna(candidate)
return parsed
def _clean_empty_cells(self, df: pd.DataFrame) -> pd.DataFrame:
"""Normalize placeholder/blank values to NA, then fill sensibly."""
for column in _CATEGORICAL_COLUMNS:
if column not in df.columns:
continue
lowered = df[column].str.lower()
df.loc[lowered.isin(_EMPTY_MARKERS), column] = pd.NA
df[column] = df[column].fillna("Unknown")
for column in _NUMERIC_COLUMNS:
if column in df.columns:
df[column] = df[column].fillna(0)
df = df.dropna(subset=[c for c in _DATE_COLUMNS if c in df.columns])
return df
def _fix_wrong_data(self, df: pd.DataFrame) -> pd.DataFrame:
"""Correct implausible values without discarding the whole row."""
if "adults" in df.columns:
df["adults"] = df["adults"].clip(lower=0, upper=_MAX_PLAUSIBLE_ADULTS)
if "children" in df.columns:
df["children"] = df["children"].clip(lower=0, upper=_MAX_PLAUSIBLE_CHILDREN)
if "babies" in df.columns:
df["babies"] = df["babies"].clip(lower=0, upper=_MAX_PLAUSIBLE_CHILDREN)
# A booking with zero occupants in every guest field is invalid;
# treat it as a single adult rather than dropping the record.
guest_columns = [c for c in ("adults", "children", "babies") if c in df.columns]
if guest_columns:
no_guests = (df[guest_columns].sum(axis=1) == 0)
if "adults" in df.columns:
df.loc[no_guests, "adults"] = 1
for column in ("lead_time", "booking_changes", "previous_cancellations", "agent"):
if column in df.columns:
df[column] = df[column].clip(lower=0)
return df
def _remove_duplicates(self, df: pd.DataFrame) -> pd.DataFrame:
"""Drop repeated bookings, ignoring the (non-business) id column."""
subset = [c for c in df.columns if c != "booking_id"]
return df.drop_duplicates(subset=subset, keep="first")
def _add_pricing(self, df: pd.DataFrame) -> pd.DataFrame:
"""Add room size and revenue, and fill in missing prices.
``prize_per_nigth`` in the data is the price the customer actually
paid, so it is kept as-is. Only a missing or negative price is replaced
with the room type's standard price from the hotel metadata. Room types
outside the metadata get no size and no invented price.
"""
if "assigned_room_type" not in df.columns:
return df
room_types = self._metadata.room_types
room_type = df["assigned_room_type"].astype("string").str.upper()
standard_price = room_type.map(
{code: room.standard_price_per_night for code, room in room_types.items()}
).astype("float64")
paid_price = df.get(_PRICE_COLUMN, pd.Series(float("nan"), index=df.index))
df[_PRICE_COLUMN] = paid_price.where(paid_price >= 0).fillna(standard_price)
df["room_size"] = room_type.map(
{code: room.size for code, room in room_types.items()}
).fillna("Unknown")
nights_columns = [
c for c in ("stays_in_weekend_nights", "stays_in_week_nights") if c in df.columns
]
if nights_columns:
# Cancelled bookings bring in no money.
billable = df["is_canceled"] == 0 if "is_canceled" in df.columns else True
df["revenue"] = df[nights_columns].sum(axis=1) * df[_PRICE_COLUMN] * billable
return df
+86
View File
@@ -0,0 +1,86 @@
import json
import re
from typing import Any
import httpx
from nf_hotel_api.domain.metadata import HotelMetadata
_PROMPT_TEMPLATE = """You work as a data analyst and in marketing to optimize hotel operations.
User cannot interact with you so do not ask questions.
Respond in markdown format.
{hotel_context}We have extract descriptive analysis:\n
{descriptive_analysis_data}"""
_HOTEL_CONTEXT_TEMPLATE = """Hotel reference data (room types the bookings refer to):
{room_catalogue}
The column prize_per_nigth is the price the customer actually paid per night.
The column revenue is nights x prize_per_nigth for non-cancelled bookings.
"""
# Reasoning models (e.g. Qwen3) may wrap their internal reasoning in
# <think>...</think>; that content must never reach the API response.
_THINK_BLOCK_PATTERN = re.compile(r"<think>.*?</think>", re.DOTALL | re.IGNORECASE)
class LLMServiceError(RuntimeError):
"""Raised when the LLM backend cannot produce a report."""
class LLMReportService:
"""Turns descriptive statistics into a marketing/ops narrative via an LLM."""
def __init__(
self,
base_url: str,
api_key: str,
model: str,
timeout_seconds: float,
metadata: HotelMetadata | None = None,
) -> None:
self._metadata = metadata
self._base_url = base_url.rstrip("/")
self._api_key = api_key
self._model = model
self._timeout_seconds = timeout_seconds
async def generate_report(self, descriptive_stats: dict[str, Any]) -> str:
hotel_context = (
_HOTEL_CONTEXT_TEMPLATE.format(room_catalogue=self._metadata.describe())
if self._metadata
else ""
)
prompt = _PROMPT_TEMPLATE.format(
hotel_context=hotel_context,
descriptive_analysis_data=json.dumps(descriptive_stats, indent=2),
)
payload = {
"model": self._model,
"messages": [{"role": "user", "content": prompt}],
"temperature": 0.3,
}
headers = {"Authorization": f"Bearer {self._api_key}"}
try:
async with httpx.AsyncClient(timeout=self._timeout_seconds) as client:
response = await client.post(
f"{self._base_url}/chat/completions",
json=payload,
headers=headers,
)
response.raise_for_status()
except httpx.HTTPError as exc:
raise LLMServiceError(f"LLM backend request failed: {exc}") from exc
data = response.json()
try:
raw_content = data["choices"][0]["message"]["content"]
except (KeyError, IndexError) as exc:
raise LLMServiceError(f"Unexpected LLM response shape: {data}") from exc
return self._strip_thinking(raw_content)
@staticmethod
def _strip_thinking(content: str) -> str:
return _THINK_BLOCK_PATTERN.sub("", content).strip()
+26
View File
@@ -0,0 +1,26 @@
from nf_hotel_api.domain.schemas import ReportResponse
from nf_hotel_api.repositories.booking_repository import BookingRepository
from nf_hotel_api.services.cleaning import DataCleaningService
from nf_hotel_api.services.llm_report import LLMReportService
from nf_hotel_api.services.statistics import DescriptiveStatsService
class ReportService:
"""Application-layer use case: raw bookings -> full analytics report."""
def __init__(
self,
cleaning_service: DataCleaningService,
stats_service: DescriptiveStatsService,
llm_service: LLMReportService,
) -> None:
self._cleaning_service = cleaning_service
self._stats_service = stats_service
self._llm_service = llm_service
async def generate(self, repository: BookingRepository) -> ReportResponse:
raw_df = repository.load()
clean_df = self._cleaning_service.clean(raw_df)
descriptive_stats = self._stats_service.compute(clean_df)
llm_report = await self._llm_service.generate_report(descriptive_stats)
return ReportResponse(descriptive_stats=descriptive_stats, llm_report=llm_report)
+27
View File
@@ -0,0 +1,27 @@
from typing import Any
import pandas as pd
class DescriptiveStatsService:
"""Computes descriptive statistics over cleaned booking data."""
def compute(self, df: pd.DataFrame) -> dict[str, dict[str, Any]]:
df = df.drop(columns=["booking_id"], errors="ignore")
described = df.describe(include="all")
return {
str(column): {
str(stat): self._to_jsonable(value)
for stat, value in described[column].items()
if pd.notna(value)
}
for column in described.columns
}
@staticmethod
def _to_jsonable(value: Any) -> Any:
if isinstance(value, pd.Timestamp):
return value.isoformat()
if hasattr(value, "item"):
return value.item()
return value
+114
View File
@@ -0,0 +1,114 @@
import os
# Force (not setdefault): an API_KEY already set in the shell/IDE/.env must not
# leak into the tests, or the authenticated API tests fail with 401.
os.environ["API_KEY"] = "test-api-key"
import pandas as pd
import pytest
from nf_hotel_api.domain.metadata import HotelMetadata, RoomType
@pytest.fixture
def dirty_bookings_df() -> pd.DataFrame:
"""A small, deliberately messy DataFrame exercising every cleaning rule."""
return pd.DataFrame(
[
{
"booking_id": 1,
"hotel": "NF Hotel",
"is_canceled": 0,
"lead_time": "10",
"arrival_date_week_number": 1,
"booking_date": "2024-01-01",
"arrival_date": "2024-02-01",
"arrival_date_day_of_month": 1,
"stays_in_weekend_nights": 1,
"stays_in_week_nights": 2,
"adults": 2,
"children": 0,
"babies": 0,
"meal": "BB",
"country": "Portugal",
"market_segment": "Direct",
"is_repeated_guest": 0,
"previous_cancellations": 0,
"assigned_room_type": "A",
"booking_changes": 0,
"deposit_type": "No Deposit",
"agent": 0,
"customer_type": "Contract (Single)",
"required_car_parking_spaces": 0,
"total_of_special_requests": 0,
},
{
# Exact duplicate of booking 1 except the id -> must be removed.
"booking_id": 2,
"hotel": "NF Hotel",
"is_canceled": 0,
"lead_time": "10",
"arrival_date_week_number": 1,
"booking_date": "2024-01-01",
"arrival_date": "2024-02-01",
"arrival_date_day_of_month": 1,
"stays_in_weekend_nights": 1,
"stays_in_week_nights": 2,
"adults": 2,
"children": 0,
"babies": 0,
"meal": "BB",
"country": "Portugal",
"market_segment": "Direct",
"is_repeated_guest": 0,
"previous_cancellations": 0,
"assigned_room_type": "A",
"booking_changes": 0,
"deposit_type": "No Deposit",
"agent": 0,
"customer_type": "Contract (Single)",
"required_car_parking_spaces": 0,
"total_of_special_requests": 0,
},
{
# Wrong/empty data: bogus adults count, blank meal, zero guests overall.
"booking_id": 3,
"hotel": "NF Hotel",
"is_canceled": 1,
"lead_time": -5,
"arrival_date_week_number": 2,
"booking_date": "2024-01-05",
"arrival_date": "not-a-date",
"arrival_date_day_of_month": 5,
"stays_in_weekend_nights": 0,
"stays_in_week_nights": 1,
"adults": 55,
"children": 0,
"babies": 0,
"meal": "",
"country": "Unknown",
"market_segment": "Groups",
"is_repeated_guest": 0,
"previous_cancellations": 0,
"assigned_room_type": "C",
"booking_changes": 0,
"deposit_type": "Non Refund",
"agent": 9,
"customer_type": "Group Contract",
"required_car_parking_spaces": 0,
"total_of_special_requests": 1,
},
]
)
@pytest.fixture
def hotel_metadata() -> HotelMetadata:
return HotelMetadata(
hotel="NF Hotel",
currency="USD",
room_types={
"A": RoomType(size="Small", standard_price_per_night=20, room_count=10),
"B": RoomType(size="Large", standard_price_per_night=25, room_count=5),
},
)
+56
View File
@@ -0,0 +1,56 @@
from fastapi.testclient import TestClient
from nf_hotel_api.api.deps import get_llm_service
from nf_hotel_api.main import app
API_KEY = "test-api-key"
class _FakeLLMService:
async def generate_report(self, descriptive_stats: dict) -> str:
return "## Fake Report"
app.dependency_overrides[get_llm_service] = lambda: _FakeLLMService()
client = TestClient(app)
def test_health_check_is_public():
response = client.get("/health")
assert response.status_code == 200
assert response.json() == {"status": "ok"}
def test_report_endpoint_rejects_missing_api_key():
response = client.post("/api/v1/report/from-file")
assert response.status_code == 401
def test_report_endpoint_rejects_wrong_api_key():
response = client.post(
"/api/v1/report/from-file", headers={"X-API-Key": "wrong-key"}
)
assert response.status_code == 401
def test_report_from_file_returns_stats_and_llm_report():
response = client.post(
"/api/v1/report/from-file", headers={"X-API-Key": API_KEY}
)
assert response.status_code == 200
body = response.json()
assert "booking_id" not in body["descriptive_stats"]
assert body["llm_report"] == "## Fake Report"
def test_report_from_json_returns_stats_and_llm_report(dirty_bookings_df):
records = dirty_bookings_df.to_dict(orient="records")
response = client.post(
"/api/v1/report/from-json",
headers={"X-API-Key": API_KEY},
json={"records": records},
)
assert response.status_code == 200
body = response.json()
assert "booking_id" not in body["descriptive_stats"]
assert body["llm_report"] == "## Fake Report"
+107
View File
@@ -0,0 +1,107 @@
import pandas as pd
from nf_hotel_api.services.cleaning import DataCleaningService
def test_clean_removes_duplicate_bookings(dirty_bookings_df, hotel_metadata):
result = DataCleaningService(hotel_metadata).clean(dirty_bookings_df)
# booking 2 is a duplicate of booking 1 (ignoring booking_id) and must go.
assert result["booking_id"].tolist() == [1]
def test_clean_coerces_dates_and_drops_unparseable_rows(hotel_metadata):
df = pd.DataFrame(
[
{"booking_id": 1, "booking_date": "2024-01-01", "arrival_date": "2024-02-01"},
{"booking_id": 2, "booking_date": "2024-01-01", "arrival_date": "garbage"},
]
)
result = DataCleaningService(hotel_metadata)._fix_wrong_format(df)
assert pd.api.types.is_datetime64_any_dtype(result["arrival_date"])
cleaned = DataCleaningService(hotel_metadata)._clean_empty_cells(result)
assert cleaned["booking_id"].tolist() == [1]
def test_clean_replaces_blank_categoricals_with_unknown_placeholder(hotel_metadata):
df = pd.DataFrame([{"hotel": "NF Hotel", "meal": "", "country": "Portugal"}])
result = DataCleaningService(hotel_metadata)._clean_empty_cells(df)
assert result.loc[0, "meal"] == "Unknown"
assert result.loc[0, "country"] == "Portugal"
def test_clean_caps_implausible_guest_counts_and_fixes_zero_guest_rows(hotel_metadata):
df = pd.DataFrame(
[
{"adults": 55, "children": 0, "babies": 0},
{"adults": 0, "children": 0, "babies": 0},
]
)
result = DataCleaningService(hotel_metadata)._fix_wrong_data(df)
assert result.loc[0, "adults"] <= 10
assert result.loc[1, "adults"] == 1
def test_clean_parses_day_first_and_iso_dates_without_dropping_rows(hotel_metadata):
df = pd.DataFrame(
[
{"booking_id": 1, "booking_date": "13-07-2018", "arrival_date": "02-06-2018"},
{"booking_id": 2, "booking_date": "2018-07-14", "arrival_date": "2018-06-03"},
]
)
result = DataCleaningService(hotel_metadata).clean(df)
assert result["booking_id"].tolist() == [1, 2]
assert result.loc[0, "booking_date"] == pd.Timestamp(2018, 7, 13)
assert result.loc[0, "arrival_date"] == pd.Timestamp(2018, 6, 2)
assert result.loc[1, "arrival_date"] == pd.Timestamp(2018, 6, 3)
def test_add_pricing_keeps_price_paid_and_fills_missing_from_standard_price(hotel_metadata):
df = pd.DataFrame(
[
{"assigned_room_type": "A", "prize_per_nigth": 15}, # discounted, kept
{"assigned_room_type": "B", "prize_per_nigth": None}, # missing -> 25
{"assigned_room_type": "A", "prize_per_nigth": -3}, # invalid -> 20
{"assigned_room_type": "A", "prize_per_nigth": 0}, # free stay, kept
{"assigned_room_type": "C", "prize_per_nigth": None}, # not in metadata
]
)
result = DataCleaningService(hotel_metadata)._add_pricing(df)
assert result["prize_per_nigth"].iloc[:4].tolist() == [15, 25, 20, 0]
assert pd.isna(result.loc[4, "prize_per_nigth"])
assert result["room_size"].tolist() == ["Small", "Large", "Small", "Small", "Unknown"]
def test_add_pricing_revenue_is_nights_times_price_paid_and_zero_when_canceled(hotel_metadata):
df = pd.DataFrame(
[
{"assigned_room_type": "A", "is_canceled": 0, "prize_per_nigth": 15,
"stays_in_weekend_nights": 1, "stays_in_week_nights": 2},
{"assigned_room_type": "B", "is_canceled": 0,
"stays_in_weekend_nights": 0, "stays_in_week_nights": 3},
{"assigned_room_type": "B", "is_canceled": 1,
"stays_in_weekend_nights": 2, "stays_in_week_nights": 2},
]
)
result = DataCleaningService(hotel_metadata)._add_pricing(df)
assert result["revenue"].tolist() == [45, 75, 0]
def test_clean_adds_pricing_columns_to_full_pipeline(dirty_bookings_df, hotel_metadata):
result = DataCleaningService(hotel_metadata).clean(dirty_bookings_df)
assert result.loc[0, "prize_per_nigth"] == 20
assert result.loc[0, "room_size"] == "Small"
assert result.loc[0, "revenue"] == 60 # 3 nights x $20
+80
View File
@@ -0,0 +1,80 @@
import httpx
import pytest
from nf_hotel_api.services.llm_report import LLMReportService, LLMServiceError
def test_strip_thinking_removes_think_block():
raw = "<think>internal reasoning that must not leak</think>## Report\nBody text"
assert LLMReportService._strip_thinking(raw) == "## Report\nBody text"
def test_strip_thinking_is_noop_when_no_think_block():
raw = "## Report\nBody text"
assert LLMReportService._strip_thinking(raw) == raw
class _FakeResponse:
def __init__(self, payload: dict, status_code: int = 200) -> None:
self._payload = payload
self.status_code = status_code
def raise_for_status(self) -> None:
if self.status_code >= 400:
request = httpx.Request("POST", "http://fake/chat/completions")
raise httpx.HTTPStatusError(
"error", request=request, response=httpx.Response(self.status_code, request=request)
)
def json(self) -> dict:
return self._payload
class _FakeAsyncClient:
def __init__(self, payload: dict, status_code: int = 200) -> None:
self._payload = payload
self._status_code = status_code
async def __aenter__(self) -> "_FakeAsyncClient":
return self
async def __aexit__(self, *args) -> None:
return None
async def post(self, *args, **kwargs) -> _FakeResponse:
return _FakeResponse(self._payload, self._status_code)
@pytest.mark.asyncio
async def test_generate_report_strips_thinking_and_returns_content(monkeypatch):
payload = {
"choices": [
{"message": {"content": "<think>hidden</think>## Insights\nBook more direct."}}
]
}
monkeypatch.setattr(
"nf_hotel_api.services.llm_report.httpx.AsyncClient",
lambda timeout: _FakeAsyncClient(payload),
)
service = LLMReportService(
base_url="http://fake/v1", api_key="k", model="m", timeout_seconds=1.0
)
report = await service.generate_report({"adults": {"mean": 2}})
assert report == "## Insights\nBook more direct."
@pytest.mark.asyncio
async def test_generate_report_raises_llm_service_error_on_bad_response(monkeypatch):
monkeypatch.setattr(
"nf_hotel_api.services.llm_report.httpx.AsyncClient",
lambda timeout: _FakeAsyncClient({"unexpected": "shape"}),
)
service = LLMReportService(
base_url="http://fake/v1", api_key="k", model="m", timeout_seconds=1.0
)
with pytest.raises(LLMServiceError):
await service.generate_report({"adults": {"mean": 2}})
+35
View File
@@ -0,0 +1,35 @@
import pytest
from nf_hotel_api.domain.metadata import HotelMetadata, RoomType
from nf_hotel_api.repositories.metadata_repository import JsonHotelMetadataRepository
def test_bundled_metadata_file_defines_small_and_large_room_prices():
from nf_hotel_api.core.config import get_settings
metadata = JsonHotelMetadataRepository(get_settings().metadata_path).load()
assert metadata.room_types["A"].size == "Small"
assert metadata.room_types["A"].standard_price_per_night == 20
assert metadata.room_types["B"].size == "Large"
assert metadata.room_types["B"].standard_price_per_night == 25
def test_repository_raises_when_file_is_missing(tmp_path):
with pytest.raises(FileNotFoundError):
JsonHotelMetadataRepository(tmp_path / "nope.json").load()
def test_describe_lists_prices_and_marks_unknown_room_counts():
metadata = HotelMetadata(
hotel="NF Hotel",
room_types={
"A": RoomType(size="Small", standard_price_per_night=20, room_count=10),
"B": RoomType(size="Large", standard_price_per_night=25),
},
)
text = metadata.describe()
assert "Room type A: Small, standard price 20 USD per night, number of rooms in hotel: 10" in text
assert "Room type B: Large, standard price 25 USD per night, number of rooms in hotel: unknown" in text
+28
View File
@@ -0,0 +1,28 @@
import pytest
from nf_hotel_api.repositories.booking_repository import JsonBookingRepository
from nf_hotel_api.services.cleaning import DataCleaningService
from nf_hotel_api.services.report import ReportService
from nf_hotel_api.services.statistics import DescriptiveStatsService
class _FakeLLMService:
async def generate_report(self, descriptive_stats: dict) -> str:
assert "booking_id" not in descriptive_stats
return "## Fake Report"
@pytest.mark.asyncio
async def test_generate_produces_stats_and_llm_report(dirty_bookings_df, hotel_metadata):
repository = JsonBookingRepository(dirty_bookings_df.to_dict(orient="records"))
service = ReportService(
cleaning_service=DataCleaningService(hotel_metadata),
stats_service=DescriptiveStatsService(),
llm_service=_FakeLLMService(),
)
result = await service.generate(repository)
assert result.llm_report == "## Fake Report"
assert "adults" in result.descriptive_stats
assert "booking_id" not in result.descriptive_stats
+27
View File
@@ -0,0 +1,27 @@
import json
from nf_hotel_api.services.cleaning import DataCleaningService
from nf_hotel_api.services.statistics import DescriptiveStatsService
def test_compute_drops_booking_id(dirty_bookings_df, hotel_metadata):
clean_df = DataCleaningService(hotel_metadata).clean(dirty_bookings_df)
stats = DescriptiveStatsService().compute(clean_df)
assert "booking_id" not in stats
def test_compute_output_is_json_serializable(dirty_bookings_df, hotel_metadata):
clean_df = DataCleaningService(hotel_metadata).clean(dirty_bookings_df)
stats = DescriptiveStatsService().compute(clean_df)
# Must not raise: every value has to be a plain JSON-compatible type.
json.dumps(stats)
def test_compute_includes_numeric_and_categorical_columns(dirty_bookings_df, hotel_metadata):
clean_df = DataCleaningService(hotel_metadata).clean(dirty_bookings_df)
stats = DescriptiveStatsService().compute(clean_df)
assert "mean" in stats["adults"]
assert "top" in stats["meal"] or "unique" in stats["meal"]