Files
thamanyah/CLAUDE.md
T
FahdShalhoub 98950aba20
Build, Push and Deploy Discovery / build-push-deploy (push) Successful in 1m45s
Deploy Infrastructure / pulumi-up (push) Successful in 2s
FEAT: Init Discovery Service
2026-08-27 22:03:41 +03:00

341 lines
18 KiB
Markdown

# CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
## Repo layout
```
.
├── cms/ Go service: video ingestion/CMS (implemented)
├── discovery/ Go service: catalogue discovery/read side (skeleton)
├── infrastructure/ Pulumi (Go) program provisioning all AWS resources
├── tests/ Go module: Gherkin/gobdd scenarios run against a live cms
└── .gitea/workflows/ Gitea Actions CI/CD pipelines
```
`discovery` is a skeleton, not a feature: it mirrors `cms`'s layout, boot
sequence and Postgres wiring, and serves `GET /health` plus the Swagger UI, but
it has no domain yet — `internal/models` and `internal/db/repositories` are
package doc comments, and migration `0001_init` creates no tables (it exists
only because `internal/db/migrate.go` embeds `migrations/*.sql`, which will not
compile against an empty directory). It has no `internal/services` and no
`internal/consumers`: those are AWS-only in `cms`, and `discovery` has no task
role, so it reaches nothing but its own database. It is deployable — its ECR
repo, ECS service and ALB have always existed in infra, and it now has an image
to put in them.
## Commands
### cms (Go 1.25, module `thamanyah/cms/v2`)
```bash
cd cms
go run . # serves on :8081 (requires DB_*/S3_*/MEDIACONVERT_* env vars — see below)
go run . migrate # applies pending DB migrations, then exits (no HTTP server)
go build ./...
go vet ./...
```
### discovery (Go 1.25, module `thamanyah/discovery`)
Same commands as `cms`, and the same two entry paths (`runServer`,
`runMigrate`) in `discovery/main.go` — only the port and the dependencies
differ. Note the module path carries no `/v2`: unlike `cms`, this module has
never had a v1.
```bash
cd discovery
go run . # serves on :8080 (requires DB_* env vars only — no S3/MediaConvert)
go run . migrate # applies pending DB migrations, then exits (no HTTP server)
go build ./...
go vet ./...
```
Its OpenAPI spec is generated into `discovery/docs` by the same `swag init`
invocation as `cms`, run from `discovery/`. In `docker-compose.yml` it is a
service of its own on `127.0.0.1:8080`, wired to the `discovery` database and
role LocalStack provisions — and, like `cms`, started in server mode only, so a
freshly created local database needs `docker exec discovery ./discovery migrate`
once.
`cms`, `discovery` and `infrastructure` hold no test files of their own. The
only tests in the repo are the black-box BDD scenarios in `tests/` — see below.
`cms` is a JSON API only — it serves no HTML and has no static assets. The
OpenAPI spec is generated from swaggo annotations on the handlers into
`cms/docs`, a committed, compiled-in Go package. If you add or change a
handler, its annotation comments, or a request/response struct, regenerate it:
```bash
go install github.com/swaggo/swag/cmd/swag@v1.16.6 # match go.mod; not installed by default
swag init --generalInfo main.go --dir ./ --parseInternal --output ./docs
```
`--parseInternal` is required — the handlers live under `internal/`, which
swag skips without it. Keep the `swaggo/swag` version in `go.mod` and the CLI
in lockstep: `http-swagger/v2` transitively pulls a much older `swag` whose
`swag.Spec` struct lacks the `LeftDelim`/`RightDelim` fields newer generators
emit, and the build breaks outright if the two drift. (The CLI's `--version`
misreports itself as v1.16.4; `go version -m $(go env GOPATH)/bin/swag` gives
the real one.)
### tests (Go, module `thamanyah/tests`)
```bash
cd tests
go test ./... # runs features/*.feature against CMS_BASE_URL (default http://localhost:8081)
go test -v ./... # -v prints the Gherkin: gobdd nests a subtest per feature/scenario/step
CMS_BASE_URL=… go test ./...
```
Black-box BDD covering the video upload feature: its own Go module importing
nothing from `cms/`, talking to a running service over HTTP only, so the same
scenarios run against compose and against a deployed environment. Skips (does
not fail) when nothing is serving. Uploading is a plain `PUT` to the presigned
URL, no AWS SDK. Scenarios can be tagged `@known-gap` to pin current behaviour
that differs from the documented contract — see `tests/README.md`.
Note compose starts `cms` in server mode only: the migrate container exists in
the ECS task definition, not in `docker-compose.yml`, so a freshly created
local database needs `docker exec cms ./cms migrate` once or every scenario
fails on `relation "categories" does not exist`.
### infrastructure (Go, Pulumi, module `thamanyah`)
```bash
cd infrastructure
pulumi preview # plan changes against stack "main"
pulumi up # apply — this touches real AWS resources, confirm with the user first
pulumi stack output # e.g. ecsClusterArn, cmsServiceArn
```
Deploys run in CI (`.gitea/workflows/infrastructure-deploy.yml`) on push to
`main` touching `infrastructure/**`. Treat local `pulumi up` as something to
confirm with the user, not a routine dev command — it mutates shared cloud
state and Pulumi state isn't safe to update concurrently with CI.
## Architecture
### cms package layout
Four `internal` packages, split by the kind of thing they hold — keep new code
on the same seams:
| package | holds |
|---|---|
| `internal/api` | the **wire contract**: request/response structs with their JSON + swaggo tags. No logic. |
| `internal/models` | the **domain types** the rest of the code passes around: `Video`, `Category`, `VideoStatus`. No JSON tags, no SQL. |
| `internal/handlers` | HTTP: decode `api.*`, validate, call repositories/services, encode `api.*`. |
| `internal/db` | the `*sql.DB` connection lifecycle and migrations; `internal/db/repositories` holds all SQL. |
| `internal/services` | **AWS only** — S3 and MediaConvert. Nothing database-related lives here. |
`api` and `models` are deliberately separate types even where the fields look
alike: `api.CompleteResponse` is the published schema, `models.Video` is the
row. Handlers translate between them field by field, so renaming a `models`
field doesn't move the published API and vice versa.
### cms service
Plain `net/http` (Go 1.22+ pattern-based `ServeMux`), no framework. Entry
point `cms/main.go` builds the dependencies and assigns them to package-level
interface vars that the handlers call through:
- `services.S3Client` (`services.S3`) ← `*services.S3Concrete`
- `services.MediaConvertClient` (`services.MediaConvert`) ← `*services.MediaConvertConcrete`
- `repositories.VideoRepo` (`repositories.VideoRepository`) ← `*repositories.ConcreteVideoRepository`
- `repositories.CatagoriesRepo` (`repositories.CatagoriesRepository`) ← `*repositories.ConcreteCatagoriesRepository`
Each interface has exactly one implementation; the indirection is what makes
the handlers substitutable in tests, even though no tests exist yet. Note the
spelling: the categories repository is `Catagories`/`CatagoriesRepo` in
`repositories/catagories.go` (and `handlers.validateCatagoryIDS`) — the
misspelling is load-bearing for compilation, so match it rather than "fixing"
it piecemeal.
On boot, `S3Concrete.AssertSuccessfulConnection` proactively exercises
head-bucket/put/get/presign against the bucket (writing a throwaway
`.s3-connectivity-check` object), and `db.AssertSuccessfulConnection` pings
Postgres — both panic on failure rather than letting the service come up in a
broken state.
`cms/main.go` has two entry paths, dispatched on `os.Args[1]`: the default
path (`runServer`) boots the HTTP server; `./cms migrate` (`runMigrate`)
only opens the DB connection, applies pending migrations via
`cms/internal/db.Migrate` (golang-migrate, `iofs` source, SQL files embedded
from `cms/internal/db/migrations/*.sql`), and exits — it does not touch
S3/MediaConvert or start the server. This is run as its own ECS container
before the main container starts (see infrastructure below), so `runServer`
never runs migrations itself, only `AssertSuccessfulConnection`.
The DB connection is opened once for the process lifetime by
`db.CreateDBConnection(connString)` and closed by `db.CloseConnection` — both
plain functions over `*sql.DB` in `internal/db/client.go`, not methods on a
wrapper type. Repositories take that `*sql.DB` as their `SQLDB` field.
Routes (`cms/main.go`): `GET /health`, `GET /api/categories`,
`POST /api/videos/presign`, `POST /api/videos`, plus Swagger UI at
`GET /swagger/` (`/swagger/doc.json` serves the spec). The UI assets are
embedded in the binary by `swaggo/files`, so nothing is read from disk and
nothing is fetched from a CDN at runtime.
Handlers live in `cms/internal/handlers``handlers.go` holds `Health` and
the shared response writers, `videos.go` the categories and upload endpoints.
The wire structs they serve live in `cms/internal/api` (`api.go` for
`HealthResponse` and `ProblemDetails`, `videos.go` for `Category`,
`CategoriesResponse`, `PresignRequest`/`PresignResponse` and
`CompleteRequest`/`CompleteResponse`), referenced by the handlers' swaggo
annotations as `api.CompleteResponse` and so on — so the generated spec's
definition names track that package, and renaming a type there changes the
published schema names.
There is no view layer: the `internal/views` templ package, the `static/`
directory, and the htmx frontend were all removed when the service became a
JSON API, along with the `templ` dependency.
Error responses are RFC 9457 Problem Details objects
(`application/problem+json`), written by
`writeProblem(w, status, title, detail)`. `type` is always `"about:blank"`;
`title` is a short summary held identical across every occurrence of a given
problem, so clients can branch on it; `detail` is the only member that varies
with request data. Success responses go through `writeJSON`
(`application/json`). Both share `writeJSONContent`. Note this deviates
slightly from RFC 9457, which pairs an `about:blank` type with a title that is
just the HTTP status phrase — meaningful titles like these are supposed to
carry a real `type` URI. Adding per-problem type URIs is the conforming fix if
it ever matters.
### Data model (Postgres, `cms/internal/db/migrations/`, two migrations: `0001`, `0002`)
- `videos` — one row per uploaded video: `title`, `description`, `tags`
(free-text, comma-separated — not normalized), `file_name`, `storage_key`
(the S3 key, unique), `mediaconvert_job_id`, `status` (written once as
`"processing"` on insert — see Known gaps), `size_bytes`, timestamps.
`id` is a `UUID` filled by the column's own `DEFAULT gen_random_uuid()`
(v4) and read back through `RETURNING id` — the application does not
generate it. (An earlier revision generated UUIDv7 application-side; that
was reverted in `785154c`, so primary keys are random, not time-ordered.)
`status` is a plain `TEXT NOT NULL DEFAULT 'processing'` column with **no
CHECK constraint** — the closed set exists only in Go, as the
`models.VideoStatus` string type in `cms/internal/models/video.go`
(`processing`, `ready`, `failed`). Extending it means adding a constant
there and extending the `enums` annotation on `api.CompleteResponse.Status`
before regenerating the spec. (The doc comment on `VideoStatus` claims a
`videos_status_check` constraint enforces it in the database — that
constraint does not exist; nothing has ever created it.)
- `categories` — a small fixed lookup table (`documentary`, `news`,
`entertainment`, `podcast`, `other`), seeded by migration `0001`. Its
`SMALLSERIAL` ids are part of the public API: `GET /api/categories` returns
`{id, name}` pairs (ordered by name, not id) and `POST /api/videos` takes
`categoryIds`, so the seed order in migration `0001` is what fixes which id
means which name — never renumber it.
- `video_categories` — join table (`video_id`, `category_id`, composite PK,
`ON DELETE CASCADE`) added in migration `0002`: a video can belong to
*multiple* categories, not just one. `ConcreteVideoRepository.CreateVideo`
inserts the `videos` row and its `video_categories` links inside a single
transaction, taking the category ids as given — `handlers.validateCatagoryIDS`
is what checks them against `CatagoriesRepo.ListCategoriesIDs` before the
transcode job is queued, leaving the foreign key and the composite PK as
backstops (see Known gaps for how those failures surface).
### Video upload → transcode pipeline
1. Client calls `POST /api/videos/presign` → cms returns a presigned S3 `PUT`
URL for `raw-uploads-bucket`, key `videos/<random-hex>.<ext>`. The
submitted `contentType` (only `video/mp4` or `video/quicktime`) is signed
into the URL, so the client's `PUT` must send the identical header.
2. Client `PUT`s the file directly to S3 (from a browser this requires the
bucket's CORS rule, set up in `infrastructure/main.go`,
its allowed origin is now stale). The file never passes through cms.
3. Client calls `POST /api/videos` with the metadata (categories given as
`categoryIds` from `GET /api/categories`) + key → cms validates the
category ids, calls `MediaConvertClient.QueueEncodingJob(key)`, submitting
a MediaConvert job `s3://raw-uploads-bucket/<key>`
`s3://encoded-bucket/<key>` (H.264/AAC → MP4, QVBR rate control — QVBR
requires `MaxBitrate` to be set explicitly), then
`repositories.VideoRepo.CreateVideo` persists the `videos` row (status
`"processing"`) and its `video_categories` links.
4. Finished output lands in `encoded-bucket`, served via CloudFront.
Config wiring: cms reads `S3_BUCKET`, `MEDIACONVERT_INPUT_BUCKET`,
`MEDIACONVERT_OUTPUT_BUCKET`, `MEDIACONVERT_ROLE_ARN`, `AWS_REGION`,
`DB_HOST`, `DB_PORT`, `DB_NAME`, `DB_USER`, `DB_PASSWORD` from env vars
injected by the ECS task definition (`extraEnv` and the unconditional DB_*
vars in `deployFargateService`, `infrastructure/main.go`) — `main.go` panics
on boot if any required var is empty.
### infrastructure (`infrastructure/main.go`, single Pulumi Go program, region `us-east-1`, stack `main`)
- **Postgres**: one shared RDS instance (`db.t3.micro`, single-AZ, no
backups — intentionally minimal). Each app (`cms`, `discovery`) gets its
own login role and same-named database via the `postgresql` provider
(`newServiceDatabase`), so services never share DB credentials.
- **ECS Fargate**: one cluster (`app-cluster`), one ALB *per service* (each
gets its own DNS name rather than sharing a load balancer on different
ports). `deployFargateService(...)` is the shared helper building a
service's ECR repo, CloudWatch log group, task definition, ECS service,
and ALB. `taskRole` is optional (nil = no AWS identity beyond the shared
execution role); `extraEnv` appends container env vars beyond the DB_* set;
`runMigrations` (true for both services) adds a second,
non-essential `<name>-migrate` container to the task — same image,
`command: ["migrate"]` — with the main container's `dependsOn` set to
`condition: "COMPLETE"` on it. This is ECS's container-dependency
mechanism, the Fargate equivalent of a Kubernetes init container: ECS runs
the migrate container to completion (exit 0) before starting the main
container, so schema migrations always finish before the service accepts
traffic. No separate CI/Docker migration step exists — `Dockerfile`'s
`ENTRYPOINT ["./cms"]` plus the container's `command` override composes to
`./cms migrate`.
- **S3 + CloudFront**: `encoded-bucket` holds finished transcoded output,
served publicly via CloudFront using Origin Access Control (OAC) — the
bucket itself blocks all public access; only CloudFront's OAC principal
can read it.
- **S3 raw uploads**: `raw-uploads-bucket` is a separate, private bucket for
pre-transcode uploads — deliberately kept apart from `encoded-bucket` so
raw source video is never reachable through the public CDN. CORS is
scoped to `PUT` only, from the `cms` ALB's own origin — correct back when
cms served the upload page itself, but stale now that it serves no UI (see
Known gaps).
- **IAM roles** — four distinct roles/users, each scoped narrowly, don't
conflate them:
- `ecs-task-execution-role` — shared by both services' ECS *agent* (image
pull, log write, Secrets Manager read for DB password). Not usable by
application code inside the container.
- `cms-task-role` — the `cms` container's own AWS identity: S3
`ListBucket`/`PutObject`/`GetObject` on `raw-uploads-bucket` only,
`mediaconvert:CreateJob`, and `iam:PassRole` scoped to the MediaConvert
service role (`iam:PassedToService` condition). `discovery` has no task
role — it doesn't touch S3 or MediaConvert.
- `mediaconvert-service-role` — trusted by `mediaconvert.amazonaws.com`,
not by ECS; the role MediaConvert itself assumes (passed as
`CreateJobInput.Role`) to read `raw-uploads-bucket` and write
`encoded-bucket`. Distinct from `cms-task-role` by design: one is "cms
calling AWS", the other is "AWS calling AWS on cms's behalf".
- `gitea-ci-user` — an IAM **user** (static access keys, not OIDC — the
Gitea Actions runner doesn't support instance-profile auth) scoped to
just ECR push (`cms`/`discovery` repos) and
`ecs:UpdateService`/`ecs:DescribeServices` on the two ECS services.
Never broaden this to `ecr:*`/`ecs:*`.
### CI/CD (`.gitea/workflows/`)
- **`cms-deploy.yml`**: push to `main` touching `cms/**`. Builds/pushes the
Docker image to ECR using `gitea-ci-user`, installs the AWS CLI (not
preinstalled on the runner image — via AWS's official install script, not
a third-party action), then `aws ecs update-service --force-new-deployment`.
Cluster/service ARNs come from `pulumi stack output` and are set as repo
*variables* (not secrets — ARNs aren't sensitive).
- **`discovery-deploy.yml`**: the same pipeline for `discovery`, on pushes
touching `discovery/**`. One difference: the ECS service ARN is read from
the `DISCOVERY_ECS_SERVICE_ARN` repo variable rather than pinned in the
workflow, so it has to be set (from `pulumi stack output
discoveryServiceArn`) before the first deploy can succeed.
- **`infrastructure-deploy.yml`**: push to `main` touching
`infrastructure/**`. Runs `pulumi up` using a separate, broader AWS
credential (`PULUMI_AWS_ACCESS_KEY_ID`/`SECRET`) than `gitea-ci-user`,
since provisioning IAM/RDS/ECS/CloudFront needs wider permissions than
pushing images and forcing deployments. Guarded with a concurrency group
since Pulumi state isn't safe to update concurrently.
- The runner's `ubuntu-latest` label maps to a docker image configured on
the runner host (outside this repo) — currently minimal, lacking the AWS
CLI, hence the manual install step in `cms-deploy.yml`.