11 KiB
CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Repo layout
.
├── cms/ Go service: video ingestion/CMS (implemented)
├── discovery/ Go service: empty placeholder, no code yet
├── infrastructure/ Pulumi (Go) program provisioning all AWS resources
└── .gitea/workflows/ Gitea Actions CI/CD pipelines
discovery is provisioned in infra (its own ECR repo, ECS service, ALB,
Postgres database/role) but has no application code yet — don't assume it's
deployable.
Commands
cms (Go 1.25, module thamanyah/cms/v2)
cd cms
go run . # serves on :8081 (requires DB_*/S3_*/MEDIACONVERT_* env vars — see below)
go run . migrate # applies pending DB migrations, then exits (no HTTP server)
go build ./...
go vet ./...
There are no test files in this repo (cms, discovery, or infrastructure) — don't assume a test suite exists.
Views are written as .templ files (github.com/a-h/templ) and compiled to
*_templ.go. If you edit a .templ file, regenerate its Go code with the
templ generate CLI before building (not installed in this environment by
default — install via go install github.com/a-h/templ/cmd/templ@v0.3.1020
to match go.mod, or check for an existing binary first).
infrastructure (Go, Pulumi, module thamanyah)
cd infrastructure
pulumi preview # plan changes against stack "main"
pulumi up # apply — this touches real AWS resources, confirm with the user first
pulumi stack output # e.g. ecsClusterArn, cmsServiceArn
Deploys run in CI (.gitea/workflows/infrastructure-deploy.yml) on push to
main touching infrastructure/**. Treat local pulumi up as something to
confirm with the user, not a routine dev command — it mutates shared cloud
state and Pulumi state isn't safe to update concurrently with CI.
Architecture
cms service
Plain net/http (Go 1.22+ pattern-based ServeMux), no framework. Entry
point cms/main.go wires up AWS SDK v2 clients (S3, MediaConvert) and a
Postgres connection from env vars and assigns them to package-level interface
vars in internal/services (services.S3Client, services.MediaConvertClient,
services.DB) — handlers call through these interfaces, and
S3Concrete/MediaConvertConcrete/DBConcrete are the only implementations,
which is what makes the handlers testable even though no tests exist yet. On
boot, S3Concrete.AssertSuccessfulConnection proactively exercises
head/put/get/presign against the bucket, and DBConcrete.AssertSuccessfulConnection
pings Postgres — both panic on failure rather than letting the service come
up in a broken state.
cms/main.go has two entry paths, dispatched on os.Args[1]: the default
path (runServer) boots the HTTP server; ./cms migrate (runMigrate)
only opens the DB connection, applies pending migrations via
cms/internal/db.Migrate (golang-migrate, iofs source, SQL files embedded
from cms/internal/db/migrations/*.sql), and exits — it does not touch
S3/MediaConvert or start the server. This is run as its own ECS container
before the main container starts (see infrastructure below), so runServer
never runs migrations itself, only AssertSuccessfulConnection.
Routes (cms/main.go): GET /, GET /health, GET /videos/new,
POST /videos/presign, POST /videos, static files under /static/.
Views live in cms/internal/views (templ components) with a shared
layouts.Layout wrapper; types.go holds view-model structs like
VideoMetadata used by the upload-success page.
Data model (Postgres, cms/internal/db/migrations/)
videos— one row per uploaded video:title,description,tags(free-text, comma-separated — not normalized),file_name,storage_key(the S3 key, unique),mediaconvert_job_id,status(written once as"processing"on insert — see Known gaps),size_bytes, timestamps.idis aUUIDgenerated application-side as a UUIDv7 (uuid.NewV7()inservices.DBConcrete.CreateVideo) rather than via the column's ownDEFAULT gen_random_uuid()(which generates v4 and is never actually relied on) — v7 keeps primary-key inserts roughly time-ordered, avoiding B-tree fragmentation as the table grows.categories— a small fixed lookup table (documentary,news,entertainment,podcast,other), seeded by migration0001.video_categories— join table (video_id,category_id, composite PK,ON DELETE CASCADE) added in migration0002: a video can belong to multiple categories, not just one.CreateVideoinserts thevideosrow and itsvideo_categorieslinks inside a single transaction.
Video upload → transcode pipeline
- Browser calls
POST /videos/presign→ cms returns a presigned S3PUTURL forraw-uploads-bucket, keyvideos/<random-hex>.<ext>. - Browser
PUTs the file directly to S3 (requires the bucket's CORS rule, set up ininfrastructure/main.go). - Browser calls
POST /videoswith the metadata + key → cms callsMediaConvertClient.QueueEncodingJob(key), submitting a MediaConvert jobs3://raw-uploads-bucket/<key>→s3://encoded-bucket/<key>(H.264/AAC → MP4, QVBR rate control — QVBR requiresMaxBitrateto be set explicitly), thenservices.DB.CreateVideopersists thevideosrow (status"processing") and itsvideo_categorieslinks. - Finished output lands in
encoded-bucket, served via CloudFront.
Config wiring: cms reads S3_BUCKET, MEDIACONVERT_INPUT_BUCKET,
MEDIACONVERT_OUTPUT_BUCKET, MEDIACONVERT_ROLE_ARN, AWS_REGION,
DB_HOST, DB_PORT, DB_NAME, DB_USER, DB_PASSWORD from env vars
injected by the ECS task definition (extraEnv and the unconditional DB_*
vars in deployFargateService, infrastructure/main.go) — main.go panics
on boot if any required var is empty.
infrastructure (infrastructure/main.go, single Pulumi Go program, region us-east-1, stack main)
- Postgres: one shared RDS instance (
db.t3.micro, single-AZ, no backups — intentionally minimal). Each app (cms,discovery) gets its own login role and same-named database via thepostgresqlprovider (newServiceDatabase), so services never share DB credentials. - ECS Fargate: one cluster (
app-cluster), one ALB per service (each gets its own DNS name rather than sharing a load balancer on different ports).deployFargateService(...)is the shared helper building a service's ECR repo, CloudWatch log group, task definition, ECS service, and ALB.taskRoleis optional (nil = no AWS identity beyond the shared execution role);extraEnvappends container env vars beyond the DB_* set;runMigrations(true forcms, false fordiscovery) adds a second, non-essential<name>-migratecontainer to the task — same image,command: ["migrate"]— with the main container'sdependsOnset tocondition: "COMPLETE"on it. This is ECS's container-dependency mechanism, the Fargate equivalent of a Kubernetes init container: ECS runs the migrate container to completion (exit 0) before starting the main container, so schema migrations always finish before the service accepts traffic. No separate CI/Docker migration step exists —Dockerfile'sENTRYPOINT ["./cms"]plus the container'scommandoverride composes to./cms migrate. - S3 + CloudFront:
encoded-bucketholds finished transcoded output, served publicly via CloudFront using Origin Access Control (OAC) — the bucket itself blocks all public access; only CloudFront's OAC principal can read it. - S3 raw uploads:
raw-uploads-bucketis a separate, private bucket for pre-transcode uploads — deliberately kept apart fromencoded-bucketso raw source video is never reachable through the public CDN. CORS is scoped toPUTonly, from thecmsALB's own origin (browser uploads directly via presigned URL). - IAM roles — four distinct roles/users, each scoped narrowly, don't
conflate them:
ecs-task-execution-role— shared by both services' ECS agent (image pull, log write, Secrets Manager read for DB password). Not usable by application code inside the container.cms-task-role— thecmscontainer's own AWS identity: S3ListBucket/PutObject/GetObjectonraw-uploads-bucketonly,mediaconvert:CreateJob, andiam:PassRolescoped to the MediaConvert service role (iam:PassedToServicecondition).discoveryhas no task role — it doesn't touch S3 or MediaConvert.mediaconvert-service-role— trusted bymediaconvert.amazonaws.com, not by ECS; the role MediaConvert itself assumes (passed asCreateJobInput.Role) to readraw-uploads-bucketand writeencoded-bucket. Distinct fromcms-task-roleby design: one is "cms calling AWS", the other is "AWS calling AWS on cms's behalf".gitea-ci-user— an IAM user (static access keys, not OIDC — the Gitea Actions runner doesn't support instance-profile auth) scoped to just ECR push (cms/discoveryrepos) andecs:UpdateService/ecs:DescribeServiceson the two ECS services. Never broaden this toecr:*/ecs:*.
CI/CD (.gitea/workflows/)
cms-deploy.yml: push tomaintouchingcms/**. Builds/pushes the Docker image to ECR usinggitea-ci-user, installs the AWS CLI (not preinstalled on the runner image — via AWS's official install script, not a third-party action), thenaws ecs update-service --force-new-deployment. Cluster/service ARNs come frompulumi stack outputand are set as repo variables (not secrets — ARNs aren't sensitive).infrastructure-deploy.yml: push tomaintouchinginfrastructure/**. Runspulumi upusing a separate, broader AWS credential (PULUMI_AWS_ACCESS_KEY_ID/SECRET) thangitea-ci-user, since provisioning IAM/RDS/ECS/CloudFront needs wider permissions than pushing images and forcing deployments. Guarded with a concurrency group since Pulumi state isn't safe to update concurrently.- The runner's
ubuntu-latestlabel maps to a docker image configured on the runner host (outside this repo) — currently minimal, lacking the AWS CLI, hence the manual install step incms-deploy.yml.
Known gaps
cms/internal/services/mediaconvert.gonever callsDescribeEndpointsand configures no custom MediaConvert endpoint — relies on the SDK's default regional endpoint.- No MediaConvert completion webhook or poller exists — a
videos.statusrow is written once as"processing"inCreateVideoand never updated, even after the transcode job actually finishes or fails. discoveryhas infra provisioned (ECR repo, ECS service, ALB, Postgres DB/role) but no application code — it doesn't touch its database at all.- No automated tests exist for
cms,discovery, orinfrastructure.