Building PR Previews Across Domains and Data Profiles on Fly.io
Keeping preview URLs aligned with the code, site and data profile QA intended to test.

QA needed to review a pull request across a storefront and several CMS-driven sites, then repeat the review against either of two Shopify stores. A generated deployment URL could show the code running, but it did not reproduce the hostname behavior and store selection those checks depended on.
Our server-rendered Remix application ran on Fly.io. GitHub Actions already coordinated code checks and Docker-based deployments. Hostnames selected the logical site within the application, while each Shopify store brought its own data and credentials. Extending previews meant carrying both kinds of context through the deployment and the request.
I designed and owned the preview architecture and delivery. After directing the initial per-PR implementation, I built out store-profile selection, wildcard subdomain routing with Fly Replay, and the deployment safeguards described here. I built this around QA's need to start and repeat those reviews. The team now uses these previews daily.
The resulting system has one permanent routing application and one disposable application per PR. GitHub PR labels select the preview's data profile or request its removal. Code checks, label changes and cleanup reach a shared workflow that verifies current PR state before changing Fly. A preview serves one selected profile at a time. Profile changes deliberately allow downtime, and external systems remain outside the isolation provided by its compute.
Domains, application names and site labels below are illustrative. The pseudocode is reconstructed to explain decisions, not copied from private source.
This account is written for frontend, platform and software engineers working with server-rendered applications, CI/CD, ephemeral environments or delivery infrastructure. I assume familiarity with pull requests, containers, and basic hostname and proxy behavior.
At the center is one engineering question: how do you keep a PR preview aligned with the code revision, hostname-derived site and selected data profile when those can change independently? The architecture treats code eligibility, request destination and configuration acceptance as separate concerns, each with its own validation boundary.
From a running branch to a useful QA environment
The first version solved a useful, narrow problem: create a Fly application for a PR and give reviewers a URL. As the requirements expanded, QA needed more of the real application's behavior.
Consider a reviewer checking a change that touches both the storefront and a CMS site. They need to open both surfaces on the same PR, follow links without leaving it, and sometimes repeat the check against the other store's configuration. Meanwhile, the PR may receive another commit, its base branch may advance, or someone may change its preview label.
That creates four connected questions:
Which PR application should receive this request?
Which logical site and data profile should that application use?
Which revision is eligible to be deployed there?
Should that PR still have a preview, and which Fly app may the automation change?
The hostname can answer the first two only if the deployment behind it remains consistent with those answers.
Fly already provides a review-apps approach for creating temporary applications. The custom work here joined that lifecycle to our hostname-dependent application and store configuration. Existing Fly hosting was a project constraint; moving platforms was not part of the work.
Other approaches would have different costs. A generated app hostname is sufficient when the application does not depend on several logical hosts. Provisioning domains and certificates for each PR could retain that behavior, but would add certificate and DNS work to creation and cleanup. I used permanent wildcard domains so that opening a PR changed the deployment, without requiring a new domain setup. That comparison explains the tradeoff within our existing Fly setup.
GitHub PR labels became the preview controls
Reviewers control the environment through three labels attached to the pull request in GitHub. They select a desired state; a label change does not bypass validation or guarantee an immediate deployment.
| GitHub PR labels | Requested preview state |
|---|---|
| Neither store-profile label | Staging-store profile, the default |
preview-store:stage only |
Explicit staging-store profile |
preview-store:prod only |
Production-store profile |
| Both store-profile labels | Remove the preview until the ambiguity is resolved |
skip-preview, with any other labels |
Remove the preview and suppress deployment |
Other PR labels, such as needs-review or bug, do not participate in this decision.
For example, suppose PR 123 is open on a branch that allows previews, its current code has passed validation, and no store-profile label is present:
The normal PR workflow deploys the staging-store profile. QA opens
pr-123.preview.example.com.A reviewer adds
preview-store:prod. The label workflow reads all current labels, confirms that the current revision has valid checks, and requests a profile change. It replaces the runtime of the same PR app. After success, the comment points topr-123.preview.example.net.Someone adds
needs-review. That label cannot change the selected profile, restart code validation, or displace a pending preview-label run.The reviewer adds
skip-preview. Cleanup removes PR 123's app even if the production-store label remains.Removing
skip-previewmakes the production-store profile eligible again. Deployment still requires current, successful validation. Removingpreview-store:prodas well restores the staging-store default.
Switching from an explicitly labeled stage profile takes two label edits. If the reviewer adds the prod label before removing the stage label, both are briefly present. A run that observes that state requests removal. A run that starts after the stage label has been removed sees only prod. The policy follows the label set it reads, rather than assuming each intermediate edit must become a deployment.
These controls change the data and credential configuration of an ephemeral app. Our Fly organization also contains a stable production app and two stable staging/pre-production apps. Selecting preview-store:prod never selects one of those as the deployment target; it means production Shopify data with pre-production domain behavior inside the PR app. This distinction is essential when both kinds of app live in the same organization.
Permanent domains in front of disposable apps
Each supported profile has its own preview base domain. These illustrative addresses show the pattern:
| Address | Meaning |
|---|---|
pr-123.preview.example.com |
PR 123, storefront, staging-store profile |
editorial--pr-123.preview.example.com |
PR 123, editorial site, staging-store profile |
pr-123.preview.example.net |
PR 123, storefront, production-store profile |
The double hyphen matters. Both the site label and PR identifier sit inside one DNS label. A certificate for *.preview.example.com covers the first two addresses. It would not cover editorial.pr-123.preview.example.com, which introduces another label. That follows the wildcard matching rules in RFC 9525.
The two wildcard domains terminate at a permanent Fly router application. Fly's custom-domain support provides the DNS and certificate setup; the router supplies the application-specific mapping.
The router accepts only recognized preview bases and hostname shapes. It also checks the site label against a fixed allowlist. It extracts the PR number and identifies the corresponding Fly app, such as the illustrative preview-123. Both profile domains lead to that same per-PR app.
The router responds with a Fly Replay header instructing Fly Proxy, Fly's HTTP proxy, to send the original request to the selected application. The browser keeps the preview URL. Our router chooses the destination; Fly's infrastructure performs the replay. See Fly's dynamic request routing documentation.
Permanent wildcard domains reach one router; Fly Proxy replays the original request to the PR application, which resolves the logical site and selected configuration.
This keeps the router small: it holds no Shopify credentials and does not query GitHub on each request. Adding a PR does not add another route entry. Adding a supported profile or site still requires explicit configuration.
The permanent router is also a shared dependency. Unsupported hosts receive a 404. If replay falls back because the target is unavailable, the router returns a 503 instead of issuing another replay and creating a loop. The per-PR Machines, the VMs that run each application, can stop when idle, so startup remains part of preview availability. Fly's documented 1 MB replay limit also constrains requests through this path.
The application had to understand preview hosts
Getting the request to the right container did not finish the work. The application needed to understand the hostname it received.
The real application used hostnames to choose CMS models and construct links. An encoded preview host therefore had to resolve to the same logical site while retaining its PR-specific address. I added host parsing and URL helpers so that the storefront and supported CMS sites could share a per-PR application. Known authored links also needed rewriting to stay within the corresponding preview.
Two domains now served different purposes: the application's configured base domain selected environment behavior, while the browser's preview host preserved the reviewer's location. Treating those as interchangeable would send links out of the preview or resolve the wrong site.
The store profile was another explicit choice. In this project it meant selecting one of two Shopify data and credential sets, together with the matching domain configuration. Selecting the production-store profile did not deploy the PR to the production storefront. It ran a separate preview application connected to that store.
QA can switch between stores, but this design does not keep both profiles running side by side for a PR.
That distinction matters operationally. Separate compute does not make production data disposable or make every operation read-only. Reviewers still need to understand which store they are using.
Cookies needed similar care. I scoped preview cookies to their preview base rather than issuing them for the storefront apex. That supports sessions across a PR's logical sites, but cookies under one preview base are also shared across different PRs on that base. This does not provide per-PR session isolation or a complete browser sandbox. QA needs a clean browser context when session carryover could change the result.
Six workflows coordinated the preview lifecycle
There are six workflows in the preview infrastructure, with different responsibilities. A workflow is the event entry point or reusable definition; its jobs are the units GitHub schedules. The normal code workflow also handles stable branch deployments, but those use separate branch-gated jobs.
| Workflow responsibility | Entry point and jobs relevant to previews |
|---|---|
| Code validation and preview dispatch | PR opened, reopened or synchronized after a code push. After the required checks, Preview label check reads live labels; Preview deploy or Preview skip cleanup calls the shared workflow. |
| Label changes | PR labeled / unlabeled events. Two jobs: Resolve preview label state, then Reconcile Preview when removal or a validated deployment is required. |
| Shared PR reconciliation | Called by other workflows. One preview job serializes operations for that PR, reads live state and either deploys, removes or leaves the app unchanged. |
| Policy cleanup | Trusted PR-target events. One cleanup caller handles closure, adding skip, and retargeting or reopening onto a branch that disallows previews. |
| Stale-app cleanup | Weekly schedule or manual dispatch. Discover selects canonical PR apps; a reconcile matrix sends each PR through the shared workflow. |
| Permanent router deployment | Manual dispatch. One job creates the fixed router app if needed and deploys its dedicated Dockerfile and Fly configuration. |
The main code path therefore has three preview orchestration jobs in addition to its checks. They are not three independent ways to mutate Fly: both deploy and skip cleanup delegate the mutation decision to the same reusable workflow.
GitHub provides labeled and unlabeled activities, but the pull_request trigger does not offer a label-name filter. The label workflow checks the changed label's name in its job condition. An unrelated label can create a lightweight skipped workflow run; it does not launch a replacement code-check suite. See GitHub's PR event reference.
I also separated concurrency groups at this entry point. The three preview-control labels share a group for their PR. Unrelated labels get unique run groups, so a burst of review or status labels cannot replace pending preview-control work. Once a relevant run starts, it reads the complete live label set. The event tells it whether to wake up; the current PR tells it what to do.
The shared deployment path then selects one of two explicit secret bindings, stage or prod. Both call the same deployment script. That keeps credential selection visible while keeping the build, runtime replacement and deployment sequence in one place.
Code, label and cleanup paths converge on one per-PR reconciliation job. Each executed job checks current intent before changing Fly; the permanent router is deployed separately.
The router's deployment is deliberately separate. Opening a PR does not redeploy the shared router or change wildcard DNS and certificates. A router code or accepted-domain change needs its own manual deployment; deploying a PR copy of the application does not update that permanent service.
The shared deployment path also leaves maintenance across components. Accepted domain and site mappings exist in both the router and application. Preview policy is read by the event workflows and checked again by the reusable workflow. Changes to those contracts have to stay aligned. Centralizing Fly mutations gave the lifecycle one coordination point, but it did not turn every configuration and policy definition into a single source.
Deployment eligibility depended on the tested merge revision
Once labels could change the store profile, the deployment workflow needed to distinguish a request to deploy from permission to deploy particular code.
The normal code path establishes eligibility after its required type, unit and end-to-end checks pass. The label path reuses that evidence. Reuse is useful only if the checks still describe the code about to be deployed.
A PR has a head commit, but CI can test a generated merge commit that combines that head with its base branch. GitHub documents that distinction for the pull request event.
Suppose head H passed checks as part of merge revision M1. The base branch then advances, producing M2, while H stays unchanged. Deploying the latest merge ref because “the PR head passed” would deploy a different integration result.
I bound the successful validation marker to both revisions. The code workflow publishes an artifact named for the PR and tested merge revision. The label workflow requires an unexpired matching artifact associated with the current head, plus successful current GitHub Actions checks. Where checks are sharded, every matching non-skipped check must have succeeded; one green shard is insufficient.
Missing, pending, failed, expired or mismatched validation stops label-triggered deployment. If a label changes while code CI is running, the code workflow reads live labels after its checks finish, so that path can pick up the selected profile.
Deployment checks out the recorded merge revision rather than resolving a moving branch ref later. If the shared workflow finds that the base advanced and replaced that merge revision, it leaves the existing preview in place and updates the PR comment to say it was not refreshed. Advancing the base does not automatically revalidate every open PR. A new code-validation run is needed before deploying the new integration result.
That choice sometimes makes QA wait for fresh validation. It preserves a more useful meaning for a preview update: the source revision being deployed is the revision whose checks established eligibility.
The marker validates source. The deployment subsequently builds its container from that revision; it does not prove that the final container was itself the artifact exercised by every CI check. Nor does reusing code checks validate every selected store credential or integration.
Queued workflows needed current PR intent
Revision matching alone does not resolve competing workflow runs. A label event, a new commit and a cleanup request can all refer to the same PR.
I put deployment and destruction through a shared per-PR concurrency group. After acquiring that slot, the workflow reads the PR again and derives its desired state from current information: whether it is open, its target branch, repository, labels, selected profile and revisions.
The event starts reconciliation. It does not remain the authority for a later mutation.
This simplified pseudocode shows the decision boundary. It omits API retries, marker retrieval, detailed eligibility rules and platform commands; it is not runnable workflow code.
caller establishes validated head + merge revision for deployment
caller passes its expected action, profile and revisions
within the PR's shared deployment slot:
current = read_current_PR()
desired = derive_preview_state(current)
if desired requires removal:
remove_only_this_PR_preview()
else if request_is_destroy_only:
leave_current_preview_unchanged()
else if desired permits deployment
and caller_has_validated_deployment
and request_matches_current_profile_and_revisions:
deploy_recorded_merge_revision()
else:
leave_current_preview_unchanged_and_report()
For example, a cleanup request may wait while someone removes the skip label. When it acquires the slot, the live PR may be valid again. The cleanup caller then leaves the app alone; it cannot turn itself into a deployment because it did not bring a validated deployment request. Conversely, an old deploy caller can remove the app if the PR is now closed or skipped.
An unknown mergeability result is retried briefly because GitHub calculates it asynchronously. After waiting, the workflow reparses the full PR state, including labels, base and revisions. A conflict or still-uncertain merge result defers deployment. It does not, by itself, mean an existing preview should be deleted.
The concurrency group prevents simultaneous mutations by these workflows for the same PR. It does not promise that every event runs in order. GitHub's concurrency behavior can replace an older pending run with a newer one; we do not cancel an already running deployment.
There is also no transaction spanning GitHub and Fly. A label can change after the live-state read while an image is building. A later eligible reconciliation must handle that change. The design reduces stale decisions at the mutation boundary; it cannot promise that the runtime reflects a label change instantly.
Stable and preview apps shared one Fly organization
The hostname allowlist and the Fly app safeguards solve different problems. The router limits which requests it will forward. Workflow guards limit which app the automation may create, deploy or destroy.
Because production, staging and preview apps share a Fly organization, cleanup cannot mean “delete anything with preview in its name.” The target is constructed from a positive PR number. In the illustrative naming used here, PR 123 can target only preview-123.
The protections are repeated at the boundaries where they matter:
Discovery: the scheduled sweep accepts only the anchored canonical name pattern, equivalent here to
^preview-[1-9][0-9]*$. A stable staging app or permanent router is not a candidate.Shared workflow: it independently validates the PR number and constructed app identity, and explicitly rejects the stable production, staging and router app names. A caller cannot supply an arbitrary app name.
Deployment script: it checks the canonical name again and requires exact agreement with the PR number. Profile, domain and required configuration checks run before building or draining.
Fly operations: commands use that exact app name. App-list reads are scoped to the configured organization, and deletion requires the exact app to be present. A failed GitHub or Fly lookup stops the operation; it is not interpreted as permission to delete.
For example, a sweep can discover preview-123, but discovery alone does not authorize removal. The shared workflow must still read PR 123 under its concurrency slot and find that a no-preview policy applies. The explicit protected-name check is a backstop alongside the narrow naming contract, not a substitute for it.
These are protections against accidental selection by this pipeline. The organization-scoped Fly token is still a broader capability, and the preview receives its selected store's credentials. These checks do not sandbox intentionally modified workflow code or make production-store operations harmless. Repository and workflow write access remain part of the trust boundary.
Profile changes accepted downtime to avoid mixed state
The deployment script had a second consistency problem: changing credentials and changing code are separate operations.
If a secret update starts a deployment of the old image, a profile switch can briefly run old code with the newly selected store's credentials. For a disposable QA environment, I preferred a period of unavailability over that mixed state.
The final sequence is deliberate:
Validate the selected profile, domain mapping, PR app identity and required configuration.
Build and push the Docker image from the recorded merge revision, using a unique tag for that run and revision.
Remove the existing preview Machine.
Stage the selected profile's configuration and secrets, including removal of an obsolete production-site flag.
Deploy the already built image with the staged values into one new preview Machine.
Building first protects the existing preview from build failure. Removing the old runtime before activating the new profile trades availability for a consistent code-and-configuration replacement.
Fly's staged secret import lets the script set values without starting a separate deployment. The final deployment names the image built before the runtime transition.
Building first preserves an existing running preview if the build fails. A new Fly app record may already have been created, but an existing preview has not yet been drained.
After the old Machine is removed, failure has a different consequence: the preview can remain unavailable until a successful deployment or retry. This is an ordered replacement with accepted downtime, not an atomic rollback mechanism or a zero-downtime production release design.
The old profile URL needed a runtime check
Permanent wildcard routing creates one more case to handle.
After PR 123 switches profiles, both wildcard domains can still route to its app. The router intentionally does not maintain live profile state. Without an application check, an old staging-profile URL could reach a runtime now connected to the production store.
I added a runtime check before static assets and Remix request handling. A recognized preview domain must match the preview base configured in that deployment. A mismatch returns 421, with responses marked non-cacheable. Missing or invalid preview-domain configuration also rejects requests on those known preview hosts.
The destination app therefore checks the part of the address that the router cannot establish from a PR number: whether this domain belongs to its active profile.
The scope is precise. This guard covers recognized custom preview domains. The raw Fly hostname remains available, and the guard is neither authentication nor a complete host allowlist. Its job is to stop an old custom preview address from silently presenting the newly selected profile.
Cleanup and the QA handoff completed the lifecycle
A preview should disappear when the PR no longer qualifies for one, including closure or merge, explicit skipping and retargeting onto a branch where previews are disabled. Eligible same-repository PRs can deploy; fork PRs do not receive this secret-backed preview path.
The policy cleanup workflow covers a specific event gap. Ordinary pull_request workflows do not run while a PR has merge conflicts. A trusted pull_request_target path can still respond to closure, adding skip-preview, or retargeting/reopening onto a disallowed base. It uses default-branch workflow code and never checks out or executes PR source. That is why it is cleanup-only. See GitHub's event model.
The filters matter here too. An ordinary title or body edit does not trigger removal. The workflow checks whether an edit changed the base branch and whether the new base disallows previews. Returning to an allowed base does not automatically validate that new merge result; a later code-validation run is still required.
This trusted path is narrow: it does not implement every label transition on a conflicted PR. For example, conflicting store labels are normally handled by the label workflow. The weekly or manually invoked sweep gives remaining no-preview states another reconciliation opportunity, so event coverage is not a claim of immediate convergence for every transition.
The sweep has two jobs. Discovery lists apps in the organization and extracts PR numbers only from canonical preview names. A matrix then reconciles each candidate through the shared per-PR workflow, with up to four PRs in parallel and without canceling the other candidates when one fails. It removes apps when live policy requires absence, keeps valid open previews, and stops a candidate on an API error. It does not deploy new code or use app age alone as deletion authority.
The handoff is part of the delivered workflow. A sticky PR comment records the successful preview's profile, URLs and source revisions, along with values needed for manual account setup. That gives QA a specific environment to identify. Workflow logs record the resolved operation, profile and reason for leaving an app unchanged or removing it. A stale base revision also produces a warning and an updated comment.
A failed deployment can still leave the previous success comment in place. In that case, the workflow result matters when deciding whether the latest change is ready; the comment alone is not a current deployment-health signal.
Account authentication registration remained a separate manual setup step. The workflow surfaced the callback, logout and origin values that QA and engineering needed, but deploying a preview did not register them automatically, and deleting the Fly app did not remove those external registrations. Shopify documents the relevant configuration in its Headless account setup guide.
What changed, and what a preview still cannot prove
The change from the initial version is concrete:
| Earlier capability | Added capability |
|---|---|
| A temporary app and a generated PR URL | Permanent wildcard addresses for the storefront and supported CMS sites |
| One configured store environment | Explicit selection between two store profiles, one active per PR |
| A deployment triggered by PR activity | Deployment eligibility tied to recorded validation and current PR state |
| A preview creation path | Shared deployment/removal coordination, runtime domain checks and cleanup reconciliation |
Regression coverage asserts the router's host parsing and fallback behavior, logical-site URL mapping, preview cookie domains, and rejection of a custom host that disagrees with the runtime profile. The runtime tests also cover missing configuration and the intended behavior of raw Fly and non-preview hosts. These assertions target application contracts; they do not establish the success of a live profile transition across GitHub, Fly and the selected store.
What I can substantiate is the expanded review capability: one PR can now be checked across the storefront and supported CMS sites, against either store profile, with explicit controls around revision eligibility, profile changes, runtime acceptance and cleanup. The implemented mechanisms and regression coverage support those capabilities and boundaries. This case does not include a quantified comparison of QA time or hosting cost.
Preview fidelity still has boundaries. Redis caching and Convert experimentation are disabled in this preview setup, so it cannot establish production cache behavior or end-to-end experiment attribution. Shared cookies and manual account configuration can affect a QA session. Connecting to a production store retains real external dependencies and their effects. A healthy application process is not proof that every store, account or checkout journey is ready.
Those limits belong in the handoff because they determine which conclusions a reviewer can draw from a successful test.
The lesson I took forward
The first deployment gave QA somewhere to open a branch. Extending it meant defining what that address represented as the code, hostname and configuration changed.
Fly supplied the application runtime, domain support and request replay. My work connected those capabilities to the application's site model and the PR's deployment lifecycle, then made the failure boundaries explicit.
The part I would carry into another system is the separation of code eligibility, request destination and configuration acceptance. Here the configuration came from two Shopify stores. Another application might choose between different data sources or service accounts, but it would still need to define that mapping and its side effects. A useful preview lets a reviewer understand what they are testing and when the environment can no longer support that conclusion.



