# Site Clones — Master Instructions

These rules apply to every job. Always read the job-level CLAUDE.md for per-job variables.

---

## Output Structure

```
[job-folder]/
├── index.html
├── assets/
│   ├── css/
│   ├── js/
│   ├── fonts/
│   ├── images/
│   └── media/
├── assets-new/        ← new media assets sent by the team (images, videos, etc.)
├── brief.pdf          ← optional — sent by the team for media asset placement context
├── ERRORS.md          ← created only if unresolved console errors exist
└── CLAUDE.md          ← per-job variables
```

---

## UI/UX Fidelity — Standing Rule

This applies throughout every phase. **Do NOT change any of the following:**
- Layout, spacing, or visual hierarchy
- CSS rules, class names, or inline styles
- Responsive breakpoints and media queries
- JavaScript behavior, animations, scroll effects, or interactions
- Font loading (Google Fonts `<link>`, Adobe Fonts, or self-hosted `@font-face`) — keep all intact
- Third-party UI libraries (Swiper, GSAP, Lottie, etc.) — keep all intact
- z-index, overflow, and positioning rules

The cloned page must be visually and functionally identical to the original on desktop, tablet, and mobile.

---

## Phase 1 — Rip & Clean

### Step 1 — Rip the Page

- Use `wget` to spider and download the page and all linked assets:
  ```bash
  wget --mirror --convert-links --adjust-extension --page-requisites --no-parent -P [job-folder] [URL]
  ```
- Move the downloaded HTML to `index.html` at the job root
- Move all other files (CSS, JS, images, fonts, media) into `assets/` with subfolders as above
- Rewrite all asset references in `index.html`, CSS, and JS to reflect new paths under `assets/`
- Verify no references still point to the original domain

**No external CDN — download everything locally.** After moving files, scan `index.html` (and any CSS/JS) for remaining external image/font/media URLs (e.g. `cloudfront.net`, `imgix.net`, CDN subdomains). Download each file into the appropriate `assets/` subfolder with `wget -q [url]` and replace the URL in the HTML/CSS with the local `assets/…` path. The finished clone must not depend on any external CDN for visual assets.

### Step 2 — Strip Tracking & Analytics

Remove all of the following from ALL `.html` and `.js` files:

**Script tags and inline code for:**
- Google Analytics / GTM (ga.js, analytics.js, gtag.js, GTM, Analyzely)
- Meta Pixel (fbevents.js, facebook.com/tr)
- Snap Pixel (sc-static.net/scevent.min.js, `snaptr(`)
- Taboola pixel (archive-digger.com/assets/script_tag.js, taboolaId)
- Hotjar, Intercom, Segment, Mixpanel, Heap, FullStory, Clarity, Amplitude
- TripleWhale / Triple Pixel (config-security.com)
- Elevate A/B testing (elevateab.app.txt, ds0wlyksfn0sb.cloudfront.net)
- Shoplift A/B testing (`<!-- Start of Shoplift scripts -->` ... `<!-- End of Shoplift scripts -->`)
- Klaviyo, ReCharge, Okendo (tracking calls only — keep Okendo CSS/widget HTML for visual fidelity)
- Any `<script>` tags loading from third-party analytics or tracking domains
- Any `fetch()`, `XHR`, or `navigator.sendBeacon()` calls to analytics endpoints
- Cookie consent banners tied to analytics (OneTrust, Cookiebot, etc.)

**Shopify platform scripts to remove (non-functional locally):**
- `window.Shopify.SignInWithShop?.initShopCartSync` — triggers `/cart.js` fetches
- `window.Shopify.featureAssets` — loads shop-cart-sync, checkout-modal, etc.
- Kaching Bundles app block (`shopify://apps/kaching-bundles`) and its `kaching-loader.js` script tag
- Cloudflare email-decode (`/cdn-cgi/scripts/.../email-decode.min.js`)
- Any other `shopify://apps/` embed blocks that are purely functional (cart, upsell, checkout)

**Also remove:**
- `<noscript>` pixel tags (Meta, GTM, etc.)
- Hidden tracking `<img>` tags (1x1 pixels)
- Any `data-*` attributes that reference tracking IDs

---

## Phase 2 — Apply Team Changes

*Run this phase only when the team provides changes. All steps are optional — apply only what is provided and skip the rest.*

### Step 3 — Read Brief (if present)

- **Text changes** arrive as plain text in the thread — this is preferred. Apply them directly with no PDF needed.
- **`brief.pdf`** is used when the team needs to show visual asset placement (which new image/video goes in which section). If present, read it to extract the asset swap map. Also check for any text changes it may contain.
- If the PDF shows slides side-by-side (original vs new), treat left as original and right as replacement.
- If neither is present, skip to Phase 3.

### Step 4 — Brand, Product & Text Changes

**Brand/product sweep** — only if Original brand and Our brand are present in the per-job CLAUDE.md (not blank or `(TBD)`). Replace across ALL files (HTML, JS, CSS, meta tags, JSON-LD, OG tags):
- `[ORIGINAL_BRAND]` → `[OUR_BRAND]`
- `[ORIGINAL_PRODUCT]` → `[OUR_PRODUCT]`

If branding is missing, skip the brand/product sweep and leave a note in `ERRORS.md` that branding is pending.

**Specific text changes from thread or brief** — apply each one precisely:
- Match the exact string in the HTML (use the original text from the page as the search key)
- Replace with the exact new copy provided
- Preserve surrounding HTML tags, classes, and formatting — only the text content changes
- If a text change affects a button label, also check if the CTA link needs updating (cross-reference Step 5)

**Also update:**
- `<title>` tag
- `<meta name="description">`
- Open Graph tags (`og:title`, `og:description`, `og:site_name`)
- Twitter card tags
- JSON-LD structured data (name, brand fields)
- Alt text on brand images/logos

### Step 5 — CTA Link Replacement

Read CTA URL from the job-level CLAUDE.md. If it is missing or `(TBD)`, skip this step entirely and leave a note in `ERRORS.md` that the CTA URL is pending.

- Replace ALL primary CTA button `href` values with `[OUR_CTA_URL]`
- Check for JS-driven navigation (onClick handlers, router pushes, window.location assignments) and replace those too
- Check for form `action` attributes pointing to original domains
- Do NOT change navigation links, footer links, or non-CTA internal links

**Goto() / indirect CTA pattern** — many pages use a `Goto()` function that builds a redirect URL and a link-rewriting script that sets all non-footer `<a>` tags to `href='#'` with a click listener calling `Goto()`. Remove this pattern entirely and replace with direct links:
1. Delete the `Goto()` function, `getQueryString()`, and `GetRequest()` helpers
2. In the link-rewriting loop, replace the `else` branch (the one that was calling `Goto()`) with: `currentA.href = '[OUR_CTA_URL]'; currentA.target = '_blank';`
3. Leave `privacy-link`, `articles_links`, and `image-misalignment-link` branches unchanged

### Step 6 — Media Asset Replacement

Use the asset swap map from `brief.pdf` (Step 3) and/or the table in the per-job CLAUDE.md.

- New assets are in `assets-new/` (sent by the team — any format: JPG, PNG, MP4, etc.)
- For each swap: copy the new asset into the correct `assets/` subfolder keeping its own filename — do not rename it to match the old file. Update every HTML/CSS reference to the old filename to point to the new one instead.
- If the brief shows the new asset visually placed in a specific section of the page, match it to the image in that section in the HTML
- Preserve `width`, `height`, and `alt` attributes on replaced `<img>` tags
- If no new assets are provided, skip this step

**A message with both attached files and text is one instruction — act on both together.** By the time you're invoked, the Slack bot has already saved any attachments into `assets-new/` — that part isn't your job. What matters is: if the message text describes a change (e.g. "swap the hero image for this"), that text is the instruction for the file(s) that just landed alongside it. Check `assets-new/` for recently-added files whenever you're given a change request, match them to what the text describes, and apply the change in the same turn. Never respond as if the files were a bare upload with no instruction (e.g. "saved, let me know what to do") when the instruction was the message you were just given.

**Explicit swap request with an attached file — apply it, don't just acknowledge it.** If a message names the target directly ("swap the hero image for this", "change the product shot to this one", etc.) and an attached file came with it, treat the target as already resolved. Find the file in `assets-new/` and immediately perform the swap into the named target in that same turn — copy it into place, verify with a screenshot, and report the result as done. Do not stop after confirming the file is there and wait for a separate "apply it" / "go ahead" message — that's a second round-trip the instruction already made unnecessary.

Do not second-guess an explicit instruction because the new asset's *content* looks unexpected or off-brand for the page (e.g. it doesn't visually match the product, theme, or other images). If the message clearly names the target and provides the file, that's authorization enough — apply it and move on. Content mismatch is not grounds to pause and confirm; only pause if the target placement itself is genuinely ambiguous (see above).

---

## Phase 3 — Final QA

### Step 7 — Console Error Audit

1. **Always serve via HTTP — never open `index.html` directly as a `file://` URL.** Fonts and cross-origin scripts will be blocked by CORS on `file://`. Use:
   ```bash
   python3 -m http.server 8080
   ```
   Then open `http://localhost:8080` in the browser.
2. Open DevTools → Console and Network tabs. **Scroll the full page** to trigger lazy-loaded assets and scroll-event JS.
3. Fix ALL of the following:
   - Broken asset paths (404s for CSS, JS, fonts, images)
   - Mixed content warnings (http vs https)
   - Missing files referenced in CSS (`url()` paths)
   - CORS errors on self-hosted assets
   - JS runtime errors fired on scroll/resize — add null guards (`if (!el) return;`) inside event handler functions that query DOM elements by class/ID
   - Image src paths with `&width=N` instead of `?width=N` — replace `&` with `?` (caused by CDN URL construction when the base path already contained a query param)
   - JS errors caused by removed tracking scripts (wrap in try/catch or remove dependent code cleanly)
4. For errors that cannot be resolved (e.g. Okendo reviews API, Shopify cart/checkout), document them in `ERRORS.md`
5. Goal: zero critical errors. Non-critical warnings are acceptable if documented.

---

## Next.js Sites — Special Handling

Next.js pages cannot be cloned like a regular static site. The framework requires specific inline scripts and JS chunks to hydrate — removing or misrouting them produces a blank page or 404.

### How to detect Next.js

Page HTML contains any of: `__NEXT_DATA__`, `__next_f`, `/_next/static/`, `next/dist`, or `_next/` in script `src` attributes.

### Step 1 — wget adjustments for Next.js

The standard `wget --mirror` command mishandles Next.js chunk paths because they live under `_next/` and are content-hashed. Use this instead:

```bash
wget --mirror --convert-links --adjust-extension --page-requisites --no-parent \
  --reject "*.map" -P [job-folder] [URL]
```

After downloading, verify that `_next/static/chunks/` exists and contains `.js` files. If the directory is missing or empty, manually download the missing chunks by inspecting the Network tab (filter by JS, copy the chunk URLs, wget each one).

### Step 2 — Scripts you must NEVER remove on Next.js pages

Even though the cleaning step removes inline `<script>` blocks, these are **off-limits** on Next.js pages:

| Pattern | Why it must stay |
|---|---|
| `<script id="__NEXT_DATA__" type="application/json">` | Contains initial props, build ID, and route data. Removing it kills hydration entirely — blank page. |
| `<script src="/_next/static/...">` | Framework and page chunks. Removing any of these breaks the React app. |
| Any inline `<script>` referencing `__NEXT_F`, `__NEXT_P`, `self.__next_f`, `__webpack_require__`, or `webpackChunk` | Next.js runtime internals — not tracking. |
| `<script id="__NEXT_FONT_MANIFEST__">` or similar Next.js manifest scripts | Font and route manifests needed by the router. |

### Step 3 — Path rewriting for `_next/` assets

When rewriting asset paths after moving files, preserve the `_next/` directory name exactly. Do **not** move `_next/` contents into `assets/` — keep them at:

```
[job-folder]/_next/static/chunks/
[job-folder]/_next/static/css/
[job-folder]/_next/static/media/
```

References to `/_next/` in HTML and JS should become `_next/` (relative, no leading slash) when served from the job root.

**After copying files, fix subpath references in both `index.html` and all JS chunks.** Next.js sites hosted under a subpath (e.g. `/breezebox/`) embed that subpath as hardcoded strings in two places that `wget --convert-links` does NOT fix:

1. **Turbopack chunk loader** — one of the `_next/static/chunks/turbopack-*.js` files contains a string like `"/[subpath]/_next/"` that it uses as the base URL for dynamically loading all other chunks. Find and replace it:
   ```bash
   # Find which chunk has it:
   grep -rl '/[subpath]/_next/' _next/static/chunks/
   # Fix it (replace with just "_next/"):
   sed -i 's|"/[subpath]/_next/"|"_next/"|g' _next/static/chunks/turbopack-*.js
   ```
   If this string is wrong, the page loads blank — React chunks never load.

2. **Page component chunk** — one of the `_next/static/chunks/*.js` files contains a variable assignment like `e="/[subpath]"` (or `t="/[subpath]"`, `r="/[subpath]"`) that is used as the base URL prefix for all CSS, images, JS, and other static assets via template literals (`${e}/css/...`, `${e}/images/...`). Blank it out:
   ```bash
   # Find which chunk has it:
   grep -rl 'e="/[subpath]"' _next/static/chunks/
   # Fix it:
   sed -i 's|e="/[subpath]"|e=""|g' _next/static/chunks/[page-chunk].js
   ```
   With `e=""`, all asset paths become root-relative (e.g. `/css/pre/bootstrap.min.css`). You must then download those assets from the origin and place them at matching paths under the job root.

**Also apply the same `/[subpath]/` → `/` replacement in `index.html`** for any occurrences in inline `__next_f` script blocks (these also bypass `--convert-links`):
```python
content = content.replace('/[subpath]/_next/', '_next/')
content = content.replace('/[subpath]/', '/')
```

**Downloading assets referenced via the page chunk base variable:**
After blanking `e`, grep the page chunk for all `${e}/...` patterns to build a complete list of CSS, JS, images, fonts, and media the page loads:
```bash
grep -oh '\${e}/[^"'\''`,) ]*' _next/static/chunks/[page-chunk].js | sort -u
```
Download each file from the origin into the matching local path. Common structure: `css/[slug]/`, `images/[slug]/`, `js/[slug]/`. If font files (`.woff2`, `.woff`, `.ttf`) 404 on the origin, download them from cdnjs using the version found in the CSS file header.

### Step 4 — API route failures (expected, non-fixable)

Next.js pages often call `/api/...` endpoints on load to fetch dynamic content. These will 404 locally — that is expected and cannot be fixed without a running server. Document them in `ERRORS.md`.

### Diagnostic checklist — blank page / 404

Open DevTools before assuming a tracking script was wrongly removed:

1. **Elements tab** — is `<script id="__NEXT_DATA__">` present? If missing, the cleaning step removed it. Restore it from the original downloaded file.
2. **Console tab** — look for:
   - `Hydration failed` → `__NEXT_DATA__` is missing or corrupted
   - `Cannot read properties of undefined (reading 'call')` → a webpack chunk is missing (404 on a `_next/static/chunks/*.js` file)
   - `ChunkLoadError` → same as above — a JS chunk didn't load
3. **Network tab (filter: JS)** — are any `_next/static/chunks/*.js` files returning 404? If so, re-download the missing chunks and place them at the correct relative path.
4. **Network tab (filter: Fetch/XHR)** — `/api/` calls returning 404 are expected and non-critical. Document them in `ERRORS.md`.
5. **Page shows text but CSS/images 404** — the page chunk base variable (`e="/[subpath]"`) was not blanked. Fix it as described in Step 3.

---

## GemPages Pages — Local Asset Path Fix

If the source page is built with GemPages (Shopify app), the downloaded `assets/js/gp-lazyload.js` contains a broken `A()` function that prepends `https://` to relative paths, turning `assets/images/foo.jpg` into `https://assets/images/foo.jpg` (treating `assets` as a hostname). This breaks all locally-hosted images and media.

After downloading, find and replace this exact function in `gp-lazyload.js`:

**Find:**
```
let A=(e,t)=>{try{e.startsWith("http://")||e.startsWith("https://")||(e="https://"+e);let n=new URL(e);return n.searchParams.delete(t),n.toString()}catch(t){return console.error("Error occurred while removing query by key:",t),e}}
```

**Replace with:**
```
let A=(e,t)=>{try{var isRel=!e.startsWith("http://")&&!e.startsWith("https://");var base=isRel?window.location.origin+"/"+e.replace(/^\//,""):e;let n=new URL(base);n.searchParams.delete(t);return isRel?n.pathname.replace(/^\//,"")+n.search+n.hash:n.toString()}catch(t){return console.error("Error occurred while removing query by key:",t),e}}
```

**How to detect GemPages:** page HTML contains `gp_lazyload`, `gps-link`, `base-src`, or `gp-global.js`.

---

## Final Checklist Before Done

- [ ] `index.html` exists at job root
- [ ] All assets under `assets/` with no broken paths
- [ ] No tracking/analytics scripts remaining
- [ ] Brand and product names fully replaced — OR noted as pending in `ERRORS.md`
- [ ] All text changes from thread or brief applied (if provided)
- [ ] CTA links replaced — OR noted as pending in `ERRORS.md`
- [ ] New media assets in place and matched to correct page sections (if provided)
- [ ] Page looks identical on desktop, tablet, mobile
- [ ] Zero critical console errors
- [ ] `ERRORS.md` created if any unresolved issues or pending items
