Updated 27 September 2026. The original July version described JSONL as the only format, treated several optional fields as required, said omitted products disappeared immediately and predated WooCommerce's native MCP developer preview. On 26 September those claims were corrected against the current OpenAI and WooCommerce documentation. On 27 September Bridge 1.0.138 aligned Bridge's generator and validator with the current stable product schema and revalidated the real 95-row Project2209 feed.
1. The model: push, not pull
This is an onboarding and delivery integration, not ordinary web crawling. The documented pipeline is:
- Prepare product records against the current stable product schema, with one row per purchasable item or variant.
- Apply through OpenAI's merchant onboarding. Product-feed onboarding is currently limited to approved partners, and approval attaches to the integration rather than to the plugin that generated a draft.
- Choose the approved delivery method: a full snapshot uploaded to OpenAI's SFTP service, or the product-feed API. OpenAI recommends starting validation with a sample of about 100 items.
- Keep it current: OpenAI generally recommends a complete file snapshot at least daily and API updates during the day. Reuse stable filenames for SFTP snapshots.
- Remove deliberately: set
is_eligible_search=falsefor prompt removal on the next processing pass. Omission from later snapshots can take up to 14 days to expire.
A download URL exposed by your own tooling remains useful for inspection, but it is not the documented delivery integration. Bridge does not submit the application, upload to SFTP or call the OpenAI product-feed API automatically.
2. What the spec actually requires
For file upload, OpenAI currently prefers Parquet and also accepts jsonl.gz, csv.gz and tsv.gz. Bridge prepares jsonl.gz. The stable OpenAI-format schema requires nine basic fields; other fields are optional or conditional:
| Field | Rule | Real-world gotcha |
|---|---|---|
item_id, title, description, url | Required | Keep the item identifier stable and use the actual purchasable product or variant URL |
brand, seller_name | Required | Use real merchant data, never placeholders or a guessed brand |
image_url | Required public URL | The image must represent the item or selected variant |
price, availability | Required | Use major currency units and the documented availability enum |
group_id, listing_has_variations, variant_dict | Conditional for variants | Each purchasable selection gets its own row and item ID |
additional_image_urls | Optional array in JSONL | The comma-separated representation belongs to CSV/TSV cells, not JSONL |
gtin | Optional | If supplied, it must be exactly 8, 12, 13 or 14 digits with a valid check digit |
seller_url, returns and market fields | Optional or setup-dependent | Country columns alone do not enable a market; confirm the integration setup with OpenAI |
3. Brand: the field that decides your feed size
On paper, brand is just another required field. On real WooCommerce stores it is the gatekeeper, because most Woo stores have never filled a brand taxonomy — the data lives in product titles ("Wolford Velvet 66…") and in category names, not in a machine-readable field. Shopify sidesteps this structurally: its product model has a vendor field that defaults to the store name, so every Shopify product has something in that slot by construction.
What we learned running this on a real multi-brand catalog (~1,100 SKUs, luxury lingerie and beachwear):
- The Woo core
product_brandtaxonomy existed but was empty — while 20+ real brands lived as top-level product categories. - The honest fix is on the data side, under the merchant's authority: assign the real brands (the merchant did — 22 brands, 1,129 product assignments, sourced from their own category names). No inference, no guessing by the tooling.
- For genuinely own-label or unbranded stores, "seller as brand" is legitimate — it is literally how Shopify's vendor default and most Etsy listings work. For a multi-brand retailer it would be false data on the exact field shoppers search by. Tooling should offer it only as an explicit, merchant-declared opt-in.
- For products with no real brand data, do not fabricate one and do not call the row conformant. The current stable OpenAI-format schema requires
brand; the honest result is to keep that row out of a deliverable feed until the merchant supplies the source data.
4. Numbers from real generations
These are historical July 2026 generation results, useful as catalog-quality measurements but not evidence of conformance to today's stable schema. The generator writes atomically, so a failed run does not destroy the previous snapshot:
- A small mixed catalog (~110 products, food and apparel): 86 conformant rows, 28 rows excluded for missing primary image, 0 schema errors.
- The ~1,100-SKU multi-brand catalog above, after the merchant assigned real brands: full-catalog feed with real brand values, 39 rows excluded for missing images, 0 schema errors — first pass.
- Generation cost matters at scale: a naive per-product implementation took ~20 seconds on 1,100 SKUs (thousands of N+1 queries) and would hit PHP execution limits near 5,000 SKUs; batch-priming product, meta and term caches collapsed it to a few seconds. If you build your own generator, budget for this.
5. Discovery-first: the checkout stays yours
The product-feed schema treats discovery and checkout as separate scopes. Every discovery row points to the merchant's product URL; checkout requires a separately enabled integration, and setting is_eligible_checkout alone does not complete that onboarding. A clean discovery feed therefore does not transfer catalog authority or checkout control away from WooCommerce.
6. The honest disclaimers nobody writes
- Approval is entirely OpenAI's decision, per merchant, with no published timeline. Being on the waitlist costs nothing; nobody can promise you acceptance.
- Shopping-surface availability varies by market. The feed format is region-neutral; where results actually show is OpenAI's call.
- A conformant feed is necessary, not sufficient. It gets you to the door with clean data; relevance does the rest.
- Products excluded from the feed (missing image, for instance) are not invisible to AI in general — agents that read your site directly, or structured catalog APIs, still see them. The feed is one channel among several.
7. What you can do today
WooCommerce now includes native MCP support in developer preview. That is an authenticated surface for WooCommerce operations, not an OpenAI product-feed delivery path. KaliCart Bridge continues to expose a separate public, read-only catalog surface and a feed-readiness workflow.
Bridge 1.0.138 (released 27 September): the optional JSONL generator and bundled validator now follow the current stable schema. Products without a merchant-declared brand or configured own-label fallback are excluded, counted and listed; additional images are emitted as a list; optional store-policy fields no longer block generation; GTINs, prices, ratings and variant options are validated against the documented contract. Approved partners should still validate the exact deliverable and use the SFTP or API method enabled for their integration.
The one-line takeaway: treat this as an approved integration and data-quality process: verify the current schema, fix the source data, validate the exact deliverable, and keep it fresh.
Current primary sources
- OpenAI onboarding and integration methods: Get Started — Agentic Commerce
- File delivery, formats and retention: File Upload overview
- Stable product fields: Products schema
- API delivery: Product-feed API overview
- WooCommerce native MCP status: WooCommerce MCP Integration