$ cat methodology.md

Methodology & Glossary

A programmatic infrastructure log for marketing AI. Every claim is sourced. Unknowns are marked Unknown rather than guessed.

#Methodology

aistacktrack is a structured registry of vendors selling AI to marketing teams. For each vendor we capture a fixed schema — architecture, data posture, agentic level, pricing, integrations, channels, and company facts — so that any two vendors can be compared on equal footing.

Research is a hybrid pipeline. An automated scraper pulls each vendor's site (pricing, security, about, docs). An LLM extracts structured fields against a strict schema with a no-hallucination instruction: missing evidence must yield null, not a guess. Every extracted field is paired with a source URL and an excerpt in our vendor_attributions store.

Editors review enriched profiles before they go live, and we re-scan on a rolling cadence so pricing changes, new compliance certifications, and product repositioning surface within days.

#Sources & attribution

Every field on a vendor profile can be traced to a source. We prioritize, in order:

  1. Vendor site, docs, and trust/security pages (primary)
  2. Vendor SOC 2 / ISO / privacy policy artifacts
  3. Independent analyst reports (Gartner, Forrester)
  4. Review platforms (G2, TrustRadius, Capterra)
  5. Reputable press / news
  6. Community signal (Reddit, HN) — clearly flagged

Each citation carries a confidence rating (high, medium, low) so you can weigh self-reported claims against third-party validation.

#Proprietary scores

Four ownable scores summarize each vendor. All are deterministic functions of the structured fields you can already see — nothing is crowd-sourced, paid-for, or hidden. Recomputed daily. Each score on a vendor profile links back here.

Transparency Index (0–100)

How much of our schema the vendor has actually disclosed, weighted by source quality. Vendor-primary sources (vendor site, docs, privacy policy / DPA) count full credit; third-party sources count 60%; disclosed but unsourced fields count 50%. Highest weights go to pricing transparency, compliance, data governance, integrations, and notable customers.

Reads: 75+ = well-documented; 50–74 = average; under 50 = thin disclosure. This is a disclosure score, not a quality score — a stealth vendor with a great product can score low until they publish.

Agentic Autonomy (0–100)

How real the "agent" claim is. Combines stated autonomy level (copilot / semi-autonomous / fully agentic), breadth of channels acted on, depth of integrations, and whether the vendor charges on outcomes — a signal of confidence in autonomous performance.

Buyer Risk (green / amber / red)

Procurement's first question, scored from compliance certifications, funding stage, company age, team size, and data-handling posture. Green = enterprise-ready signal; amber = viable with diligence; red = early-stage or thin disclosure — proceed with eyes open.

Change Velocity (trailing 90 days)

Count of changelog entries detected by our nightly scanner in the last 90 days — pricing changes, new compliance, repositioning, new integrations. High velocity means the product is moving; zero may mean stable, dormant, or unindexed.

None of these scores are reviews. They measure observable signals about how a vendor presents itself and what we can detect — not whether you should buy.

#How we treat unknowns

Where a vendor does not publish a fact and no third party corroborates it, we render Unknown instead of a plausible-sounding guess. This is deliberate: a marketing director comparing four CDPs should be able to trust that a populated cell is sourced and that a blank cell is genuinely unknown — not a scraping failure dressed up as data.

#Corrections & claims

Vendors can claim their profile to suggest edits — every change still requires a source. Anyone can submit a correction with a link; we re-verify before publishing. Claim a profile →

$ cat glossary.md

Glossary

Every term we use across the matrix, vendor pages, and comparisons. Deep-link any definition.

#Architecture terms

§Pure Wrapper
Thin layer over a third-party foundation model. No proprietary weights.
§Fine-Tuned
Customized/fine-tuned models on top of base providers.
§Sovereign AI
Owns the full model stack and inference infrastructure.
§Copilot
Suggests; human approves every action.
§Semi-Autonomous
Acts within guardrails; human reviews outputs.
§Fully Agentic
Plans and executes multi-step workflows end-to-end.
§Model stack
The chain of models a vendor uses — foundation model(s), any fine-tuned layers, retrieval, and orchestration. Disclosed when the vendor publishes it.
§Inference infrastructure
Where model calls actually execute (vendor-hosted GPUs, a hyperscaler, or your own VPC). Matters for latency, residency, and cost pass-through.

#Data & governance

§Synthetic Only
Trains and operates on synthetic data; no PII ingress.
§Zero-Copy Warehouse
Runs inside your warehouse. No data leaves your perimeter.
§External Processing
Data is sent to vendor infrastructure for processing.
§PII (Personally Identifiable Information)
Any data that can identify an individual — emails, names, device IDs, IPs. Drives compliance scope.
§Zero retention
Vendor commits in writing that prompts and outputs are not stored or used for training. Verify in the DPA, not the marketing page.
§Subprocessor
A third party (often a foundation-model provider) that processes your data on the vendor's behalf. Listed in the subprocessor log.
§Compliance certifications
Independent attestations — SOC 2 Type II, ISO 27001, HIPAA, GDPR readiness. We record only certifications with a public artifact.

#Commercial terms

§Usage-based
Pay per token, message, generation, or compute unit. Scales with success but harder to forecast.
§Seat-based
Pay per active user/seat. Predictable; can penalize broad enablement.
§Outcome-based
Pay per resolved ticket, qualified lead, conversion. Aligns incentives; rare and complex to audit.
§Enterprise-only
No self-serve tier; pricing is bespoke and contract-driven.
§Hybrid
Mix of seats + usage or platform fee + consumption.
§Pricing transparency: Public
Numeric pricing published on the website.
§Pricing transparency: Gated (signup required)
Pricing visible only after signup or form fill.
§Pricing transparency: Contact sales
No public numbers — sales conversation required.
§Time to value: Self-serve
Sign up and produce value the same day.
§Time to value: Weeks to deploy
Light implementation — integrations, data mapping, a pilot.
§Time to value: Months (implementation)
Real implementation project: data work, change management, SOWs.
§Funding stage
Most recent disclosed round (Seed, Series A–E, PE-backed, public, bootstrapped). Signal for runway and roadmap risk.
§Team size band
Approximate headcount band from LinkedIn or vendor disclosure. Bands avoid false precision.

#Source types

§Vendor site
Vendor's marketing site. Primary for product claims; treat positioning critically.
§Privacy policy
Vendor's published privacy policy or DPA. Authoritative for data handling.
§Vendor docs
Vendor's developer documentation. Authoritative for APIs, models, and limits.
§G2
G2 reviews and category pages. User-reported feature presence and sentiment.
§Capterra
Capterra reviews. Similar to G2; weight against vendor self-claims.
§TrustRadius
TrustRadius — typically deeper, enterprise-skewed reviews.
§News / press
Press articles, funding announcements, product launches.
§Reddit
Reddit threads — useful signal but unverified; flagged as community.
§Analyst (Gartner/Forrester)
Gartner, Forrester, IDC reports and Magic Quadrant placement.
§Community
Other community sources — HN, Slack, Discord, blogs.
§Other
Anything that doesn't fit the above; URL always provided.
§Confidence (high / medium / low)
high: primary source, unambiguous. medium: secondary or partial source. low: inferred or community signal — treat as a lead, not a fact.

#Marketing fit

§Channels
Where the tool operates — email, SMS, paid social, paid search, web, in-product, CRM, sales outbound, organic content.
§Use cases
The marketing jobs the tool is built for: lifecycle, lead scoring, content generation, ad creative, attribution, SEO, RevOps, etc.
§Integrations
Native connections to the surrounding stack (warehouses, CDPs, CRMs, ad platforms, analytics). Native > Zapier > CSV.
§Output types
What the tool produces — copy, images, video, audio, structured data, decisions/actions, or analyses.
§Notable customers
Publicly disclosed logos (case studies or vendor site). We do not list customers from sales decks or rumor.
§Headquarters
Primary legal/operational HQ. Useful as a first-pass proxy for data residency posture; always confirm in the DPA.
See a mistake?

Definitions and vendor facts evolve. Send a correction with a source and we'll update.