AI Memory is what the AI knows about a client. A large part of it is built automatically from the crawl of the client's website. This article explains the pipeline, which fields are filled in for you, and which parts always need a person.
The pipeline for each crawled page
Fetch the page. See How the app crawls and categorizes pages.
Clean it: menus, footers and boilerplate are removed so only the readable content remains.
Classify and extract. The page type is confirmed from the content, and structured data for that page type is extracted.
Extract memories: separate facts from the page, such as contact details, locations, products, services and team members.
Save the memories. Memories that already exist are skipped, so a re-crawl doesn't create doubles.
Index the page for search. The content is split into overlapping chunks (about 1,000 characters, overlapping by 200) and added to a hybrid search index that combines meaning-based search with keyword search.
Pages whose type never becomes memories (Legal, Thank You, Login, Local SEO and Article / Blog) skip steps 4 and 5, but they are still indexed for search.
How the AI finds it back
When someone uses AI Chat or another AI tool, the platform searches the index for the parts of the client's own website that match the request and uses them as source material, next to the company profile and the instructions in AI Memory.
What is filled in automatically
Company Branding
Branding is read from the homepage when the website is registered, and topped up again when the crawl finishes. It fills blank fields only: brand name, tagline, brand description, logo, favicon, colors, fonts and images. See Setting up Company Branding.
The company profile
When at least one page was crawled, the profile on the Memory tab is filled in from the crawled content. Only empty fields are filled, so nothing someone typed is overwritten.
Field | Where it comes from |
|---|---|
Content & About | The most relevant crawled pages |
What we do | The most relevant crawled pages |
Mission | Only when the site states it explicitly |
Company name, Phone, General email, Address | Extracted contact and location details, then the header and footer of the homepage and contact page |
Tax number | The header and footer of the homepage and contact page |
Web Services | The company's website and social links. Links already on the card are kept; only missing ones are added. |
Contacts | Extracted contact details |
Keywords | Up to 8 search phrases a customer would use |
The profile is written from up to 8 of the most relevant crawled pages, with the homepage first, followed by about, contact and location pages. Someone can also run Pre-fill from website at any time to review suggested changes, including improvements to fields that are already filled in. See Completing the company profile.
Products and Services
The Products and Services cards list the pages categorized as Product or Service on the Sitemap tab. If a card is empty, it says so ("No products. Categorize pages as Product in the Sitemap."). Correct the page types and the cards follow.
What stays manual
Some things can't be read reliably from website text, so they are left for your team:
target audience details (audience, gender, age, generation)
language, country and region
sector
timing and events
Exclusions (topics, competitors and content the AI should avoid)
the house rules in How we write and the Content Examples
Tip: A crawl gives the AI the facts. The manual cards and How we write make the difference between generic and on-brand output. Plan time to complete them after every new crawl. See How we write: house rules for the AI.
The crawl-complete email
When a crawl started from Accounts finishes with at least one page crawled, the person who clicked Crawl receives an email. It lists how many pages were crawled (and how many couldn't be reached), how many new memories were extracted and which fields were auto-filled.
The approval step
A finished crawl moves the company's memory status to To verify. An admin reviews the page list on the Sitemap tab and clicks Approve, which sets the status to Ready.
Approval is refused when no page was categorized as Product or Service, because that usually means only part of the site was found. The warning offers Look for more sitemaps, and an option to approve anyway for organizations that sell neither. When a later sitemap refresh finds new pages, the status goes back to To verify. See Approving a website crawl.