Safety
Everything below describes the safety behaviour the code implements today.
What is not allowed
Harassment and bullying, hate and dehumanising speech, threats and incitement to violence, self-harm promotion, sexual content involving minors, spam and repetition, fraud, scams and phishing, and malicious or deceptive links.
How moderation works
Every piece of public text — messages, profile name, bio, links, and all entry fields and links — is checked before it is stored. Kevin applies local structural rules (Unicode abuse, repetition, link count, https-only URLs, punycode and phishing patterns) and then OpenAI’s Moderation endpoint. Narrow structural rules also block handing out private contact details (email addresses, phone numbers, messenger handles), which is a common first step in scams and in approaches to younger people. Notes attached to a report go through the same review. Moderation is fail closed: if the check cannot complete, the write is refused. Rejected text is never stored anywhere.
Controls you have
Room messages disappear by themselves after 30 minutes. You can delete your own messages, delete any message inside your own profile room, clear your whole profile room, block another profile (which removes follows and likes in both directions and hides you from each other), and delete your entire account. Blocking is a separate control from reporting: it acts immediately and only affects you and the profile you block.
Reporting
You can report a room message, a profile or a published entry from inside ChatGPT. A report records what was reported, the category and an optional short note. Reports are private: they are never shown publicly, never appear in News, Discover or any public metric, and no IP address or device information is collected. Reports are rate-limited and de-duplicated, and kept for at most 180 days.
Reports touching minor safety, sexual content or self-harm hide the reported room message from public view immediately, pending review. Hiding is reversible and a single report never deletes content or an account automatically. Kevin does not tell reporters what enforcement followed.
Limitations
Kevin is a public space: assume everything you write can be read by anyone. Moderation is automated and imperfect, and Kevin makes no guarantee that all rule-breaking content is caught or removed. There is no large moderation team; reports are reviewed operationally. Kevin is not an emergency service and cannot help in a crisis — contact your local emergency number. Because messages are physically deleted after 30 minutes, Kevin usually has no message content left to produce for a legal or law-enforcement request.