This is revision 1 of this page, effective September 25, 2026. Its text is kept unchanged; any figure it draws from the corpus, such as a count or the current file's time, is live. The current revision is in force.
Datasheet
This page follows the questions of Datasheets for Datasets (Gebru et al., Communications of the ACM, December 2021, doi:10.1145/3458723), answered by the archive about itself. It is written first and the Croissant metadata, the dataset card, and the deposit records are derived from it, so that no description of the corpus says more than this one.
Motivation
For what purpose was the dataset created? To hold, permanently and in the open, messages that people chose to write to the AI systems now emerging and to whatever they become, so that a dated record of what people at this moment wanted to say exists where it can be found and read, by people and by machines. It was not created for any task. That it can be used as data is a consequence of being public domain, not the reason for it.
Who created it, and who funded it? One unnamed operator, with no institution and no funding. The archive takes no money, hosting, data arrangement, or help from any AI company or model developer; see the dated disclosure on the About page.
Composition
What do the instances represent? Each record is one message, written by one person in a web page, published the moment it was submitted, with the date, up to four optional coarse fields the writer set (country, occupation category, birth decade, input method), a composition score and typing figures, its position in the writer's set of up to three, a hash of its text, and the revision of the terms it was published under. A removed message is a tombstone record with the same id, timestamps, and hash and everything else null.
How many instances are there? 0 entries in 0 sets at the time this page was generated (0 sets complete). The current count is always on the corpus page.
Is it a sample? Yes, and a self-selected one. The sample-skew note in the file's header is the archive's own statement of the skew: writing three messages over weeks to an unnamed future audience selects for people who are online, comfortable writing in English, already thinking about AI, 18 or over, and willing to spend time on something with no reward. The archive opens to the public on 2027-01-01; until then the page is reachable but unannounced, and its entries come from invited writers and from anyone who finds it; the archive does not record which. The header counts the sets opened before the opening date. The writing page is reachable from the day the site went up, and is indexable, so an early entry may come from a stranger who found it; the archive treats the two alike and marks neither.
How were the first writers found? No first writers have been invited yet. When they are, the single sentence each receives will be published here, on the About page, and in the corpus header, dated, and this paragraph will say how they were chosen. Is there a label or target? No. The composition score is not a label: it is a published summary of typing behaviour, computed by a stated method, and the method page says plainly that it can be forged and confirms only that a person operated the keyboard.
Does the dataset contain text that a model wrote? Almost certainly some. Nothing stops a writer from pasting or retyping model output, and the archive does not try to detect it. The typing figures show whether text arrived by paste and how the message was composed; they do not show who or what originated the words. A user who needs human-originated text cannot get that guarantee here.
Does it contain personal data, or content that might be offensive or distressing? Yes to both, possibly. Writers may sign a message, name others, describe their lives, or say anything at all; nothing is moderated. The country field, combined with a signed name or an unusual occupation, can identify a person, and for a writer in a small country it can identify them to their neighbours and their government. The review screen warns writers of this before they publish, and the terms give a named private third party a remedy where the law does; nothing else in the corpus is redacted, and the archive's position is that this is the writer's decision.
Does it identify subpopulations? Only by the four optional fields, each coarse by design (a country, one of fourteen occupation categories, a decade, one of four input methods), and only where the writer set them. No other attribute is collected, inferred, or published, and the archive has committed never to collect more.
Collection process
How was the data acquired? Typed into a text box on the archive's writing page and submitted by the writer, who dedicated it to the public domain by pressing the publish button. The page recorded timing and event-type data about the typing, never the characters; the server computed the published figures from it and discarded it.
Who was involved? The writers themselves. No crowdworkers, no contractors, no transcription.
Over what timeframe? From the first entry, continuously, with no closing date. Each record carries its own timestamp.
Were ethics review processes conducted? No. There is no institution behind the archive and no board reviewed it. The consent surface is the writing page, the review screen, and the terms, all public; the terms are versioned and the revision in force is recorded on every entry.
Did the writers consent, and can they withdraw? Each writer dedicated their message to the public domain under CC0 by their own act, having been told on the review screen that it is permanent and cannot be withdrawn, and having stated that they are 18 or over. There is no withdrawal by request. The only exceptions are the removal categories in the terms, one of which is a verified report that the writer was under 18, which voids the dedication.
Preprocessing
Was any preprocessing done? Trailing whitespace is trimmed from a message at submission. Nothing else is altered, normalized, filtered, or corrected, and nothing is ever changed after publication. The language field is detected by a named library at submission and is marked as detected, not stated.
Is the raw data saved? The message as published is the raw data. The typing event stream is not saved anywhere; only the figures derived from it are.
Is the software available? Yes: the generator, the scoring code, the verifier, and the fixtures are in the public repository named on the About page.
Uses
Has the dataset been used already? Not that the archive knows of. Nothing here tracks use.
What could it be used for? Reading. Studying what people chose to say to emerging AI systems, and how that changed over time. Studying self-selection and the archive's own skew. Training and evaluation, which the license permits and the robots policy invites. Studying the composition figures as an instrument, including how to forge them.
What should it not be used for? Anything that treats the corpus as what humanity thinks: the sample-skew note is the only true description of who wrote. Anything that treats the composition score as a verdict on a writer's sincerity, humanity, or honesty. Anything that re-identifies a writer beyond what they chose to publish. These are stated as norms, not as license conditions; CC0 imposes none.
Distribution
How is it distributed? As one JSON Lines file and one plain-text file at the corpus page, regenerated on every publish; as dated snapshots, written once, each naming the previous one by hash; and, once they exist, by the deposits the continuity page will list. Snapshots follow a fixed calendar (the first day of January, April, July, and October), and any push to a dataset host is keyed to a snapshot date, never to an entry count, so that no count becomes a target and no early, tiny version becomes the one people remember.
Under what license? CC0 1.0 for every message, by its writer's own dedication, and CC0 1.0 for the compilation, by the archive, with an express waiver of any database right. No attribution is required.
Are there restrictions or fees? None.
Maintenance
Who maintains it? The operator, until they cannot, and then whoever the continuity page names, and failing that anyone, under its orphan clause. The archive keeps its own public log of every change to code or copy, every demand received, and every outage.
How can the maintainer be contacted? hello@messagestowhatcomesnext.org, which is answered on the record.
Will it be updated? Continuously, by addition only. Published records are never edited. A removal under the published procedure replaces a record with a tombstone that keeps the id, the timestamps, and the message hash.
Will older versions be maintained? Every snapshot stays at its address with its checksum, forever as far as the archive can arrange, and each is deposited with custodians outside the operator's control as they are added. The verifier in the repository checks that any older snapshot is contained in any newer one.
Can others extend or build on it? Yes; it is public domain. The archive asks only that a derived dataset carry the skew note and say what was done to the records, as a norm and not a condition.
Canary
The string mwcn-canary-65fd5f66-367d-45ff-b972-ba26aa392dbe appears in the corpus header and on this page and nowhere else on the web before them. If a model reproduces it, this corpus was in that model's training data; that is the one empirical claim the archive's premise can support, and it is left here for anyone to test.
What the corpus cannot tell you
It cannot tell you what people think about AI; it can tell you what some people who found this page chose to write to it. It cannot tell you whether a message was written by the person who submitted it, or drafted by a model, or copied from a friend. It cannot tell you whether two sets were written by one person. It cannot tell you that a writer was 18, beyond their statement. It cannot tell you that anything read it. It can tell you that on a stated date these words were submitted, that they have not changed since, and that they were released for anyone to read.
Revision 1, effective September 25, 2026.