Until September 24, 2026, Folio Surveyor could draft the Methods and Results sections of a survey study. A model read the study's numbers and wrote paragraphs, and one click put them into the manuscript. That feature is gone. In its place, Surveyor gives researchers structured notes: the facts to report, each one computed from the study's own data, and a short list of what the researcher still has to supply.
This post explains the decision and shows what replaced it. The reasoning holds whether or not you ever use Folio, and the last part walks through turning a block of notes into a Methods paragraph by hand.
A Methods section is a signed statement
Most of a paper argues. The Methods section testifies. It says who took part and how they were found, what they were asked, how the answers were scored, what approval the study had, and which responses were set aside. Reviewers check the claims in the Discussion against it, and editors look there for the ethics approval and consent details a journal requires. When a published paper is questioned, the Methods is often where the questions start.
That makes it the part of a paper where authorship carries the most weight. The ICMJE recommendations, which many medical journals follow, say AI tools should not be listed as authors because they cannot take responsibility for the accuracy and integrity of the work, and that the humans who submit a paper are responsible for anything a tool helped produce. The same recommendations ask journals to require authors to disclose AI use at submission. A Methods paragraph is a list of exactly those responsibilities, one sentence at a time.
Why generated prose does the most damage here
A language model writes the most plausible next sentence. In a Discussion, a plausible sentence that is wrong is a weak argument, and a careful reader can push back on it. In a Methods section, a plausible sentence that is wrong is a false statement of fact about a study the model never observed.
Methods sentences are also the easiest kind for a model to produce. "Informed consent was obtained from all participants." "Internal consistency was good." Sentences like these recur almost word for word across published papers. A model can write them fluently for any study at all, and that is the problem: fluency is no evidence that the sentence is true of this one.
Folio found this in its own output. A review of the old drafts on a test study, built from two validated scales (the PSS-10 and the GAD-7) plus a handful of questions such as age and gender, turned up a Measures paragraph about twenty-two custom Likert items "developed to assess additional dimensions of stress, coping, and emotional regulation." The survey had no custom items. Part of that error was Folio's: a bug handed the scale items to the model as if they were custom questions. The rest was the model's. Nobody had told it how those items came to exist, so it supplied the most plausible story, down to how they were "selected and grouped thematically." The same draft said demographic characteristics were not systematically recorded, in a survey that asked participants their age and gender. Neither sentence looks wrong to a reader who does not have the data open, and both would have gone into a manuscript under the researcher's name.
That is why the fix was not a better prompt. Surveyor stopped writing the sentences and started handing over the facts.
What Surveyor gives instead
The Methods, Participants, Measures and Results deliverables in Surveyor are now notes. Each note is a label and a value, grouped by topic. Every value comes from the study's own records (the instrument as it was fielded, the responses and the study details the researcher entered) or from the scale records in Folio's library, which supply each validated scale's citation, scoring rule and published reliability. No model is involved. The numbers are computed in code, so the notes are available on every plan and use no AI allowance. They can be copied one at a time, by section, or all at once, and they are never inserted into a document.
The Methods notes cover:
- Design and procedure. The project type, the research question, how the survey was administered, each distribution with its count of complete responses, the collection window, the dates responses came in, the median completion time and the number of questions.
- Sample. How many people started and finished (with the completion rate), how many complete responses were analyzed, and how many were held out as flagged or excluded. Recruitment and compensation appear when the researcher has entered them.
- Participant characteristics. Every choice or number question with its answers, as counts and percentages or as mean, SD, range and n, so the researcher can keep the ones that describe participants.
- Each validated scale. The citation from Folio's scale library, the number of items fielded, the response format and anchors, the scoring rule applied, how many items were reverse-scored, any subscales, the published reliability from the scale's record, and this sample's Cronbach's alpha with its n and a 95% confidence interval.
- Ethics and data handling. The approving committee, protocol number and approval date when entered, whether the survey had a consent question and whether its text mentions the right to withdraw, and what was stored with each response.
The Results notes give each published scale score with its n, mean, SD, observed range and reliability, plus interpretation bands where the scale defines them, with the source of the cut points when Folio has one on record. Then come the statistics for the other questions, and up to two tables that paste into a word processor or a spreadsheet.
Both sets end with a list headed "To supply yourself". It names what the data cannot tell Folio: how the sample size was decided, how participants were recruited and compensated if that was never entered, missing ethics approval details, and, in the Results notes, the tests of the study's hypotheses. The notes are descriptive on purpose. The tests are the ones the researcher planned, and Surveyor's Comparisons and Regression views compute them when asked.
Why the two alphas sit on separate lines
One detail in the notes guards against a common mistake in survey write-ups. Every validated scale comes with a reliability coefficient from its validation sample, and it is tempting to report that number as if it described your own data. It does not.
The APA Task Force on Statistical Inference put it directly in 1999: "a test is not reliable or unreliable." Reliability belongs to the scores from a particular sample, so authors should report the coefficient for the data they actually analyzed. APA's current reporting standards for quantitative research (JARS-Quant) ask for the same thing, and they add that when a paper reports coefficients from other samples, such as a test manual's, it should describe those samples' basic demographics too. Citing the manual's alpha in place of your own has a name in the measurement literature, reliability induction, and studies of the practice have compared the samples that borrow a coefficient with the samples it came from.
So the notes print both, labeled: "Published reliability (scale record)" and "Reliability in this sample". The sample line carries its n and a 95% interval computed with Feldt's method, because a coefficient without its precision invites over-reading. The interval depends heavily on sample size. For a 10-item scale with α = .86, it runs from .83 to .89 with 212 respondents, and from .76 to .93 with 24. Below 30 respondents, the notes add a reporting note: give the interval, and leave out labels such as "good" or "excellent" that the interval cannot support.
From notes to a Methods paragraph
Here is what that looks like for an illustrative study: an online survey of graduate students using the 10-item Perceived Stress Scale. The study and its numbers are made up for this example. The layout and labels are what Surveyor produces. An excerpt from the Methods notes:
Sample
- Started: 248
- Finished (completion rate): 219 (88%)
- Complete responses analyzed: 212
- Held out of the analysis: 7 flagged, 0 excluded
- Recruitment: Email list (Graduate program mailing lists)
Perceived Stress Scale (PSS-10)
- Citation: Cohen, S., & Williamson, G. M. (1988). Perceived stress in a
probability sample of the United States. In S. Spacapan & S. Oskamp
(Eds.), The social psychology of health (pp. 31-67). Sage.
- Items in this survey: 10
- Response format: 5-point frequency scale (Never / Almost never /
Sometimes / Fairly often / Very often)
- Scoring: PSS-10 total, 0-40, Cohen 1988 scoring
- Items reverse-scored: 4
- Published reliability (scale record): α = .78
- Reliability in this sample (Total): α = .86, 95% CI [.83, .89], n = 212
To supply yourself
- Whether participants were compensated, and how.
- How the sample size was decided, for example a power analysis.
And a paragraph the researcher might write from it, once the two open items have answers:
Graduate students were invited through their programs' mailing lists and were not compensated. A target of at least 200 complete responses was set before data collection, based on a power analysis for the study's main comparison. Of the 248 people who began the survey, 219 finished it (88%). Seven finished responses were set aside because they failed the survey's attention check, leaving 212 for analysis. Perceived stress was measured with the 10-item Perceived Stress Scale (PSS-10; Cohen & Williamson, 1988), answered on a 5-point scale from 0 (never) to 4 (very often). The four positively worded items were reverse-scored and the items summed to a total from 0 to 40. Cohen and Williamson reported α = .78 in a U.S. national probability sample; in the present sample, α = .86, 95% CI [.83, .89].
The counts, the scoring range and both reliability figures came straight from the notes. The order and emphasis came from the researcher, and so did the sample-size target, the compensation and the reason those seven responses were set aside. The researcher is the only person in a position to know them.
A few habits from this carry over to any survey write-up, with or without Folio:
- Report your own sample's reliability, not only the published figure, and give its interval when the sample is small.
- Before drafting, list what your data cannot tell a reader (recruitment, compensation, how the sample size was chosen) and answer each item.
- Report participant flow as numbers: started, finished, set aside and why, analyzed. JARS-Quant asks for the flow of participants through each stage of the study.
- Name the scoring rule, including which items were reverse-scored, so a reader can reproduce the total. The PSS-10 scale page and the BFI-44 reporting guide show what that detail looks like for two common instruments.
What Surveyor still assembles
Not every deliverable became notes. The instrument appendix, the scale citations and the short ethics and data availability statements are still assembled from templates in code, using only what the study records, and they can be inserted into a connected document. Where a fact is not on record, such as an approval number, the statement leaves a bracketed placeholder for the researcher instead of a guess.
"Start the write-up" in Results has changed to match. It creates a document with an empty Method and Results outline and a table of the survey's descriptive statistics, and it opens the notes beside the document in the editor. The document holds no generated sentences. The route that creates it rejects any draft with text outside headings and table cells, so no client can slip prose in through it. The Surveyor help article covers the rest of the workflow.
The line this draws
Folio's rule for AI in research writing has been the same across the product: assist, never ghostwrite. The Methods notes apply that rule where it matters most. Software does what it is reliable at, which is counting, scoring and computing intervals from data it can see. The researcher does the part that carries their name, which is saying what the study was.
Sources: Appelbaum, M., Cooper, H., Kline, R. B., Mayo-Wilson, E., Nezu, A. M., & Rao, S. M. (2018). Journal article reporting standards for quantitative research in psychology: The APA Publications and Communications Board task force report. American Psychologist, 73(1), 3-25. · American Psychological Association. (2024). JARS-Quant Table 1: Information recommended for inclusion in manuscripts that report new data collections. · Feldt, L. S. (1965). The approximate sampling distribution of Kuder-Richardson reliability coefficient twenty. Psychometrika, 30(3), 357-370. · International Committee of Medical Journal Editors. Use of AI by authors. Recommendations for the Conduct, Reporting, Editing, and Publication of Scholarly Work in Medical Journals. · Vacha-Haase, T., Kogan, L. R., & Thompson, B. (2000). Sample compositions and variabilities in published studies versus those in test manuals: Validity of score reliability inductions. Educational and Psychological Measurement, 60(4), 509-522. · Wilkinson, L., & the Task Force on Statistical Inference. (1999). Statistical methods in psychology journals: Guidelines and explanations. American Psychologist, 54(8), 594-604.
The worked example is illustrative. Its study and numbers are invented; the intervals were computed with the same code Surveyor uses.