Whitepaper: Can AI Improve Learning? Evidence from a Randomized Study of Kyron’s AI-Powered Instructional Experience
As AI becomes increasingly common in higher education, institutions face an important question: does AI actually improve how students learn? While much of the current conversation focuses on model capabilities, efficiency, or automation, relatively little evidence addresses whether AI-powered experiences lead to better student learning. We designed this study to address that gap directly.
The Question
Can Kyron’s AI-powered, interactive instructional experience produce better learning outcomes compared to a text-based instructional material?
The Study
To answer this question, we designed a pre/post, A/B study (see figure 1). Using Prolific’s participant pool and randomization capabilities, we recruited college students in the United States who were pursuing a bachelor’s, technical, or vocational degree whose highest completed level of education was a high school diploma or GED. A total of 201 participants1 were recruited for the pretest and 158 participants completed all parts of the study.
Participants first completed a 20-question pretest about memory concepts then were randomly assigned to complete either an interactive AI lesson or a traditional text-based reading covering the same material, with both conditions requiring comparable instructional time. Finally, students completed a 20-question posttest covering the same content as the pretest.

Figure 1
We chose memory concepts (from introductory psychology) as the study topic because it was likely to be somewhat familiar to most participants, given their experience as learners, while still leaving meaningful room for improvement. The content was drawn from a chapter of a widely used introductory psychology textbook, making it representative of the type of instructional material students commonly encounter in the early years of higher education. The topic was also accessible enough to be broadly relevant regardless of participants' academic focus and could reasonably be learned within a single 20–30 minute session.
The Intervention
Kyron Learning provided the AI-powered instructional experience evaluated in this study. Kyron is an AI-powered instructional platform that combines brief instructional video followed by a guided, interactive dialogue designed to help learners actively construct understanding rather than passively consume information (see table 1).
For this study, participants in the text-based condition read a portion of a chapter directly from the introductory psychology textbook covering core concepts of memory. To ensure that both instructional experiences were delivering the same material, the participants in the Kyron condition took a lesson that was designed using a targeted content alignment process. This process intentionally captured content, instructional methods, and academic vocabulary from the textbook and incorporated those elements into the lesson allowing us to isolate the effect of the AI instructional approach while holding topic and content constant.
|
Example Learner-Kyron AI Interaction The excerpt below shows a real interaction between a learner and Kyron’s AI during the study. Rather than simply evaluating responses as correct or incorrect, Kyron’s AI uses targeted follow-up questions and immediate feedback to help learners develop more complete explanations. |
||
Question: Working memory can only hold a limited amount of information at once. How does grouping information into chunks help you work within this limitation? |
||
|
Learner |
Grouping information into chunks makes it easier to remember. |
|
|
Kyron AI |
That’s a solid beginning — it does make memory easier. What happens to the number of separate items working memory has to handle? |
AI acknowledges what the learner knows before prompting deeper explanation |
|
Learner |
It becomes lower. |
|
|
Kyron AI |
You’re on the right track. It lowers the number of separate items in working memory. What does that grouping do to those individual pieces—how are they treated? |
AI uses a follow-up question to deepen the learner's understanding |
|
Learner |
We treat them as the same type of piece, meaning that we think there is less information to remember overall. |
|
|
Kyron AI |
Nice — you’ve captured the idea that several pieces can function as one chunk. One more part: does chunking change the actual capacity of working memory, or does it help you use it more efficiently? |
AI checks understanding of an important concept before moving on |
|
Learner |
It makes it more efficient. |
|
|
Kyron AI |
That’s correct! You’ve explained how chunking makes information more manageable without increasing working memory capacity. |
AI confirms understanding and reinforces the key takeaway |
|
Table 1 |
||
The Findings
After accounting for learners' starting knowledge, participants assigned to the Kyron condition scored higher on the posttest than participants assigned to the reading condition.
Participants in both groups improved from the pretest to the posttest. To determine whether one instructional approach led to better learning outcomes than the other, we compared posttest scores while accounting for participants' starting knowledge (pretest performance). After accounting for pretest scores, participants in the Kyron condition scored statistically significantly higher on the posttest than those in the reading condition (F(1, 155) = 9.61, p = .002)1 (see table 2). The estimated average posttest score (out of 20) was 18.21 for the Kyron group compared with 17.37 for the reading group—a difference of 0.84 points in favor of the Kyron condition (i.e., approximately 5% higher than the reading condition) (see table 3). This difference represented a moderate effect size (partial η² = .058).

Table 2

Table 3
What These Results Mean
These findings contribute to a broader question facing higher education: how can AI be used to support student learning?
One possible explanation for these findings is the nature of the learning experience itself. Rather than simply reading the material, learners using Kyron actively engaged with concepts through guided dialogue, received immediate feedback on their thinking, and were supported with scaffolded instruction as they worked through misconceptions. These instructional supports were designed to encourage learners to actively process and apply new ideas rather than passively consume information. The higher posttest performance observed in the Kyron group is consistent with that instructional approach.
These characteristics of Kyron’s AI learning experience are grounded in decades of research from cognitive science and the learning sciences, long before the emergence of modern generative AI. AI makes it possible to deliver these evidence-based instructional practices in a personalized, interactive, and scalable way, but the underlying principles of effective learning are not new.
Survey Responses and Open-Ended Feedback
Although this study focused on objective learning outcomes, learners also reported positive perceptions of the instructional experience through both survey responses and open-ended feedback.
84% enjoyed the lesson.
91% said it kept them thinking rather than just watching or clicking.
89% felt more confident after the lesson.
Many learners specifically highlighted the value of the AI-guided dialogue, opportunities to think through concepts, and interactive format.
"It felt like having an instructor to discuss and clarify concepts with."
"The back-and-forth conversation helped me really think critically."
"I wish my university taught concepts that way."
For higher education, the question may be less about whether to use AI and more about how to design AI-powered experiences that meaningfully support learning.
Limitations and Next Steps
This study was designed to evaluate learning within a single instructional session under controlled conditions. As such, it does not capture the broader context of classroom instruction, repeated use of AI-powered learning experiences across a semester, or long-term retention of knowledge. Future studies should examine how these findings generalize to authentic classroom settings, sustained use over time, and delayed measures of learning. In addition, although the instructional principles evaluated in this study (e.g., guided questioning, immediate feedback, scaffolded instruction) are not unique to psychology, future research should explore how these findings generalize across subject areas.
This study also evaluated the overall instructional experience rather than its individual design features. Understanding which aspects of AI-supported instruction most strongly influence learning remains an important direction for future research. As higher education continues to explore the role of AI in teaching and learning, rigorous evaluations such as this can help ensure that decisions are guided by evidence rather than enthusiasm alone.
Footnote
¹ Before collecting data, we conducted an a priori power analysis to determine the sample size needed to detect a moderate effect (Cohen's d = 0.5) using an independent-samples t-test, which yielded a target sample size of 128. We selected this approach as a conservative benchmark: our planned analysis, an ANCOVA controlling for pretest scores, would likely require a smaller sample to achieve the same power, but we did not want to rely on that assumption when setting our recruitment target. We ultimately set a recruitment target of 200 college students to account for anticipated attrition, and obtained complete data from 158 participants. An analysis of covariance (ANCOVA) was used because it accounts for participants' starting knowledge and compares posttest performance among learners with similar baseline knowledge (i.e., similar pretest scores). Because pretest performance was a strong predictor of posttest performance in this study, ANCOVA provided a more precise estimate of the difference between instructional conditions than a comparison of posttest scores alone. We also confirmed that the assumption of homogeneous regression slopes was met, as the interaction between condition and pretest score was not significant, F(1, 154) = 0.203, p = .653, supporting the use of the standard ANCOVA model.