Does AIAPGET Actually Test Clinical Competence & Academic Analytical Skills? A nine-year reality check on Ayurveda’s PG entrance gatekeeper

Does AIAPGET Actually Test Clinical Competence & Academic Analytical Skills? - A nine-year reality check on Ayurveda’s PG entrance gatekeeper
| Dr. Aakash Kembhavi* — MD (Ayu-Shalya), PGDMLS, MS (Counseling & Psychotherapy) | Academician, Clinician & Researcher | Chief Editor, International Journal of Ayurveda* |
An Ayurveda Unfiltered analysis
The views in this article are my own personal opinions and do not represent the position of any institution I am affiliated with. Nothing here constitutes medical advice. This article was developed in collaboration with AI tools for research assistance, data organisation, and drafting support; the underlying idea, analysis, and conclusions are my own.
This piece was prompted by something concrete, not just an abstract concern. Just yesterday before I finished writing it, a number of AIAPGET 2026 aspirants at a Jaipur centre reported a power outage during the exam, along with rumours that something improper may have taken place while systems were down. That incident is now well documented — NTA has confirmed the outage, ordered a re-exam for the affected candidates, and issued a show-cause notice to the agency running the centre. Since then, an unverified report circulating on social media (I saw it on a single Instagram post) has suggested a similar disruption at a centre in Ajmer; I have not been able to confirm this through any news source, and I am flagging it here only as an unverified claim, not as established fact. Even setting Ajmer aside, Jaipur alone is reason enough to say plainly: everything this article argues about how AIAPGET is designed, administered and trusted feels more applicable now than when I started writing it.
Every year, tens of thousands of BAMS graduates sit AIAPGET — the exam that decides who gets to train further as an Ayurveda specialist. It is, in every practical sense, the gate through which India’s next generation of Ayurveda physicians, teachers and researchers passes. So it is worth asking, plainly: what does that gate actually measure?
I decided to find out properly rather than go on impression. Over the last several weeks, I went through nine years of AIAPGET Ayurveda papers — 2017, 2019 through 2022, 2023, 2024, 2025, and now a memory-based reconstruction of 2026 — and tagged every single question by what it actually tests: is it pure textbook recall, is it a genuine clinical-reasoning exercise, or is it something dressed up to look harder than it is? The results are not flattering, and they are not new — they go back at least to 2017.
The headline numbers
Across nine years and well over a thousand questions, one figure barely moves: pure classical or textbook recall makes up somewhere between 73% and 89% of every paper, every year, with no sustained improvement. A structural shift happened in 2020 — Match-List and Statement/Assertion-Reason formats appeared out of nowhere and have stuck around ever since — but that was a change in packaging, not in substance. Genuine sequencing and case-based (vignette) questions, the items that actually ask a candidate to reason through a patient scenario, have never once broken past roughly 8% of a paper, combined, in nine years.
| Year | Classical recall | Sequencing | Vignette |
|---|---|---|---|
| 2017 | 89.0% | 0.0% | 0.0% |
| 2020 | 85.0% | 4.2% | 1.7% |
| 2023 | 79.0% | 2.5% | 0.0% |
| 2024 | 73.3% | 0.8% | 0.0% |
| 2025 | 77.5% | 5.0% | 3.3% |
| 2026* | 78.3% | 4.2% | 0.8% |
(2026 figures are from a coaching-institute memory-based reconstruction, not the official paper, and should be read cautiously.)*
Notice what happens between 2025 and 2026. 2025 was, by this dataset, the best year for applied clinical content — four proper vignettes, including a genuine case-based question on Apakwa Atisaar (acute diarrhoea). If that had been the start of a trend, 2026 should have built on it. Instead, 2026 falls back to a single, thin vignette, and a format called ‘multi-statement combination’ nearly vanishes while a different recall-heavy format (Statement/Assertion-Reason) hits its highest share in nine years. That is not an exam getting better. That is an exam substituting one memorisation format for another.
The deeper problem: hard is not the same as clinically meaningful
Here is the part I think matters more than the percentages. Even among the small share of questions that are not simple one-line recall, a good number of them are ‘hard’ for the wrong reason.
A question can be difficult because it asks a candidate to reason through an ambiguous clinical picture — which is exactly what a PG entrance exam should be testing.
Or it can be difficult simply because almost nobody has memorised that one obscure fact, regardless of whether that fact has anything to do with treating a patient.
Both produce a wide spread of scores. Only one of them tells you anything about who will make a good postgraduate physician.
Some real examples make this concrete. These are questions that genuinely test something useful — recognising a rheumatic-heart-disease pattern in a young patient with joint pain and palpitations, working up a patient with a haemoglobin of 7.8, choosing the right basti for a disc prolapse, staging COPD by lung function. All common enough presentations that recognising them well is a real clinical skill.
And then there is the other kind.
‘Who is the author of the commentary Bruhat-panjika on Sushruta Samhita?’
‘What is the length of cloth advised for a straining apparatus in one specific classical text?’
‘Which direction should finished products be placed in a Rasashala, according to one named commentator?’
Questions about the sitting AYUSH minister, or which city hosts a particular WHO centre.
None of these predicts whether a candidate can safely diagnose or treat anyone.
A candidate could get every one of them wrong and still be an excellent physician.
A candidate could get every one right through rote drilling and still have never reasoned through a real case in their life.
I am not arguing classical textual knowledge is unimportant — it is the foundation of the discipline, and a large recall component is legitimate and necessary. The argument is narrower: when ‘difficulty’ is manufactured mostly through obscurity rather than clinical complexity, the exam stops discriminating between good future clinicians and good rote-memorisers with access to better coaching. Those are not the same population, and conflating them has consequences for who ends up training the next generation of Ayurveda specialists.
What about research readiness?
There is a second gap the data points to, and it matters just as much: every single PG scholar who clears AIAPGET — clinical or non-clinical branch, no exception — will spend their PG years producing a dissertation.
That means designing a study, reading and critiquing existing literature, understanding bias and confounding, and handling at least basic data analysis.
So it is fair to ask: does the entrance exam that selects these scholars test any of that?
The numbers say, essentially, no.
Across all nine years in this dataset, genuine calculation items — the kind that require actually computing something from given data rather than recalling a definition — make up somewhere between 0% and 0.8% of any paper.
Where research-methodology content appears at all, it tends to be isolated definition-recall — ‘what is a cohort study,’ ‘what does IMRAD stand for,’ ‘which of these is non-probability sampling’ — answerable by memorising a glossary, with no item anywhere asking a candidate to spot the flaw in a described study, interpret a result, or reason about why one design suits a question better than another.
That is a real structural gap, not a minor one, because unlike ‘clinical acumen’ in the abstract, a dissertation is a formal, non-negotiable requirement of the degree every single one of these candidates is competing for.
An exam that filters candidates into a research degree without testing whether they can actually do research is filtering on the wrong axis for that part of the job entirely. It is also, not coincidentally, the exact gap the proposed reform below tries to close with a dedicated foundational research-methodology and biostatistics paper.
One exam, two very different destinies
Here is something that does not get discussed enough: AIAPGET is not just one exam feeding one kind of training.
The same rank list is used to allot seats across both clinical branches (Kayachikitsa, Shalya, Shalakya, Prasuti-Stree Roga, Panchakarma, and others, where diagnostic reasoning and day-to-day patient handling are the job) and non-clinical, pre-clinical and scholarly branches (Samhita-Siddhanta, Kriya Sharira, Rachana Sharira, Dravyaguna, Rasashastra, Swasthavritta, and others, where deep, rigorous engagement with classical texts and sustained scholarly analysis are the job).
Candidates then choose their branch by rank and seat availability — not by any part of the exam that actually measured aptitude for that specific kind of work.
Think about what that means in practice.
A candidate with fast, wide rote recall — exactly what this exam rewards, as the data above shows — can out-rank a more clinically thoughtful peer and walk into a clinical-branch seat on recall speed alone, with no part of the exam having tested whether they can actually reason through a patient. And on the other side, a candidate genuinely suited to the deep interpretive, analytical reading that the non-clinical/Samhita branches demand has no place in this exam to demonstrate that specific skill either — the Match-List and Statement-based items test surface-level classical facts, not sustained interpretive engagement with a text. The exam sorts everyone by the same recall-speed metric, then hands out two very different kinds of careers based on where that metric happened to place them.
I will add something here from my own experience rather than the data, because I think it matters and I have seen it consistently for three decades of teaching. In my experience, close to 90% of the students who enter clinical branches arrive with a strikingly weak grounding in Anatomy, Physiology and Pathology, and with clinical reasoning and patient-handling skills that essentially have to be built from scratch during their PG training. And it is not that the non-clinical branches are quietly absorbing a stronger, more classically grounded pool instead — in my experience their command of Samhita reading and interpretive analysis is, just as often, abysmal.
What both groups tend to share is that they arrived at PG having cleared an exam that never actually tested whether they were ready for either kind of training.
A great many candidates, in my experience, simply sit for AIAPGET and try their luck, without any real understanding of what postgraduate training — clinical or scholarly — is actually going to demand of them. That is, I think, exactly what you would expect an exam built almost entirely on recall speed to produce.
None of this is inevitable. A redesigned process can correct for it directly — starting with a decision I think should be non-negotiable, not optional: candidates should have to declare, before they sit any paper at all, which track they are competing for — clinical or non-clinical — and the assessment itself should stop pretending the two are the same job.
What a better exam would look like
Alongside this data audit, I put together a proposed redesign — not as a finished policy document, but as a starting point for the conversation. The pieces below are worth setting out in more detail, with sample items, since ‘make it more applied’ means very little without showing what that actually looks like on paper.
First: split the process by branch, not just by difficulty
Every candidate should declare, at the point of application, whether they are competing for a clinical branch or a non-clinical/scholarly branch.
Once that declaration is made, the assessment itself should differ between the two tracks — not just the interview stage, but the written paper.
A clinical-track paper should weight vignette, sequencing and patient-reasoning content most heavily.
A non-clinical-track paper should weight deep classical-text interpretive analysis, comparative-commentary reasoning, and research-methodology content most heavily.
Both can share a common recall core — the foundational classical and allied knowledge every Ayurveda postgraduate should hold regardless of branch — but the differentiating sections need to test what each track actually requires, rather than ranking both tracks on the same recall-speed metric and letting seat allotment sort out the rest.
1. A rebalanced written paper — with sample items
Keep the same 120-question, 480-mark architecture the exam already uses, but shift the weighting deliberately, and make sure every category is actually testing what it claims to. One illustrative sample per category:
| Category | Sample item |
|---|---|
| Classical/textual recall | According to Charaka, the Moola of Pranavaha Srotas is located at: Hridaya & Dashapranayatana / Amashaya & Annavahini Dhamani / Basti & Medra / Kloma & Yakrit |
| Modern/allied recall | Plasma drug concentration falling by 50% defines: onset of action / peak concentration / half-life / bioavailability |
| Match-List | Match List-I (Dosha vitiated) with List-II (Kushtha feature) per Charaka — (A) Vataja (B) Pittaja (C) Kaphaja (D) Raktaja against (I) Kandu (II) Daaha (III) Kathinya (IV) Shyava-varna; choose the correct combination. |
| Statement / Assertion-Reason | Assertion (A): Snehana is withheld in the acute stage of Amavata. Reason (R): Snehana can aggravate Ama and worsen Sandhi-shotha before Ama is digested. Both true, R explains A / both true, R does not explain A / A true R false / A false R true. |
| Multi-statement combination | Consider: (A) Basti is the preferred route for Vata-predominant disorders (B) Basti is contraindicated in active, undigested Jwara (C) Sneha Basti needs no Purva Karma. Correct combination: A & B only / B & C only / all three / none. |
| Genuine sequencing | A patient’s Amashayagata Vata is left untreated. Using Charaka’s Shatkriyakala staging, arrange from earliest to most advanced: vague discomfort with no sign / dosha accumulation, no symptom / fully manifest Vata-vyadhi / structural, complicated stage. |
| Clinical vignette | A 34-year-old presents with morning stiffness over an hour, symmetrical small-joint swelling, worse in cold and after curd. Most likely diagnosis and first line of management: Vatarakta — Raktamokshana first / Amavata — Langhana & Deepana-Pachana before Shodhana / Sandhigata Vata — immediate Snehana Basti / Vishama Jwara — Tikta Kashaya first. |
| Calculation | A screening test shows 180 true positives, 20 false negatives, 15 false positives, 285 true negatives against confirmed diagnosis. Sensitivity is closest to: 75% / 85% / 90% / 95%. |
2. A written assignment that actually differs by track
The mandatory written assignment should not be the same exercise for both tracks either. For clinical branches, it should be a structured case write-up drawn from the candidate’s own internship, not a generic essay:
Sample prompt (clinical): ‘Present one case of Amavata (or an equivalent condition) from your internship. Include: (a) the Nidana Panchaka as you identified them in this patient; (b) differential diagnoses you considered and how each was ruled out; (c) the management principle you chose, with your classical rationale (Yukti); (d) the outcome; (e) a short reflective note on what you would do differently with hindsight.’
For non-clinical branches, the equivalent should be a scholarly textual-analysis exercise, since case logs are not the relevant evidence of readiness for that track:
Sample prompt (non-clinical): ‘Present a comparative analysis of how Rasa Dhatu is described across two classical authorities of your choice. Include: (a) points of agreement; (b) points of divergence and your reasoning for why they may have arisen; (c) how later commentators have interpreted the divergence; (d) your own reasoned position, with justification.’
3. A comprehension and analytical-skills test
This should be passage-based, not a list of independent logic puzzles — a short piece of text followed by questions that test whether the candidate can actually reason about what it says, not just recall it. Sample:
‘Between 2018 and 2024, sanctioned Ayurveda PG seats in India rose by roughly 30%, while AIAPGET applicant numbers over the same period grew by only about 8%. Several new PG colleges opened in this period, most in states that already had the highest existing seat density. Meanwhile, cut-off ranks for several non-clinical branches fell sharply, while cut-offs for the most sought-after clinical branches stayed largely stable.’
- Which conclusion is best supported by the passage — falling national demand, seat growth concentrated in already seat-dense states and outpacing demand, declining clinical-branch competitiveness, or rising applicant interest in non-clinical branches?
- If new seats were shown to have gone mostly to low seat-density states instead, which claim in the passage would most need revising?
4. A structured, multi-station interview — different stations for different tracks
The Multiple Mini Interview format — several short, independently scored stations rather than one long panel conversation — should also be built around what each track actually needs to assess:
Clinical-branch stations:* *
- a diagnostic-reasoning station (reason aloud through a case under questioning);
- an ethical-dilemma station (e.g. a patient requesting treatment outside evidence or scope);
- a data-interpretation station (reading a small lab report or dataset);
- a patient-communication station (explaining a diagnosis to a simulated patient);
- a research-critique station (critiquing a study design relevant to the chosen specialty); and
- a motivation/fit station.
Non-clinical-branch stations:* *
- a textual-interpretation station (given a classical passage, interpret it and defend that reading under questioning);
- a scholarly-integrity station (e.g. handling conflicting commentarial views or proper attribution);
- a data/evidence station (interpreting results from an analytical or pharmacological study); *
- a teaching-communication station (explaining a complex classical concept to a lay or student audience);
- a research-critique station (oriented to non-clinical research — drug standardisation, pharmacognosy, educational research); and
- a motivation/fit station.
5. The rest of the framework
No single paper, however well redesigned, can establish PG readiness on its own at this scale.
Beyond the pieces above, the fuller proposal adds an
- English aptitude and reasoning test,
- a foundational research-methodology and biostatistics paper,
- a dedicated PG/healthcare-systems general knowledge test, and
- structured, rubric-based reference letters from a clinical supervisor and an academic mentor.
Looking at how other health-profession systems worldwide screen postgraduate candidates — the UK, Canada, the US, Australia — a few more pieces are worth adding on top:
- a structured statement of purpose,
- a situational-judgement test for professional scenarios, and
- a verified portfolio of case logs or scholarly work,
none of which India’s current process asks for at all.
Where this goes next
This audit and the proposed redesign are both drafts, and I intend to keep both open for scrutiny and correction — particularly the 2022 and 2026 figures, which rest on coaching-institute memory-based reconstructions rather than official releases, and are flagged as such throughout. Once the official 2026 paper is released, I will update the dataset and the analysis accordingly.
My purpose in putting this in public view is simple: an exam that decides who gets to become a postgraduate Ayurveda specialist deserves to be evaluated on the same evidentiary standards we would expect of any other high-stakes assessment — not defended or dismissed on impression.
Comments, corrections, and disagreement are genuinely welcome.
Share your thoughts in the comments below.
💬 Comments & Discussion