To assess HIV knowledge, stigma and perceived adequacy of HIV curricular coverage among medical students in Egypt and to identify factors associated with HIV stigma.
An online-based cross-sectional study.
Medical schools across Egypt. Data were collected in August 2025 using a bilingual (Arabic/English) online questionnaire using convenience sampling.
First- through fifth-year students enrolled in Egyptian medical schools.
HIV knowledge was assessed using the Brief HIV Knowledge Questionnaire (HIV-KQ-18); HIV-related stigma was assessed using the Healthcare Providers HIV/AIDS Stigma Scale (HPASS) and perceived adequacy of HIV curricular coverage.
A total of 1503 students participated (mean age 20.6 years; 57.4% female), half of whom (48.9%) rated curricular coverage of stigma and psychosocial aspects of HIV as inadequate. The mean HIV-KQ-18 score was 8.96/18 (SD 4.26). Only 39.9% recognised that HIV cannot be transmitted through kissing, 47% believed washing after sex is protective and just 41.3% knew that not all infants born to mothers with HIV will have AIDS. The mean HPASS score was 60.1/108 (SD 17.6). Most students (76.8%) worried about contracting HIV from patients, 52% believed patients acquired HIV through risky behaviours and 43.6% endorsed a right to refuse providing care. Knowledge and stigma were inversely but weakly correlated (r = –0.17, p<0.001), and higher knowledge was independently associated with lower stigma on multivariable regression (B=–0.16, p<0.001). Despite higher knowledge, males reported significantly higher stigma (B=0.25, p<0.001) compared with their female counterparts. Similarly, participants who completed the Arabic form had significantly lower knowledge and higher stigma (B=0.24, p<0.001).
HIV stigma is prevalent among medical students in Egypt, with significant variations observed across gender, survey language and levels of HIV knowledge. These findings call for multifaceted interventions and curriculum reform to reduce stigma among future clinicians.
To assess and compare the performance of four contemporary frontier large language models (LLMs)—GPT-5.2 (OpenAI), Gemini 3 Pro (Google DeepMind), Claude Sonnet 4.6 (Anthropic) and Grok 4.1 (xAI)—on a simulated Fellowship of The Royal College of Surgeons Urology (FRCS(Urol)) Part A examination, evaluating overall accuracy, subspecialty-level performance, output consistency and response time.
Controlled comparative evaluation study using a standardised simulation framework with repeated independent testing runs per model.
All models were accessed via their respective consumer-facing interfaces. No clinical setting or patient data were involved. Testing was conducted under uniform conditions with conversational memory disabled across all sessions.
Four large language models were evaluated. No human participants were involved. Models were selected to represent the current frontier of publicly accessible LLMs from four distinct commercial developers. No models were excluded following selection.
Each model was presented with 240 FRCS (Urol) Part A single best answer questions, mapped to the Joint Committee on Intercollegiate Examinations' Urology Syllabus Blueprint (2023). A standardised prompt was delivered at the start of each session. Each model completed five independent examination runs. No fine-tuning or system-level modification was applied to any model.
The primary outcome was overall examination accuracy for each model, benchmarked against an indicative pass threshold for the FRCS (Urol) Part A examination. Secondary outcomes were performance across 18 individual urology subspecialty topics; response time reported as mean total and per-question elapsed time; and consistency of performance quantified by SD and 95% CIs derived from a sequential Monte Carlo sampling procedure. All outcomes were prospectively planned and fully measured as specified.
Three of four models exceeded the indicative 74% pass threshold: Gemini 3 Pro (82.4%±0.9%; 95% CI 81.3 to 83.6%), Claude Sonnet 4.6 (79.3%±1.1%; 95% CI 77.9 to 80.6%) and GPT-5.2 (76.1%±2.4%; 95% CI 73.1 to 79.1%). Grok 4.1 failed (70.4%±0.6%; 95% CI 69.6 to 71.2%), with its entire CI below 74%. All models completed the assessment in under 3 min. Strong performance was observed in research methodology (90–98%) and andrology (92–98%), with the weakest results in paediatric urology (38.7–54.7%) and testicular cancer (48.2–67.3%). Substantial within-model output instability was identified across several domains, most notably GPT-5.2 in female urology (SD±22.8%) and anatomy (SD±14.2%).
Three of four frontier LLMs achieved scores consistent with passing the FRCS (Urol) Part A examination, representing a substantial advance since ChatGPT-3.5. Aggregate accuracy alone, however, obscures important subspecialty weaknesses and output instability. LLMs should be regarded as adjunctive revision aids rather than authoritative knowledge sources and always used alongside expert-led teaching. Future work should evaluate performance on Part B and viva-style assessments.