Skip to content
Darwin, NT · relocating from SydneyUpdated · August 2026

Current · Doctoral Researcher · CDU · Responsible Generative AI

Applied AI · Responsible Generative AI · Digital Systems

Shaurav Khadka

Systems lens

I build reliable AI systems and research how generative AI can be evaluated, governed, and integrated without losing human agency.

My work spans production document AI, semantic retrieval and RAG, temporal graph learning, computer vision and robotics, reinforcement learning, data workflows, and technical governance. I bring more than three years of CTO-level technology leadership, industry AI/ML R&D at Truuth, and doctoral research at Charles Darwin University on standards-aligned responsible Generative AI.

Applied systems and research interests

Built systems lead. Emerging interests are stated separately and with scope.

Applied Systems Base

Built, evaluated, and delivered systems

  • Production AI Reliability
  • Document Intelligence
  • Semantic Retrieval & RAG
  • Temporal Graph Learning
  • Computer Vision
  • Robotics & Sim2Real
  • Reinforcement Learning
  • Data & API Workflows
  • Technology Strategy

Current Research & Governance

Doctoral direction and active research interests

  • Responsible Generative AI
  • AI Governance
  • Standards-Aligned Systems
  • Human Agency & Oversight
  • Privacy & Data Governance
  • Transparency & Traceability
  • AI Literacy
  • Socio-Technical Evaluation
  • Decision Support & Digital Transformation

Measured Highlights

Start with what changed.

Three benchmarks across deployment adaptation, reinforcement learning, and retrieval.

Deployment-specific Sim2Real adaptation

2.38% 95.24%

Baseline2.38%
Adapted95.24%

Robot-image accuracy after deployment-specific Sim2Real adaptation

The model looked strong on curated data and degraded sharply on robot-camera images. The recovery came from treating domain shift as a deployment problem, not a footnote.

Why it matters: the adaptation restored useful robot-camera performance under changed lighting, viewpoint, scale, and background conditions.

Robot-camera deployment prediction comparison after Sim2Real adaptation

Robot-camera predictions · published team-level result

Evaluation context
Baseline
2.38% before deployment-specific adaptation.
Measured
95.24% robot-image accuracy after targeted collection, augmentation, and fine-tuning.
Conditions
Robot-camera inputs with lighting, viewpoint, scale, and background differences.
Attribution
Collaborative team-level result with exported notebook figures.

Supporting benchmark 02

300 → 1,925

AirRaid PPO mean reward after temporal observation changes

Observation design materially changed what the policy could learn. Frame skipping and frame stacking improved the benchmark result without pretending algorithm choice was the only lever.

Why it matters: the result shows that observation design can materially change what a policy learns before the algorithm itself is replaced.

Supporting benchmark 03

P@5 = 0.68 · R@5 = 0.68

RedditPulse semantic retrieval quality

The retrieval layer was measured before generation was treated as useful. That matters because grounded insight quality depends on which sources the system surfaces first.

Why it matters: downstream summaries are only as useful as the source material retrieved before generation begins.

Working Method

From raw inputs to dependable deployment.

A practical evaluation loop grounded in the way I build and inspect applied AI systems.

  1. 01

    Ingest

    Map inputs, schemas, edge cases, constraints, and the operational path around the model.

  2. 02

    Benchmark

    Establish reproducible baselines and measurable success conditions before tuning the system.

  3. 03

    Error-analyse

    Trace failures across data, model, transformation, validation, and review boundaries.

  4. 04

    Adapt

    Change the representation, workflow, threshold, or model only where the findings justify it.

  5. 05

    Deploy

    Document limitations, preserve traceability, and translate results into a workflow people can inspect.

Applied Systems

Selected applied systems.

Production reliability, temporal learning, retrieval, robotics, evaluation, and active doctoral research in responsible Generative AI.

Prior industry workflowTRUUTH · Former AI/ML R&D Internship

Production AI Reliability and Document Intelligence

Problem: Document intelligence can fail long before or after OCR. Real reliability depends on the complete path from ingestion to extraction, transformation, validation, and review.

Contribution: Built repeatable evaluation workflows across OCR configurations, mappings, confidence scores, error codes, and reruns while preserving traceability and review boundaries.

PythonpandasAWS S3boto3Azure Document IntelligenceJSON
Inspect case study

01

OCR

02

Map

03

Validate

04

Trace

Shared here: sanitised workflow record. Confidential operational data and internal metrics are excluded.

AI engineering, data, digital systems, governance, research, and technology leadership connected through one systems-oriented practice.

Supporting builds

Five supporting systems across research, prototypes, and evaluation.

Each card states the system, category, and route for deeper inspection.

Temporal Graph Learning

Temporal graph-learning research build

02

Temporal GNN for Blockchain Fraud Detection

Fraud is relational and time-dependent. Static tabular features can miss how transactions evolve across a network.

t0 → t1 → t2

Inspect case study

Generative AI · Conversational Systems

Scoped conversational-AI prototype

03

LLM-Based Financial Assistant Prototype

Conversational assistants can produce fluent but poorly scoped responses. This prototype explores structured prompting, model comparison, synthetic profiles, and explicit safety boundaries.

profile → prompt → compare → respond

Inspect case study

NLP · Information Retrieval

Modular retrieval research toolkit

04
R01
R02
R03

Semantic Search and Information Retrieval Engine

Keyword matching is transparent but limited when meaning varies across phrasing. The system needed a modular comparison path from classical retrieval to dense semantic search.

clean → encode → rank → evaluate

Inspect case study

Machine Learning · Data Science

Reusable experimental evaluation pipeline

05
Benchmark 1
Benchmark 2
Benchmark 3
Benchmark 4
Benchmark 5

Machine-Learning Evaluation and Data-Science Pipeline

A model result is only useful when the path from raw data to evaluation is reproducible, comparable, and explicit about failure cases.

data → features → compare → inspect

Inspect case study

Responsible Generative AI · Governance

Doctoral research scope · framework under development

06

Standards-Aligned Responsible Generative AI and Human Agency

Responsible Generative AI principles are not enough unless they become concrete controls, evaluation criteria, human decision boundaries, and governance mechanisms that work in real settings.

standards → controls → oversight → evidence

Inspect case study

Project index

Browse all project routes.

A compact index of results, methods, and prototypes.

09 routes

Research Profile

Research Program & Directions

The centre of gravity has changed: my current trajectory is responsible Generative AI, human agency, standards-aligned system design, and evidence-based governance — built on a technical foundation in AI evaluation and deployment.

My doctoral work begins with teacher education as the empirical domain, but the underlying systems problem is broader: how to translate responsible-AI principles and standards into controls, evaluation criteria, governance mechanisms, and human-oversight boundaries that can actually be tested in practice.

Discuss research or collaboration
Current doctoral research

Human Agency and Responsible Generative AI

Designing and evaluating a standards-aligned socio-technical reference framework for responsible Generative AI, with teacher education and professional learning as the initial empirical domain. The research connects technical controls with governance, human oversight, privacy and data governance, transparency, AI literacy, and institutional decision-making.

Doctoral research program

Questions I am building toward

  • How can responsible-AI principles and standards be translated into testable technical and organisational controls?
  • Which decisions should remain meaningfully human, and how should oversight and escalation boundaries be designed?
  • How should privacy, transparency, traceability, AI literacy, and system effectiveness be evaluated together rather than in isolation?

Applied research foundation

Reliable AI Systems and Production Evaluation

Evaluation of AI pipelines where traceability, robustness, confidence handling, validation dependencies, regression risk, latency, cost, and human-review boundaries matter alongside headline accuracy.

Questions I want to pursue

  • How should reliability be measured across the full decision pipeline?
  • How can failure analysis distinguish data, model, transformation, validation, and workflow faults?

Active research area

Retrieval, Grounding and Knowledge-Centred AI

Semantic retrieval, RAG, multilingual discovery, and evidence-gated generation — with emphasis on whether the system retrieves the right evidence before generated language is treated as useful.

Questions I want to pursue

  • How should retrieval quality, provenance, and abstention shape downstream generation?
  • What evaluation designs distinguish fluent output from grounded and decision-useful output?

Applied systems direction

Human-Centred AI, Decision Support and Digital Transformation

A systems direction connecting AI engineering, organisational workflows, governance, decision support, and technology adoption. The focus is not automation for its own sake, but designing digital systems that improve decisions, preserve meaningful human control, and produce measurable operational value.

Questions I want to pursue

  • Where should AI automate, augment, recommend, or deliberately defer to human judgement?
  • How can digital transformation be evaluated through workflow quality, adoption, traceability, decision outcomes, and operational value rather than novelty alone?
Research mapResearch architectureThree foundation groupsOpen map

The program connects governance and standards to technical evaluation rather than treating them as separate conversations. Existing work in production reliability, retrieval, temporal modelling, and deployment adaptation provides the applied base.

Responsible GenAI and governance

Current doctoral layer: operationalising responsibility into system and organisational design.

  • Standards alignment
  • Human agency
  • Human oversight
  • Risk controls
  • Privacy and data governance
  • Transparency
  • Traceability
  • AI literacy

Evaluation and deployment

Applied base built through production-oriented AI R&D and benchmarked systems work.

  • Experiment design
  • Baseline comparison
  • Error analysis
  • Confidence analysis
  • Regression testing
  • Domain shift
  • Human-review boundaries
  • Reproducibility

Systems and leadership

Engineering and organisational experience that supports socio-technical research rather than model-only analysis.

  • AI/ML engineering
  • Data workflows
  • APIs
  • Digital systems
  • Technology strategy
  • Stakeholder coordination
  • Operational workflows
  • Decision-ready reporting

Decision systems and digital transformation

Bridging technical capability with organisational adoption, workflow design, decision quality, and measurable operational outcomes.

  • Human-centred AI
  • Decision support
  • Digital transformation
  • Workflow redesign
  • Automation assessment
  • Technology adoption
  • Data-informed operations
  • Socio-technical systems

Experience

Technical Experience & Leadership

Production-oriented AI R&D, more than three years of CTO-level technology leadership, and hands-on software engineering across AI, data, APIs, GIS, and operational systems.

  1. Truuth

    AI/ML Research and Development Intern

    Feb 2026 — Jun 2026

    Sydney, NSW, Australia · Hybrid

    Completed a 13-week industry AI/ML R&D major project on production document-intelligence reliability and adversarial fraud-detection evaluation. Built repeatable workflows across ingestion, OCR configuration, field mapping, transformations, validation, confidence review, reruns, and structured error analysis using Python, pandas, AWS S3/boto3, Azure Document Intelligence, and JSON; the project was awarded 83/100 (Distinction).

    • Document AI
    • Fraud Evaluation
    • Python
    • AWS S3
    • Azure Document Intelligence
    • Reliability
  2. Picpoint Nepal Pvt. Ltd.

    Chief Technology Officer

    Jun 2021 — Jun 2024

    Kathmandu, Nepal · Hybrid

    Owned technology strategy and continuous improvement across web platforms, databases, APIs, GIS/mapping inputs, data flows, export-logistics workflows, customer management, and digital operations. Translated organisational requirements into roadmaps, SOPs, implementable systems, and decision-ready recommendations while coordinating technical and non-technical stakeholders in a resource-constrained environment.

    • Technology Strategy
    • Systems Delivery
    • APIs
    • Databases
    • GIS/Mapping
    • Operational Workflows
  3. Thakur International

    Junior Full Stack Developer

    Jun 2019 — May 2020

    Kathmandu, Nepal · On-site

    Developed and maintained web and mobile components using PHP, Python, and JavaScript; integrated REST/SOAP APIs, OAuth authentication, and Google Maps/geolocation workflows; and contributed to debugging, refactoring, performance analysis, and agile sprint delivery.

    • PHP
    • Python
    • JavaScript
    • REST/SOAP APIs
    • OAuth
    • Google Maps
Additional experienceSupporting work completed while studying

Ingleburn Convenience Store

Operations and Digital Support Assistant · Part-time

Oct 2024 — Jun 2026

Supported transaction and inventory accuracy, POS troubleshooting, basic network and hardware issues, digital administration, and customer-facing operations while completing postgraduate study in Australia.

Foundation

Education & Research

Formal academic progression from software engineering and computing into applied AI and standards-aligned responsible Generative AI research.

Education

Charles Darwin University

Doctor of Philosophy (PhD) · Responsible Generative AI

2026 — Present · Casuarina Campus · Darwin, NT, Australia

Doctoral research supervised by Jon Mason: a standards-aligned technical reference framework for responsible Generative AI in teacher education, with focus on risk controls, human oversight, privacy and data governance, transparency, and AI literacy.

Education

Macquarie University

Master of Information Technology · Artificial Intelligence

Qualified 8 Jul 2026 · Sydney, NSW, Australia

Industry AI/ML R&D major project at Truuth: 83/100 (Distinction). Relevant study included Advanced Machine Learning, AI for Text and Vision, Data Science, AI Ethics and Law, Automated Decision Making, Knowledge, Planning and Decision Making under Uncertainty, and Advanced Topics in AI.

Education

London Metropolitan University · Islington College

BSc (Hons) Computing · First Class Honours

Awarded Mar 2021 · Kathmandu, Nepal

Final-year applied software-engineering project: an integrated trip-planning and travel-experience platform using PHP/Laravel, MySQL, HTML/CSS, and JavaScript.

Continuous learning

Selected Professional Development

Targeted learning that strengthens the bridge between AI engineering, responsible use, teaching, governance, computer science, and leadership.

Artificial Intelligence (AI) Education for Teachers

Macquarie University + IBM · Coursera

Completed 15 Aug 2026

AI literacy, teaching practice, and responsible classroom integration.

University Teaching

The University of Hong Kong · Coursera

In progress · Aug 2026

Higher-education teaching, learning design, and reflective academic practice.

Ethics and Governance of Artificial Intelligence for Health

World Health Organization

Completed 2023

Ethics, governance, accountability, and responsible AI in a high-stakes domain.

CS50x: Introduction to Computer Science

Harvard University

Completed

Computer-science foundations, programming, algorithms, data structures, and systems thinking.

Design Thinking

Macquarie University Incubator × KPMG

Completed 2026

Human-centred problem framing, ideation, validation, and collaborative innovation.

UPG Sustainability Leader

United People Global

Certified with Distinction

Sustainability leadership, citizen action, community impact, and SDG-oriented initiatives.

Technical Outputs

Capability Map

A deeper inventory of methods and tools, each linked to the closest inspectable system or research route instead of presented as an ungrounded keyword cloud.

Responsible Generative AI, Governance and Standards

  • Responsible Generative AI
  • AI governance
  • Standards-aligned system design
  • Human oversight
  • Human agency
  • Privacy and data governance
6 additional methods
  • Transparency
  • Traceability
  • AI literacy
  • Risk controls
  • Socio-technical evaluation
  • Technical communication

Generative AI, RAG and Multilingual Interaction

  • Large Language Models
  • Retrieval-Augmented Generation
  • Document chunking
  • Vector embeddings
  • Context injection
  • Prompt design
4 additional methods
  • Grounded generation
  • Language detection
  • Translation workflows
  • Cross-lingual retrieval

NLP, Information Retrieval and Social Intelligence

  • Text preprocessing
  • Tokenisation and lemmatisation
  • Named-entity recognition
  • TF-IDF baselines
  • Sentence-transformer embeddings
  • FAISS retrieval
4 additional methods
  • Ranking logic
  • Sentiment modelling
  • Topic classification
  • Trend and community analysis

Machine Learning, Data Science and Evaluation

  • Supervised learning
  • Classification
  • Predictive modelling
  • Clustering
  • Exploratory data analysis
  • Feature engineering
6 additional methods
  • Cross-validation
  • Hyperparameter tuning
  • Confusion-matrix analysis
  • Comparative benchmarking
  • Error analysis
  • Regression testing

Computer Vision, Robotics and Sim2Real

  • Convolutional neural networks
  • Transfer learning
  • Fine-tuning
  • Image augmentation
  • Fine-grained recognition
  • Domain-shift analysis
4 additional methods
  • Robot-camera adaptation
  • ROS2 integration
  • Confidence thresholds
  • Vision-to-action pipelines

Reinforcement Learning and Dynamic Decision Systems

  • Q-learning
  • DQN
  • PPO
  • Sparse-reward environments
  • Reward shaping
  • Frame stacking
6 additional methods
  • Frame skipping
  • Temporal observations
  • Training-curve analysis
  • Policy evaluation
  • Temporal graph learning
  • Dynamic risk modelling

Production AI, Data and Cloud Workflows

  • OCR evaluation
  • Azure Document Intelligence
  • Structured extraction
  • AWS S3
  • boto3
  • JSON transformation
8 additional methods
  • Confidence analytics
  • Error-code analysis
  • Adversarial testing
  • REST/SOAP APIs
  • SQL
  • Docker
  • Linux
  • Auditability and deployment risk

Technology Leadership and Delivery

  • Technology strategy
  • Systems implementation
  • Digital operations
  • Stakeholder coordination
  • Roadmapping
  • SOPs
4 additional methods
  • Operational workflows
  • Automation assessment
  • GIS and mapping
  • Decision-ready reporting

Closest output

Toolkit

Research Methods & Engineering Toolkit

Methods, implementation tools, and delivery capabilities used across research, production-oriented evaluation, and digital systems work.

Actively developing

Research, evaluation and governance methods

18 methods

Methods used to move from technical performance toward reliable, auditable, and responsible system decisions.

  • Experimental Design
  • Baseline Comparison
  • Benchmarking
  • Error Analysis
  • Confidence Analysis
  • Failure-Mode Analysis
  • Retrieval Evaluation
  • Temporal Modelling
  • Graph Learning
  • Domain-Shift Evaluation
  • Regression Testing
  • Traceability
  • Reproducibility
  • Risk Analysis
  • Human-Oversight Design
  • Standards Mapping
  • Socio-Technical Analysis
  • Technical Governance

Used in applied systems

Applied engineering and delivery toolkit

30 tools

Languages, libraries, platforms, and workflow tools used across AI, data, software, cloud, robotics, and digital operations.

  • Python
  • SQL
  • JavaScript
  • PHP
  • Java
  • C#
  • PyTorch
  • TensorFlow
  • scikit-learn
  • Hugging Face
  • Sentence Transformers
  • FAISS
  • NetworkX
  • pandas
  • NumPy
  • Matplotlib
  • Streamlit
  • Jupyter
  • ROS2
  • Docker
  • Linux
  • AWS S3
  • boto3
  • Azure Document Intelligence
  • JSON
  • REST/SOAP APIs
  • MySQL
  • Oracle
  • Git
  • GIS/Google Maps APIs

Community Impact

Community Technology and Leadership

Field technology support and selected leadership programs connected to sustainability, peer guidance, design thinking, and cross-cultural collaboration.

Field support

Solar and IT systems

Service continuity

2015 — Present

Leadership layer

4 selected programs

Long-term field technology

Technical Volunteer and Systems Support

Swogun Energy

Supported field deployment, testing, and troubleshooting of small-scale solar-power and IT systems in remote and off-grid settings in Nepal. Continues to provide occasional remote technical and digital support while based abroad.

  • Solar Systems
  • Field Support
  • IT Troubleshooting
  • Remote Communities

Supporting programs

Leadership, sustainability, peer support, and design thinking.

Selected record

United People Global

Certified UPG Sustainability Leader

Completed sustainability-leadership training focused on community-driven initiatives, positive citizen action, and the United Nations Sustainable Development Goals.

2024 — 2025Global online program

Aspire Institute

Aspire Leaders Program Alumnus and Peer Support Contributor

Completed leadership-development training and continues to support emerging participants through occasional peer guidance and resource sharing.

2023 — PresentGlobal online program

Macquarie University

MQ Incubator × KPMG Design Thinking

Applied human-centred problem solving, opportunity framing, and collaborative ideation within an innovation-focused program.

2026Sydney, NSW, Australia

Macquarie University

Postgraduate Global Leadership Program Graduate

Completed a university leadership-development program focused on reflective practice, cross-cultural collaboration, and professional growth.

2025Sydney, NSW, Australia

About

AI Systems, Governance and Execution

I am an applied AI and digital-systems professional, technology leader, and Doctoral Researcher at Charles Darwin University. My background combines hands-on AI/ML evaluation, software engineering, data and cloud workflows, more than three years of CTO-level technology leadership, and research communication.

The consistent thread is systems thinking. I care about what happens before and after a model: data quality, representations, baselines, failure modes, confidence, validation, human review, governance, operational constraints, and the decisions that the system eventually influences.

My current research direction asks how standards, technical controls, human agency, privacy, transparency, and AI literacy can be connected in real Generative AI deployments. Teacher education is the initial empirical domain; the design problem is intentionally broader and socio-technical.

Alongside the doctoral program, I continue to work across applied AI, data, decision support, digital transformation, and system evaluation. The aim is to connect technical depth with the organisational realities that determine whether technology is actually useful.

Research stance

Build what can be inspected.

Measure before claiming.

Treat governance as part of system design.

Keep human decision boundaries explicit.

Make the evidence trail stronger than the rhetoric.

Charles Darwin University

Doctor of Philosophy (PhD) · Responsible Generative AI

2026 — Present

Macquarie University

Master of Information Technology · Artificial Intelligence

Qualified 8 Jul 2026

London Metropolitan University · Islington College

BSc (Hons) Computing · First Class Honours

Awarded Mar 2021

Research Writing

Research Notes and Engineering Decisions

Four on-site notes on research questions, evaluation choices, and engineering decisions, with four DOI-linked technical outputs below.

03 notesORCID iD ↗
Research output trailPreprints, reports, and DOI-linked outputs are maintained on the ORCID record and linked directly below.0009-0009-9874-8239
DOI-Linked Technical OutputsEarly research outputs deposited with persistent DOI identifiers. These are non-peer-reviewed technical preprints and reports.04 outputs

Independent Publishing

Books & Independent Publishing

Independent authorship, illustration, and editorial credits presented as a compact publishing record.

Featured authored publication

The Digital Equilibrium

Navigating Technological Advancement for Optimal Well-Being

An independent authored work exploring how technological progress can be balanced with human well-being and intentional living.

A small publishing trail spanning technology, well-being, and selected creative collaboration.

Collaborative editions6 selected creative credits

Illustrated and editorial work

Selected illustration and editorial credits across children’s stories and reflective writing.

Browse author page
Browse titles

Illustrator · Creative contributor

Joyful Stories

Joyful Stories

Illustrator · Creative contributor

Joyful Stories

Mazzako Katha · Alternate edition

Illustrator · Creative contributor

2 in 1 Joyful, Children Stories

Combined children’s-story edition

Contact

Systems, research, and ideas worth building.

I am establishing my professional and research base in Darwin. This portfolio is a working record of the systems I have built, the results I can defend, the research I am developing, and the technical problems I care about next.

Research

Research collaboration

For responsible Generative AI, AI governance, human oversight, retrieval and grounding, evaluation, or standards-aligned system design, start with the research program and technical outputs.

Work

Applied AI and digital systems

My applied work sits across AI engineering, data and cloud workflows, digital systems, evaluation, automation, and technology leadership. The portfolio is designed to make that work inspectable rather than reduce it to a list of claims.

Continue exploring

Explore more of the portfolio before reaching out.

Current status

Doctoral Researcher at Charles Darwin University · Master of Information Technology (Artificial Intelligence) · Darwin, NT.

Darwin focus

Darwin, NT · establishing a long-term professional and research base.

Cross-domain range

AI engineering, data, digital systems, governance, research, and technology leadership connected through one systems-oriented practice.

Darwin, NT · applied AI · data · digital systems · researchGitHub ↗LinkedIn ↗ORCID iD ↗