Measure what emerges over time.
Mohamed A M Elansary, PhD — long-horizon evaluation under uncertainty, scientific/HPC execution, and production agent regression evaluation.
Scientific measurement
- Designed multi-model forecast comparisons across basins and hydroclimates.
- Quantified uncertainty and validated imperfect USGS, NOAA, and NASA observations.
- Ran reproducible Python, R, Bash, Linux, and HPC workflows.
Production systems
- Builds GPT, Claude, and Gemini agent workflows at Vertexium.
- Maintains regression evaluation sets for production agent behavior.
- Ships retrieval, routing, tenant isolation, provenance, and validation systems.
Proposed evaluation approach
Define observable indicators of useful novelty and competence progress; establish simple baselines; stratify outcomes across seeds, environments, and horizons; quantify uncertainty; and distinguish exploration from collapse, reward exploitation, or trivial objective discovery.
Honest fit boundary
Direct reinforcement-learning and open-endedness research depth is a stretch. I have not developed RL training methods, automatic curricula, intrinsic-motivation systems, self-play systems, or LLM post-training methods. My contribution is rigorous evaluation under uncertainty, scientific/HPC execution, and production agent regression evaluation.
Role and location
Member of Technical Staff - Open Endedness · “this role is based in our San Francisco office; for exceptional candidates we are willing to consider a hybrid arrangement”. Relocation with a support package is an honest discussion point; hybrid is not assumed.
“The expected salary range for this position is $300,000 - $500,000 USD” · Official role posting