Job Description
We are seeking a visionary AI Infrastructure Lead to architect the backbone of our next-generation systems for the 2026 era. At Quantum Leap Systems, we are not just building software; we are defining the future of intelligent computing. In this pivotal role, you will bridge the gap between cutting-edge Machine Learning models and scalable, resilient cloud architectures. You will lead the strategic roadmap to ensure our infrastructure is future-proof, sustainable, and capable of handling exponential data growth.
Why Join Us?
- Shape the infrastructure strategy for the year 2026 and beyond.
- Work with state-of-the-art Large Language Models (LLMs) and Generative AI.
- Competitive equity package and performance bonuses.
If you are a technical leader passionate about performance, scalability, and the future of AI, we want to hear from you.
Responsibilities
- Architect Scalable Solutions: Design and implement high-availability, distributed systems that can scale to millions of concurrent users.
- Model Deployment: Oversee the end-to-end deployment lifecycle of AI models, including containerization, orchestration, and continuous integration/continuous deployment (CI/CD).
- Cloud Optimization: Manage cloud resource allocation (AWS/Azure) to ensure cost-efficiency without compromising performance.
- Security & Compliance: Enforce rigorous security protocols and data governance standards across all infrastructure layers.
- Technical Leadership: Mentor a team of senior engineers and provide technical guidance on complex architectural challenges.
- Performance Tuning: Analyze system bottlenecks and implement optimizations to reduce latency and improve throughput.
Qualifications
- Education: Bachelor’s degree in Computer Science, Engineering, or a related field (Master’s preferred).
- Experience: 7+ years of experience in software engineering, with at least 3 years in a leadership or architectural capacity.
- Technical Stack: Deep expertise in Kubernetes, Docker, Python, and cloud platforms (AWS/GCP/Azure).
- AI Knowledge: Strong understanding of Machine Learning operations (MLOps) and experience deploying LLMs in production environments.
- Problem Solving: Proven ability to troubleshoot complex, multi-layered technical issues under pressure.
- Communication: Excellent verbal and written communication skills, capable of translating technical concepts to non-technical stakeholders.