AI Systems Engineering
End-to-end model engineering across training, evaluation, quantization, inference serving, and application integration.

Founder & CEO · AI Systems Architect
I connect enterprise software engineering with model training, inference optimization, and AI product delivery. My current focus includes production-grade vLLM and SGLang serving, RAG, agentic systems, and multimodal applications.
I am a hands-on technology leader with experience across system architecture, product delivery, distributed platforms, data systems, and generative AI. Starting from business outcomes, I design end-to-end solutions spanning data governance, training and evaluation, inference optimization, application orchestration, and cloud-native operations.
End-to-end model engineering across training, evaluation, quantization, inference serving, and application integration.
Scalable distributed, data, and cloud-native systems designed for availability, concurrency, and complex collaboration.
Connecting business, design, and engineering to turn technical prototypes into measurable enterprise products.
Beyond calling model APIs, I build evolvable AI systems around data, models, compute, serving, and applications.
Efficient domain adaptation, alignment, and evaluation with enterprise data
System optimization for throughput, latency, memory efficiency, and reliability
Orchestrating model capabilities into traceable and testable workflows
Backend, frontend, and infrastructure capabilities for continuous AI delivery
Selecting and tuning inference engines around model size, context patterns, concurrency, and cost objectives.
OpenAI-compatible serving, PagedAttention, continuous batching, quantized deployment, and multi-GPU scaling for high-throughput inference.
RadixAttention, prefix caching, structured generation, and efficient execution for complex agentic and multimodal workflows.
A long-term entrepreneurial and technical leadership track, complemented by hands-on enterprise roles and industry projects.
Wenku Technology is the continuous entrepreneurial track. Other entries represent concurrent enterprise roles, advisory work, or project engagements. The two-track format makes those responsibilities and timelines explicit.
Leading technology strategy, product direction, and core architecture for AI-native enterprise solutions. Deeply embedded with the RTE (Release Train Engineer) team, partnering with product and delivery teams to introduce AI Agents across requirements analysis, architecture, coding, testing, and operations, and to establish an AI Native development model. The technology portfolio spans LLM fine-tuning and high-performance inference, RAG, MCP, agentic long-term memory and evaluation, computer vision, OCR, and cloud-native distributed systems. Driving platformization, engineering standards, observability, and governance to turn project knowledge into reusable organizational intelligence, while remaining hands-on in critical architecture, prototyping, and production implementation.
Contributed to centralized payment platforms and enterprise architecture, focusing on transaction security, availability, scalability, and cross-team delivery.
Built data-platform and analytics capabilities that turned fragmented information into reusable data assets and decision insight.
Supported enterprise data and digital initiatives through business analysis, data modeling, and system delivery for large-scale transformation.
Advanced study in scalable computing, enterprise cloud applications, big data, machine learning, and modern system architecture.
Focused on language understanding, text analytics, machine learning models, and applied NLP engineering.
Built a rigorous foundation in computer science, operating systems, networking, and computer architecture.
If you are planning an enterprise AI program, inference platform, or complex system modernization, let’s discuss the technology path and delivery strategy.