Back to the team
Zhewen Li
Technology FocusLLM Systems & AI Infra
Current RoleFounder & CEO
AI Engineering · System Architecture · Leadership

Zhewen Li李 哲闻

Founder & CEO · AI Systems Architect

I connect enterprise software engineering with model training, inference optimization, and AI product delivery. My current focus includes production-grade vLLM and SGLang serving, RAG, agentic systems, and multimodal applications.

LLM Training & AlignmentHigh-Performance InferenceAI Agent & RAGEnterprise Architecture
Tokyo · ShanghaiChinese · Japanese · EnglishWenku Technology
Professional Profile

Turning frontier model capabilities into dependable business systems

I am a hands-on technology leader with experience across system architecture, product delivery, distributed platforms, data systems, and generative AI. Starting from business outcomes, I design end-to-end solutions spanning data governance, training and evaluation, inference optimization, application orchestration, and cloud-native operations.

I stay close to critical prototypes and architecture decisions—balancing model quality with latency, throughput, cost, observability, and long-term maintainability.

AI Systems Engineering

End-to-end model engineering across training, evaluation, quantization, inference serving, and application integration.

Enterprise Architecture

Scalable distributed, data, and cloud-native systems designed for availability, concurrency, and complex collaboration.

Product & Delivery

Connecting business, design, and engineering to turn technical prototypes into measurable enterprise products.

AI Technology Stack

A complete capability stack from training to online inference

Beyond calling model APIs, I build evolvable AI systems around data, models, compute, serving, and applications.

Training & Alignment

Efficient domain adaptation, alignment, and evaluation with enterprise data

PyTorchTransformersPEFT / LoRA / QLoRADeepSpeedFSDPAccelerateTRLDPO / RLHFFlashAttentionMegatron-LM

Inference & Serving

System optimization for throughput, latency, memory efficiency, and reliability

vLLMSGLangTensorRT-LLMTriton Inference ServerAWQ / GPTQ / FP8Continuous BatchingSpeculative DecodingKV CacheTensor / Pipeline Parallel

AI Applications & Agents

Orchestrating model capabilities into traceable and testable workflows

RAGAgentic WorkflowMCPMultimodal AIEmbedding & RerankerVector DatabaseEvaluationLLM Observability

Cloud-Native & Full Stack

Backend, frontend, and infrastructure capabilities for continuous AI delivery

PythonGoJavaReact / Next.jsAWSKubernetesDockerDistributed Systems
Inference Engineering

High-performance LLM inference

Selecting and tuning inference engines around model size, context patterns, concurrency, and cost objectives.

vvLLM

OpenAI-compatible serving, PagedAttention, continuous batching, quantized deployment, and multi-GPU scaling for high-throughput inference.

SSGLang

RadixAttention, prefix caching, structured generation, and efficient execution for complex agentic and multimodal workflows.

Career Architecture

A dual-track career

A long-term entrepreneurial and technical leadership track, complemented by hands-on enterprise roles and industry projects.

Why do some dates overlap?

Wenku Technology is the continuous entrepreneurial track. Other entries represent concurrent enterprise roles, advisory work, or project engagements. The two-track format makes those responsibilities and timelines explicit.

01
Entrepreneurship & leadershipContinuous operation · Long-term ownership
Apr 2020 — Present

Wenku Technology

Founder & Chief Executive Officer

Leading technology strategy, product direction, and core architecture for AI-native enterprise solutions. Deeply embedded with the RTE (Release Train Engineer) team, partnering with product and delivery teams to introduce AI Agents across requirements analysis, architecture, coding, testing, and operations, and to establish an AI Native development model. The technology portfolio spans LLM fine-tuning and high-performance inference, RAG, MCP, agentic long-term memory and evaluation, computer vision, OCR, and cloud-native distributed systems. Driving platformization, engineering standards, observability, and governance to turn project knowledge into reusable organizational intelligence, while remaining hands-on in critical architecture, prototyping, and production implementation.

Technology StrategyAI ProductSolution ArchitectureTeam Leadership
02
Concurrent professional practiceEnterprise roles · Projects and industry depth
From Aug 2022 · Concurrent

NTT DATA

Deputy Manager / System Architect

Contributed to centralized payment platforms and enterprise architecture, focusing on transaction security, availability, scalability, and cross-team delivery.

System ArchitecturePaymentHigh Availability
From Apr 2021 · Concurrent

Sense AI

Data Scientist

Built data-platform and analytics capabilities that turned fragmented information into reusable data assets and decision insight.

Data PlatformMachine LearningAnalytics
From Nov 2020 · Concurrent

Accenture

Senior Analyst

Supported enterprise data and digital initiatives through business analysis, data modeling, and system delivery for large-scale transformation.

ConsultingData EngineeringTransformation
Education

Academic background

May 2022

Arizona State University

Computing and Technology

Advanced study in scalable computing, enterprise cloud applications, big data, machine learning, and modern system architecture.

Apr 2018

Tokyo Institute of Technology

Natural Language Processing

Focused on language understanding, text analytics, machine learning models, and applied NLP engineering.

Sep 2013

Soochow University

Computer Science

Built a rigorous foundation in computer science, operating systems, networking, and computer architecture.

Credentials

Professional certifications

Jun 2025

Big Data Professional Certification

Dec 2024

AI and Machine Learning Professional Certification

Oct 2024

AWS Certified Solutions Architect – Professional

Dec 2020

Oracle Certified Java Programmer Gold SE 8

Aug 2020

Google Data Analytics Certification

Feb 2016

Japanese Language Proficiency Test N1

Build What Matters

Bring AI into real business systems

If you are planning an enterprise AI program, inference platform, or complex system modernization, let’s discuss the technology path and delivery strategy.

Start a conversation
Zhewen Li | CEO & AI Systems Architect | Wenku Technology