Canopy Wave & SAIHEACanopy Wave & SAIHEAT Announce Merger Agreement | Learn MoreArrow
GLM-5.3-Flash API
VISIONCODELLM

GLM-5.3-Flash API

All You Need to Know About GLM-5.3-Flash API

Overview

Model Provider:Zai-org
Model Type:VISION/CODE/LLM
State:Ready

Key Specs

Quantization:FP8
Parameters:321B
Context:1M
Pricing:$0.15 input / $0.50 output / $0.03 cache
Try Model API
Quick Start
Reserve Dedicated Endpoint

Introduction

GLM-5.3-Flash is the first native multimodal model in the GLM-5 series, delivering stronger intelligence than GLM-5.2 while maintaining an exceptionally cost-efficient architecture.

  • Highly Efficient Hybrid Architecture: GLM-5.3-Flash has 320B total parameters, with 18B activated parameters. It is the first open-source frontier model to adopt a hybrid architecture combining sparse attention and linear attention. This architecture significantly reduces computational and serving costs while maintaining precise long-context capabilities. Compared with GLM-5.3, it reduces attention computation and KV cache size by 3.01× and 4.44×, respectively.
  • Native Highly Efficient Hybrid Archi Visual Coding: Visual capabilities are natively integrated into the coding loop, enabling the model to actively observe interfaces, rendered results, and interaction feedback, and continuously test and improve its work accordingly. From frontend development and game creation to Blender 3D scenes and real-world environment operation driven by BUA and CUA, the model can seamlessly coordinate tasks across code, browsers, and graphical user interfaces.
  • A Professional Work Partner Beyond Coding: GLM-5.3-Flash further extends its capabilities to a wide range of professional workflows, including Office tasks, financial research, and professional document processing. It can autonomously break down complex objectives, invoke the appropriate tools, and review and optimize its outputs, completing end-to-end workflows from research and analysis and model building to delivering finished PPTX, PDF, DOCX, and XLSX files.

GLM-5.3-Flash API Usage

Model

Endpoint

zai/glm-5.3-flash


        1
        curl -X POST https://inference.canopywave.io/v1/chat/completions \
      
        2
          -H "Content-Type: application/json" \
      
        3
          -H "Authorization: Bearer $CANOPYWAVE_API_KEY" \
      
        4
          -d '{
      
        5
            "model": "zai/glm-5.3-flash",
      
        6
            "messages": [
      
        7
              {"role": "user", "content": "tell me a story"}
      
        8
            ],
      
        9
            "max_tokens": 1000,
      
        10
            "temperature": 0.7
      
        11
          }'