← 返回教程目录

IBM Granite

一句话简介:IBM Granite 4.2(2026 年 8 月 25 日发布)。

主流功能:企业治理、合规场景

可用性:可用性:国内可直接使用。

适用人群:想本地跑一个 Apache 2.0 可商用、带原生推理与工具调用的 Agent 模型的开发者

前置条件:

  • 网络:能访问 Ollama 模型库与 hf-mirror.com
  • 账号:本地部署无需账号
  • 费用:本地免费
  • 硬件:3B Q4:约 2.2GB(笔记本可跑)

分步骤教程

方式一:Ollama 本地部署(推荐)

步骤 1:安装 Ollama

  • Windows / macOS:官网下载安装包
  • Linux:
curl -fsSL https://ollama.com/install.sh | sh

步骤 2:拉取 Granite 4.2

ollama pull granite4:3b # 轻量,笔记本/边缘设备
# ollama pull granite4:8b # 均衡,单卡 Agent 推荐
# ollama pull granite4:30b # 旗舰,需 24GB+ 显存

Ollama 官方库模型名为 granite4(Granite 4.x 系列,含 4.2 推理能力);ollama.com/library/granite4 可查档位标签。若拉取后 ollama show granite4 显示非 4.2,以当时库内最新版本为准。

步骤 3:运行对话

ollama run granite4:3b

>>> 出现后直接对话。退出按 Ctrl+d;查看已装模型用 ollama list。

步骤 4(可选):OpenAI 兼容接口供 Agent 框架调用

Ollama 自带 OpenAI 兼容接口(http://localhost:11434/v1),Python 调用:

from openai import OpenAI
client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
resp = client.chat.completions.create(
model="granite4:8b",
messages=[{"role": "user", "content": "用 Python 写一个快速排序"}],
temperature=1.0, top_p=0.95) # IBM 官方建议全场景用 temperature=1.0 / top_p=0.95
print(resp.choices[0].message.content)

方式二:HuggingFace transformers 部署(官方示例)

pip install -i https://pypi.tuna.tsinghua.edu.cn/simple torch transformers accelerate
export HF_ENDPOINT=https://hf-mirror.com
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_path = "ibm-granite/granite-4.2-8b" # 或 granite-4.2-3b / granite-4.2-30b
tokenizer = AutoTokenizer.from_pretrained(model_path)
model = AutoModelForCausalLM.from_pretrained(
model_path, device_map="cuda", torch_dtype=torch.bfloat16)
model.eval()

messages = [{"role": "user", "content": "How many r's are in the word 'strawberry'?"}]
# enable_thinking=True 开启思考模式;False 则直接回答;再加 low_effort=True 为轻量思考
text = tokenizer.apply_chat_template(messages, tokenize=False,
add_generation_prompt=True, enable_thinking=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=8192,
temperature=1.0, top_p=0.95, do_sample=True)
print(tokenizer.decode(output[0][inputs.input_ids.shape[-1]:], skip_special_tokens=False))

输出中 <think>...</think> 即模型的思考过程。

方式三:vLLM 高性能服务(多卡/生产,未验证细节)

vllm serve ibm-granite/granite-4.2-8b \
--dtype bfloat16 \
--max-model-len 131072 \
--reasoning-parser granite_thinking_parser \
--reasoning-parser-plugin./granite_thinking_parser.py \
--tool-call-parser qwen3_coder \
--enable-auto-tool-choice

服务在 http://localhost:8000/v1 暴露 OpenAI 兼容接口,可直接对接 OpenHands / OpenCode 等 Agent 框架(参数据 IBM 官方文档,vLLM 需 v0.20+)。

方式四:IBM watsonx API(企业云端调用,未验证)

  1. 注册 IBM Cloud 账号,开通 watsonx.ai 服务
  2. 在 watsonx 控制台创建 Project,获取 API Key 与 Project ID
  3. 按 watsonx 官方文档的调用示例(与 OpenAI 格式类似,需传 Project ID)调用 Granite 4.2 模型

国内访问 IBM Cloud 与计费情况未验证;已有 IBM 企业协议的团队优先走内部商务通道。

常见问题

  1. thinking 模式怎么开关? transformers 用 enable_thinking=True/False(chat template 参数);Ollama 下按模型模板默认行为,复杂任务它会自动先思考再答。

  2. 3B / 8B / 30B 怎么选? 关键区别:8B 和 30B 经过了 agentic RL(真实沙盒里练代码编辑、终端操作、联网搜索),3B 没有。做 Agent/工具调用优先 8B;纯问答、边缘设备用 3B。

  3. Ollama 里 granite4 和 granite3.3 是什么关系? granite4 是 4.x 新一代(含 4.2 推理能力),granite3.x 是上一代指令模型。新项目直接用 granite4。

  4. GGUF 量化版哪里找? 官方发布了低至 Q4_K_M 的 GGUF 量化,Ollama 拉取即自动处理;手动下载可去 HF ibm-granite 组织下各模型的 GGUF 文件(经 hf-mirror.com 镜像)。

  5. 中文支持怎么样? 官方称在中、英、德、法、日、韩等 12 种语言上测试过多语言对话,中文日常问答可用;但其中文综合能力弱于国产主力模型,中文主力场景建议对比实测。