ollama

mirror of https://github.com/ollama/ollama.git synced 2026-03-27 02:58:43 +07:00

Author	SHA1	Message	Date
Shivam Tiwari	f3f31a8192	anthropic: close thinking block before tool_use when no text in between (#14825 ) Root cause: StreamConverter.Process() only incremented contentIndex when closing a thinking block if text content was present. When a model emitted thinking followed directly by a tool_use block (no text in between), thinkingDone was never set and contentIndex was not incremented, causing the tool_use content_block_start to reuse index 0. Clients expecting sequential indices would then fail to find the tool content block. Fix: In the tool call loop, close any open thinking block (thinkingStarted && !thinkingDone) and increment contentIndex before opening the tool_use block, mirroring the existing logic that closes an open text block. Fixes #14816 v0.18.0-rc1	2026-03-13 13:12:05 -07:00
Devon Rifkin	9e7ba835da	cmd: still populate `ollama ls` when using `ollama run <model:cloud>` (#14824 ) This is temporary until `api/tags` supports cloud natively v0.18.0-rc0	2026-03-13 12:24:45 -07:00
Parth Sareen	347f17b8d1	launch: add compact window for claude code (#14823 )	2026-03-13 12:09:23 -07:00
Devon Rifkin	081b9eb423	api/create: always propagate `:cloud` source for cloud models (#14822 ) Otherwise, using `/save` would try to run the local model instead	2026-03-13 11:58:00 -07:00
Parth Sareen	bb867c6fdb	launch: fix headless --yes integration flow and policy scoping (#14815 )	2026-03-13 11:45:36 -07:00
Cadu	81f4506a61	docs: document reasoning_effort support in OpenAI-compatible API (#14821 ) Add reasoning_effort and reasoning to the supported features and request fields for /v1/chat/completions. These fields control thinking on thinking-capable models but were previously undocumented. Closes #14820	2026-03-13 10:57:14 -07:00
Parth Sareen	76925f1284	cmd: TUI model ordering (#14814 )	2026-03-13 10:19:22 -07:00
Devon Rifkin	f676231de9	server: remove experimental aliases support (#14810 ) v0.17.8-rc4	2026-03-12 20:27:24 -07:00
Parth Sareen	af5f7c0a9e	cmd: refactor tui and launch (#14609 )	2026-03-12 18:39:06 -07:00
Daniel Hiltgen	a6b27d776b	ci: fix missing windows zip file (#14807 ) Use 7z compression (better compression rate) if found in path. That alone isn't sufficient to get us under 2G, so MLX is now split out as a discrete download. Fix CI so it will fail if artifacts fail to upload. v0.17.8-rc3	2026-03-12 16:14:00 -07:00
Daniel Hiltgen	539741199e	mlx: perf improvements (#14768 ) * mlx: perf improvements Fix nn.go to call mlx_fast_layer_norm instead of manually implementing (mean, subtract, variance, rsqrt, multiply, add — 6 ops) Fix llama.go, gemma3.go to remove RepeatKV to tile K/V tensors to match the Q head count, since scaled_dot_product_attention natively handles GQA (it just requires n_q_heads % n_kv_heads == 0) * review comments v0.17.8-rc2	2026-03-12 12:01:28 -07:00
Eva H	8f45236d09	middleware: enable local tool model for web search (#14787 )	2026-03-11 17:51:39 -04:00
Parth Sareen	97013a190c	openai: split mixed thinking stream chunks via ToChunks (#14648 )	2026-03-11 14:21:29 -07:00
Daniel Hiltgen	c222735c02	mlx: only log load errors when MLX is needed (#14764 ) This suppresses irrelevant/noisy errors in the GGML runner.	2026-03-11 10:31:31 -07:00
Daniel Hiltgen	87d21c7fc0	MLX: harden for init failures (#14777 ) The CLI now links to the lazy-load MLX code, but that still happens in init functions. On internal MLX errors, the CLI exits before it has a chance to start. This change re-wires the MLX error handling so it doesn't exit by default. The MLX based runners currently expect exits on failure, so they re-initialize the default error handling. We can refine error handling for better go stack traces in the future.	2026-03-10 22:52:23 -07:00
Jeffrey Morgan	54e05172a0	Revert "runner: add token history sampling parameters to ollama runner (#14537 )" (#14776 ) This reverts commit `86513cb697`.	2026-03-10 21:07:52 -07:00
Parth Sareen	464186e995	config: qwen3.5 recommendations (#14758 )	2026-03-10 18:04:57 -07:00
Devon Rifkin	8c4d5d6c2f	cloud_proxy: send ollama client version (#14769 ) This was previously included in the user agent, and we've made use of it in the past to hotpatch bugs server-side for particular Ollama versions.	2026-03-10 15:53:25 -07:00
Parth Sareen	bc72b14016	docs: update claude code docs (#14770 )	2026-03-10 15:52:41 -07:00
Parth Sareen	61086083eb	server: add experimental web search and web fetch routes (#14753 )	2026-03-09 21:52:12 -07:00
Daniel Hiltgen	62d1f01ab4	ci: Fix windows build (#14754 ) Instead of relying on sh for wildcard, do it in Go for better windows compatibility. v0.17.8-rc1	2026-03-09 19:27:59 -07:00
Daniel Hiltgen	10e51c5177	MLX: add header vendoring and remove go build tag (#14642 ) * prefer rocm v6 on windows Avoid building with v7 - more changes are needed * MLX: add header vendoring and remove go build tag This switches to using a vendoring approach for the mlx-c headers so that Go can build without requiring a cmake first. This enables building the new MLX based code by default. Every time cmake runs, the headers are refreshed, so we can easily keep them in sync when we bump mlx versions. Basic Windows and Linux support are verified. * ci: harden for flaky choco repo servers CI sometimes fails due to choco not actually installing cache. Since it just speeds up the build, we can proceed without. * review comments v0.17.8-rc0	2026-03-09 17:24:45 -07:00
Patrick Devine	3e06bde643	mlx: get parameters from modelfile during model creation (#14747 )	2026-03-09 15:33:24 -07:00
Eva H	6be2de8214	app: auto update should be enabled when reset to defaults (#14741 )	2026-03-09 15:02:36 -04:00
Daniel Hiltgen	ebb1b9ec14	rocm: update linux to v7.2 (#14391 ) * rocm: update linux to v7.2 * review comments	2026-03-09 08:26:55 -07:00
Patrick Devine	d126467d5d	x/mlxrunner: replace sampler interface chain with single stateful Sampler (#14652 ) - Collapse MLX sampling state into a single sample.Sampler struct (options + history). - Replace interface-based sampler chain (TopP, TopK, penalty, etc.) with function-based transforms. - Update request/pipeline wiring to use *sample.Sampler, seed history from prompt tokens, and append generated tokens each step. - Implement top_p, min_p, repeat_penalty, and frequency_penalty	2026-03-07 17:50:57 -08:00
Devon Rifkin	afb4c62fbf	cloud_proxy: handle stream disconnects gracefully (#14685 ) Previously we were printing out bad errors for expected cases like clients disconnecting. Now we only debug log when that happens (which still might help in cases where we're figuring out why an integration isn't working). For other errors, we print out a proper warning now	2026-03-06 19:18:52 -08:00
Patrick Devine	e790dc435b	mlx: int4 groupsize 64 (#14682 ) Change affine 4bit integers to use groupsize 64	2026-03-06 16:39:47 -08:00
Daniel Hiltgen	288077c3a3	build: smarter docker parallelism (#14653 ) Our Dockerfile leverages parallel stages for more efficient builds. However, our old parallel settings were naive and lead to under/over utilization depending on the capabilities of your build system. This change switches to using Ninja for all our docker cmake builds to leverage its smarter parallel logic. We tell Ninja to target a load of nproc so each of the build stages will share the load on the system aiming for full CPU use without oversaturation. The GPU parallelism settings are also adjusted to 4 to avoid a long-tail for the last few GPU targets as they work through the long list of GPU architectures. This also fixes the Dockerfile to move Vulkan install to just the stage that needs it instead of blocking most other GPU installs. This should speed up CI which always has a clean build cache.	2026-03-06 16:36:22 -08:00
Daniel Hiltgen	4425c54eda	create: fix localhost handling (#14681 )	2026-03-06 16:35:58 -08:00
Michael Yang	778899a5d2	docs: format compat docs (#14678 )	2026-03-06 14:53:17 -08:00
Jeffrey Morgan	4eab60c1e2	Reapply "don't require pulling stubs for cloud models" again (#14608 ) * Revert "Revert "Reapply "don't require pulling stubs for cloud models"" (#14606)" This reverts commit `39982a954e`. * fix test + do cloud lookup only when seeing cloud models --------- Co-authored-by: ParthSareen <parth.sareen@ollama.com>	2026-03-06 14:27:47 -08:00
Bruce MacDonald	1af850e6e3	parsers: repair unclosed arg_value tags in GLM tool calls (#14656 ) GLM models sometimes omits </arg_value> closing tags in tool call XML, causing xml.Unmarshal to fail with "element <arg_value> closed by </tool_call>". This is a known issue across the GLM family. Sanitize the input to fix closing arg_key values so encoding/xml can handle it.	2026-03-06 14:08:34 -08:00
Parth Sareen	9b0c7cc7b9	cmd: override stale entries for context window pi (#14655 ) v0.17.7-rc2 v0.17.7	2026-03-05 16:30:24 -08:00
Daniel Hiltgen	6928630601	mlx: prevent remote creation mismatch (#14651 ) If the user is pointing at a remote OLLAMA_HOST, fail experimental safetensor based create operations as we only support local creation currently.	2026-03-05 14:59:00 -08:00
Parth Sareen	9896e3627f	cmd/config: fix cloud model limit lookups in integrations (#14650 ) v0.17.7-rc1	2026-03-05 13:57:28 -08:00
Bruce MacDonald	15732f0ea7	cmd: use native Ollama API endpoint for OpenClaw (#14649 ) Remove the /v1 suffix from the OpenClaw provider baseUrl so it uses the native Ollama API instead of the OpenAI-compatible endpoint. The /v1 endpoint my break tool calling in OpenClaw.	2026-03-05 13:29:17 -08:00
Parth Sareen	562c76d7cc	cmd: add qwen3.5 context length for launch (#14626 ) v0.17.7-rc0	2026-03-04 14:10:52 -08:00
Parth Sareen	122c68c151	server: loosen thinking level constraint (#14625 )	2026-03-04 13:42:18 -08:00
Jeffrey Morgan	82848a7806	model: fix renderer and parser for qwen3.5 (#14605 ) v0.17.6	2026-03-03 20:58:29 -08:00
Jeffrey Morgan	39982a954e	Revert "Reapply "don't require pulling stubs for cloud models"" (#14606 ) This reverts commit `799e51d419`.	2026-03-03 20:56:10 -08:00
Patrick Devine	e9f6ea232f	Add qwen3.5-next-moe support to MLX runner and models (#14417 ) This change adds support for qwen3.5-next-moe models (qwen3-next/qwen3.5-next/qwen3-coder) to the MLX runner. It also: * introduces recurrent cache support and related MLX ops * updates pipeline/runner integration and adds tests * properly quantizes stacked expert tensors * a Gated Delta Metal kernel for fast SSM inference * adds new MLX calls for Conv1d, DepthwideConv1d, Contiguous, Exp, Log, SoftmaxAxis	2026-03-03 16:39:22 -08:00
Patrick Devine	110eff01a9	chore: remove old imagegen LLMs models (#14597 ) These models are implemented in the x/mlxrunner instead.	2026-03-03 13:23:40 -08:00
Jeffrey Morgan	799e51d419	Reapply "don't require pulling stubs for cloud models" This reverts commit `97d2f05a6d`.	2026-03-03 13:17:10 -08:00
Victor-Quqi	e8fcb29586	model/renderers: fix glm-ocr image tags in renderer prompts (#14584 )	2026-03-03 12:51:34 -08:00
Jeffrey Morgan	97d2f05a6d	Revert "don't require pulling stubs for cloud models (#14574 )" (#14596 ) This reverts commit `8207e55ec7`.	2026-03-03 12:51:23 -08:00
Devon Rifkin	8207e55ec7	don't require pulling stubs for cloud models (#14574 ) * don't require pulling stubs for cloud models This is a first in a series of PRs that will better integrate Ollama's cloud into the API and CLI. Previously we used to have a layer of indirection where you'd first have to pull a "stub" model that contains a reference to a cloud model. With this change, you don't have to pull first, you can just use a cloud model in various routes like `/api/chat` and `/api/show`. This change respects <https://github.com/ollama/ollama/pull/14221>, so if cloud is disabled, these models won't be accessible. There's also a new, simpler pass-through proxy that doesn't convert the requests ahead of hitting the cloud models, which they themselves already support various formats (e.g., `v1/chat/completions` or Open Responses, etc.). This will help prevent issues caused by double converting (e.g., `v1/chat/completions` converted to `api/chat` on the client, then calling cloud and converting back to a `v1/chat/completions` response instead of the cloud model handling the original `v1/chat/completions` request first). There's now a notion of "source tags", which can be mixed with existing tags. So instead of having different formats like`gpt-oss:20b-cloud` vs. `kimi-k2.5:cloud` (`-cloud` suffix vs. `:cloud`), you can now specify cloud by simply appending `:cloud`. This PR doesn't change model resolution yet, but sets us up to allow for things like omitting the non-source tag, which would make something like `ollama run gpt-oss:cloud` work the same way that `ollama run gpt-oss` already works today. More detailed changes: - Added a shared model selector parser in `types/modelselector`: - supports `:cloud` and `:local` - accepts source tags in any position - supports legacy `:<tag>-cloud` - rejects conflicting source tags - Integrated selector handling across server inference/show routes: - `GenerateHandler`, `ChatHandler`, `EmbedHandler`, `EmbeddingsHandler`, `ShowHandler` - Added explicit-cloud passthrough proxy for ollama.com: - same-endpoint forwarding for `/api/`, `/v1/`, and `/v1/messages` - normalizes `model` (and `name` for `/api/show`) before forwarding - forwards request headers except hop-by-hop/proxy-managed headers - uses bounded response-header timeout - handles auth failures in a friendly way - Preserved cloud-disable behavior (`OLLAMA_NO_CLOUD`) - Updated create flow to support `FROM ...:cloud` model sources (though this flow uses the legacy proxy still, supporting Modelfile overrides is more complicated with the direct proxy approach) - Updated CLI/TUI/config cloud detection to use shared selector logic - Updated CLI preflight behavior so explicit cloud requests do not auto-pull local stubs What's next? - Cloud discovery/listing and cache-backed `ollama ls` / `/api/tags` - Modelfile overlay support for virtual cloud models on OpenAI/Anthropic request families - Recommender/default-selection behavior for ambiguous model families - Fully remove the legacy flow Fixes: https://github.com/ollama/ollama/issues/13801 * consolidate pull logic into confirmAndPull helper pullIfNeeded and ShowOrPull shared identical confirm-and-pull logic. Extract confirmAndPull to eliminate the duplication. * skip local existence checks for cloud models ModelExists and the TUI's modelExists both check the local model list, which causes cloud models to appear missing. Return true early for explicit cloud models so the TUI displays them beside the integration name and skips re-prompting the model picker on relaunch. * support optionally pulling stubs for newly-style names We now normalize names like `<family>:<size>:cloud` into legacy-style names like `<family>:<size>-cloud` for pulling and deleting (this also supports stripping `:local`). Support for pulling cloud models is temporary, once we integrate properly into `/api/tags` we won't need this anymore. * Fix server alias syncing * Update cmd/cmd.go Co-authored-by: Parth Sareen <parth.sareen@ollama.com> * address comments * improve some naming --------- Co-authored-by: ParthSareen <parth.sareen@ollama.com>	2026-03-03 10:46:33 -08:00
Jesse Gross	ad16bffc7d	mlx: Remove peak memory from the API This is still in flux so it is better to just log it for now.	2026-03-02 15:56:18 -08:00
Jesse Gross	c1e3ef4bcc	mlxrunner: Refcount pinned tensors Otherwise, it is error prone to manage multiple components working with the same tensor.	2026-03-02 15:56:06 -08:00
Parth Sareen	a3093cd5e5	cmd/opencode: rename provider from "Ollama (local)" to "Ollama" (#14566 ) The "(local)" qualifier is unnecessary since there's only one Ollama provider. Existing configs with the old name are migrated automatically; custom names are left unchanged.	2026-03-02 14:17:18 -08:00

1 2 3 4 5 ...

5178 Commits