0.00.139.517 I cmn common_param: common_params_print_info: verbosity = 3 (adjust with the `-lv N` CLI arg) 0.00.251.982 I srv load_model: loading model '/models/gguf/Mistral-Large-3/UD-IQ1_S/Mistral-Large-3-675B-Instruct-2512-UD-IQ1_S-00001-of-00004.gguf' 0.00.744.979 W llama_model_loader: tensor overrides to CPU are used with mmap enabled - consider using --no-mmap for better performance 4.04.906.271 W llama_context: setting new yarn_attn_factor = 1.0000 (mscale == 1.0, mscale_all_dim = 1.0) 4.06.456.616 I srv load_model: initializing, n_slots = 1, n_ctx_slot = 32768, kv_unified = 'false' 4.06.465.102 I srv llama_server: model loaded 4.06.465.109 I srv llama_server: listening on http://:8104 4.14.689.001 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1 4.14.689.020 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0 6.49.521.033 I slot print_timing: id 0 | task 0 | prompt eval time = 149099.07 ms / 547 tokens ( 272.58 ms per token, 3.67 tokens per second) 6.49.521.036 I slot print_timing: id 0 | task 0 | eval time = 5732.09 ms / 36 tokens ( 159.22 ms per token, 6.28 tokens per second) 6.49.521.037 I slot print_timing: id 0 | task 0 | total time = 154831.16 ms / 583 tokens 6.49.521.037 I slot print_timing: id 0 | task 0 | graphs reused = 0 6.49.522.700 I slot release: id 0 | task 0 | stop processing: n_tokens = 582, truncated = 0 7.20.391.675 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, f_sim_best = 0.941 (> 0.100 thold), f_keep = 0.924 7.20.392.608 I slot launch_slot_: id 0 | task 37 | processing task, is_child = 0 7.57.239.419 I slot print_timing: id 0 | task 37 | n_decoded = 100, tg = 5.35 t/s, tg_3s = 5.35 t/s 8.00.445.367 I slot print_timing: id 0 | task 37 | n_decoded = 120, tg = 5.48 t/s, tg_3s = 6.24 t/s 8.03.460.567 I slot print_timing: id 0 | task 37 | n_decoded = 139, tg = 5.58 t/s, tg_3s = 6.30 t/s 8.06.586.466 I slot print_timing: id 0 | task 37 | n_decoded = 159, tg = 5.67 t/s, tg_3s = 6.40 t/s 8.09.641.949 I slot print_timing: id 0 | task 37 | n_decoded = 179, tg = 5.76 t/s, tg_3s = 6.55 t/s 8.12.768.440 I slot print_timing: id 0 | task 37 | n_decoded = 199, tg = 5.82 t/s, tg_3s = 6.40 t/s 8.15.802.869 I slot print_timing: id 0 | task 37 | n_decoded = 219, tg = 5.88 t/s, tg_3s = 6.59 t/s 8.18.837.255 I slot print_timing: id 0 | task 37 | n_decoded = 239, tg = 5.93 t/s, tg_3s = 6.59 t/s 8.21.844.167 I slot print_timing: id 0 | task 37 | n_decoded = 259, tg = 5.98 t/s, tg_3s = 6.65 t/s 8.22.002.500 I slot print_timing: id 0 | task 37 | prompt eval time = 18164.93 ms / 34 tokens ( 534.26 ms per token, 1.87 tokens per second) 8.22.002.503 I slot print_timing: id 0 | task 37 | eval time = 43444.95 ms / 260 tokens ( 167.10 ms per token, 5.98 tokens per second) 8.22.002.503 I slot print_timing: id 0 | task 37 | total time = 61609.87 ms / 294 tokens 8.22.002.504 I slot print_timing: id 0 | task 37 | graphs reused = 0 8.22.003.466 I slot release: id 0 | task 37 | stop processing: n_tokens = 831, truncated = 0 8.48.135.289 I srv operator(): operator(): cleaning up before exit...