Saksham Rai
UC San Diego
Download PDF
http://doi.org/10.37648/ijrst.v15i04.008
Large language model services are increasingly offering heterogeneous portfolios of models, rather than a single universal endpoint. Each model inhabits a different point in a space of cost, latency, and quality, varying with prompt type, output length, system load and provider conditions. This paper formulates adaptive multi-LLM routing as a constrained contextual decision problem. A router estimates request-level quality, monetary cost, and tail latency for candidate models, selects the least costly feasible action, and escalates when confidence or verification signals are insufficient. The proposed architecture unifies an offline performance matrix, calibrated predictors, constrained policy, and online monitoring, and treats abstention, cascades and test-time sampling as first-class actions. Previous systems demonstrate massive savings, but production deployment involves explicit service-level constraints, drift-handling, and risk-sensitive evaluation beyond average accuracy. A practical evaluation protocol based on Pareto frontiers, constraint-violation rates, calibration and tail latency is presented and applied in an offline replay of GSM8K and MMLU with a weak and a strong model. At a quality floor of 90% of the strong model’s accuracy, a simple calibrated router reduces cost by 45% on GSM8K and 69% on MMLU, but thresholds tuned to just meet the floor violate it in about half of the evaluation splits; a conservative threshold rule lowers the violation rate to at most 1% while retaining savings of 29% and 62%, respectively. The analysis suggests that robust routing gains are derived less from a universally superior router, than from accurate uncertainty estimates, diverse candidate models, and disciplined feedback loops.
Keywords: large language models; model routing; constrained optimization; latency; inference cost; quality assurance; contextual bandits
Disclaimer: Indexing of published papers is subject to the evaluation and acceptance criteria of the respective indexing agencies. While we strive to maintain high academic and editorial standards, International Journal of Research in Science and Technology does not guarantee the indexing of any published paper. Acceptance and inclusion in indexing databases are determined by the quality, originality, and relevance of the paper, and are at the sole discretion of the indexing bodies.