BH akademski imenik

Jonathan De Brusse, Mathieu Granzotto, R. Postoyan, D. Nešić

1 27. 3. 2024.

Policy iteration for discrete-time systems with discounted costs: stability and near-optimality guarantees

Given a discounted cost, we study deterministic discrete-time systems whose inputs are generated by policy iteration (PI). We provide novel near-optimality and stability properties, while allowing for non-stabilizing initial policies. That is, we first give novel bounds on the mismatch between the value function generated by PI and the optimal value function, which are less conservative in general than those encountered in the dynamic programming literature for the considered class of systems. Then, we show that the systems in closed-loop with policies generated by PI are stabilizing under mild conditions, after a finite (and known) number of iterations.

Preuzmi PDF

Mathematics
Computer Science
+ 1

Pretplatite se na novosti o BH Akademskom Imeniku