LaTex2Web logo

Documents Live, a web authoring and publishing system

If you see this, something is wrong

Table of contents

First published on Monday, Sep 21, 2026 and last modified on Monday, Sep 21, 2026 by François Chaplais.

Like what you see? Register!
Data-Driven Policy Iteration Without an Initial Stabilizing Policy: A Finite-Horizon Bootstrap Method

Jiacheng Wu College of Control Science and Engineering, Zhejiang University, Hangzhou 310027, China Email

Yang Zhu College of Control Science and Engineering, Zhejiang University, Hangzhou 310027, China Email

Keywords: Policy iteration, data-driven control, off-policy reinforcement learning, optimal control.

Abstract

1 Introduction

2 Preliminaries and problem statement

\[ \begin{aligned} (\mathbf{A}-\mathbf{B}\mathbf{K})^{T}\mathbf{P}_{K} +\mathbf{P}_{K}(\mathbf{A}-\mathbf{B}\mathbf{K}) +\mathbf{Q} +\mathbf{K}^{T}\mathbf{R}\mathbf{K} =\mathbf{0}. \end{aligned} \]
\[ \mathbf{P}_{K} = \int_{0}^{\infty} e^{(\mathbf{A}-\mathbf{B}\mathbf{K})^{T}t} \left(\mathbf{Q}+\mathbf{K}^{T}\mathbf{R}\mathbf{K}\right) e^{(\mathbf{A}-\mathbf{B}\mathbf{K})t} \,\mathrm{d}t. \]
\[ \alpha(\mathbf{A}-\mathbf{B}\mathbf{K}_\rho^\ast) = \alpha(\mathbf{A}_\rho-\mathbf{B}\mathbf{K}_\rho^\ast)-\rho <-\rho<-\varepsilon . \]

3 Model-based finite-horizon bootstrap algorithm


Algorithm 1 Model-Based Finite-Horizon Bootstrap Algorithm
1.Require: \( \varepsilon \geq 0\) , \( \rho > \varepsilon\) , \( \epsilon_{\mathrm{PI}} > 0\) , \( T_0 > 0\) , \( \gamma > 1\) , and \( \mathbf{M}=\mathbf{M}^{T}\geq\mathbf{0}\) .
2.Ensure: An \( \varepsilon\) -admissible gain \( \mathbf{K}^{[0]}\) .
3.Set \( s=0\) and \( T_s=T_0\) . \Loop
4.Choose a bounded policy \( \mathbf{K}_{\rho,T_s}^{[0]}(t)\) on \( [0,T_s]\) .
5.Set \( j=0\) . \Loop
6.Solve (12) on \( [0,T_s]\) .
7.Update \( \mathbf{K}_{\rho,T_s}^{[j+1]}(t)\) according to (13).
8.Set \( j \gets j+1\) .
9.Set \( \mathbf{K}_{c} = \mathbf{K}_{\rho,T_s}^{[j]}(0)\) .
10.if \( \displaystyle \max_{\lambda\in\sigma( \mathbf{A}-\mathbf{B}\mathbf{K}_{c})} \operatorname{Re}(\lambda) < -\varepsilon\) then
11.Set \( \mathbf{K}^{[0]}=\mathbf{K}_{c}\) .
12.return \( \mathbf{K}^{[0]}\) .
13.end if
14.if \( \displaystyle \left\| \mathbf{K}_{\rho,T_s}^{[j]} - \mathbf{K}_{\rho,T_s}^{[j-1]} \right\|_{\infty,T_s} \leq \epsilon_{\mathrm{PI}}\) then
15.break.
16.end if
17.Set \( T_{s+1}=\gamma T_s\) and \( s \gets s+1\) . \EndLoop

4 Data-driven finite-horizon bootstrap algorithm

\[ \begin{align*} \boldsymbol{\omega}_{P,k,\ell} ={}& \phi_{\ell}(t_{k+1}) \mathrm{vecv}^{T}\bigl(x(t_{k+1})\bigr) -\phi_{\ell}(t_{k}) \mathrm{vecv}^{T}\bigl(x(t_{k})\bigr) \\ &+2\rho\int_{t_{k}}^{t_{k+1}} \phi_{\ell}(t)\mathrm{vecv}^{T}\bigl(x(t)\bigr) \,\mathrm{d}t, \\ \boldsymbol{\omega}_{K,k,\ell}^{[j]} ={}& -2\int_{t_{k}}^{t_{k+1}} \psi_{\ell}(t) \left[ x^{T}(t)\otimes \bigl(\mathbf{z}^{[j]}(t)\bigr)^{T}\mathbf{R} \right] \,\mathrm{d}t, \\ b_{k}^{[j]} ={}& -\int_{t_{k}}^{t_{k+1}} x^{T}(t)\mathbf{Q}_{\rho,T}^{[j]}(t)x(t) \,\mathrm{d}t \\ &-\left[ x^{T}(t_{k+1})\mathbf{M}x(t_{k+1}) -x^{T}(t_{k})\mathbf{M}x(t_{k}) \right] \\ &-2\rho\int_{t_{k}}^{t_{k+1}} x^{T}(t)\mathbf{M}x(t) \,\mathrm{d}t, \\ r_{k}^{[j]} ={}& -\left[ x^{T}(t_{k+1}) \Delta_{P}^{[j]}(t_{k+1}) x(t_{k+1}) \right. \\ &\left.~~ -x^{T}(t_{k}) \Delta_{P}^{[j]}(t_{k}) x(t_{k}) \right] \\ &-2\rho\int_{t_{k}}^{t_{k+1}} x^{T}(t)\Delta_{P}^{[j]}(t)x(t) \,\mathrm{d}t \\ &+2\int_{t_{k}}^{t_{k+1}} \bigl(\mathbf{z}^{[j]}(t)\bigr)^{T} \mathbf{R}\Delta_{K}^{[j+1]}(t)x(t) \,\mathrm{d}t. \end{align*} \]
\[ \begin{align*} \boldsymbol{\omega}_{k}^{[j]} ={}& \bigl[ \boldsymbol{\omega}_{P,k,1},\ldots, \boldsymbol{\omega}_{P,k,N_{p}}, \\ &~~ \boldsymbol{\omega}_{K,k,1}^{[j]},\ldots, \boldsymbol{\omega}_{K,k,N_{k}}^{[j]} \bigr], \\ \theta^{[j]} ={}& \operatorname{col}\bigl( \operatorname{vecs}(\mathbf{P}_{1}^{[j]}),\ldots, \operatorname{vecs}(\mathbf{P}_{N_{p}}^{[j]}), \\ &~~ \operatorname{vec}(\mathbf{K}_{1}^{[j+1]}),\ldots, \operatorname{vec}(\mathbf{K}_{N_{k}}^{[j+1]}) \bigr). \end{align*} \]
\[ \begin{align*} \boldsymbol{\Omega}_{\rho,T}^{[j]} &= \operatorname{col}\left( \boldsymbol{\omega}_{1}^{[j]},\ldots, \boldsymbol{\omega}_{N_{d}}^{[j]} \right),\\ \mathbf{b}_{\rho,T}^{[j]} &= \operatorname{col}\left( b_{1}^{[j]},\ldots,b_{N_{d}}^{[j]} \right),\\ \mathbf{r}_{\rho,T}^{[j]} &= \operatorname{col}\left( r_{1}^{[j]},\ldots,r_{N_{d}}^{[j]} \right), \end{align*} \]
\[ \begin{align*} \mathbf{d}_{xx,k} &= \operatorname{vecv}\bigl(x(t_{k+1})\bigr) -\operatorname{vecv}\bigl(x(t_{k})\bigr),\\ \mathbf{I}_{xx,k} &= \int_{t_{k}}^{t_{k+1}} \operatorname{vecv}\bigl(x(t)\bigr)\,\mathrm{d}t,\\ \mathbf{I}_{xz,k}(\mathbf{K}_{c}) &= \int_{t_{k}}^{t_{k+1}} \bigl(x(t)\otimes z_{c}(t)\bigr)\,\mathrm{d}t. \end{align*} \]
\[ \begin{align*} \theta_{c} &= \operatorname{col}\bigl( \operatorname{vecs}(\mathbf{P}_{c}), \operatorname{vec}(\mathbf{L}_{c}) \bigr),\\ \boldsymbol{\Psi}_{c}(\mathbf{K}_{c}) &= \bigl[ \mathbf{D}_{xx}+2\varepsilon\mathbf{I}_{xx}, \;-2\mathbf{I}_{xz}(\mathbf{K}_{c}) \bigr],\\ \mathbf{b}_{c}(\mathbf{W}_{c}) &= \mathbf{I}_{xx} \operatorname{vecs}(\mathbf{W}_{c}), \end{align*} \]
\[ \begin{align*} \mathbf{D}_{xx} &= \rm col \bigl( \mathbf{d}_{xx,1}^{T},\ldots, \mathbf{d}_{xx,N_{d}}^{T} \bigr),\\ \mathbf{I}_{xx} &= \rm col \bigl( \mathbf{I}_{xx,1}^{T},\ldots, \mathbf{I}_{xx,N_{d}}^{T} \bigr),\\ \mathbf{I}_{xz}(\mathbf{K}_{c}) &= \rm col \bigl( \mathbf{I}_{xz,1}^{T}(\mathbf{K}_{c}),\ldots, \mathbf{I}_{xz,N_{d}}^{T}(\mathbf{K}_{c}) \bigr). \end{align*} \]

Algorithm 2 Data-Driven Finite-Horizon Bootstrap Algorithm
1.Require: \( \varepsilon \geq 0\) , \( \rho > \varepsilon\) , \( \epsilon_{\mathrm{PI}} > 0\) , \( T_0 > 0\) , \( \gamma > 1\) , \( \mathbf{M}=\mathbf{M}^{T}\geq\mathbf{0}\) , and \( \mathbf{W}_c=\mathbf{W}_c^{T}>\mathbf{0}\) .
2.Ensure: A certified \( \varepsilon\) -admissible gain \( \mathbf{K}^{[0]}\) .
3.Set \( s=0\) and \( T_s=T_0\) . \Loop
4.Collect data on \( [0,T_s]\) .
5.Choose a bounded \( \widehat{\mathbf{K}}_{\rho,T_s}^{[0]}(t)\) on \( [0,T_s]\) .
6.repeat
7.Construct \( \boldsymbol{\Omega}_{\rho,T_s}^{[j]}\) and \( \mathbf{b}_{\rho,T_s}^{[j]}\) from (30).
8.Solve \( \widehat{\boldsymbol{\theta}}^{[j]} = \left( \boldsymbol{\Omega}_{\rho,T_s}^{[j]} \right)^{\dagger} \mathbf{b}_{\rho,T_s}^{[j]}\) .
9.Set \( j \gets j+1\) .
10.until \( \left\| \widehat{\mathbf{K}}_{\rho,T_s}^{[j]} - \widehat{\mathbf{K}}_{\rho,T_s}^{[j-1]} \right\|_{\infty,T_s} \leq \epsilon_{\mathrm{PI}}\)
11.Set \( \bar{j}_s=j\) and \( \mathbf{K}_c = \widehat{\mathbf{K}}_{\rho,T_s}^{[\bar{j}_s]}(0)\) .
12.if (60) admits a solution with \( \mathbf{P}_c>\mathbf{0}\) then
13.Set \( \mathbf{K}^{[0]}\gets\mathbf{K}_c\) .
14.return \( \mathbf{K}^{[0]}\) .
15.end if
16.Set \( T_{s+1}=\gamma T_s\) and \( s\gets s+1\) . \EndLoop

5 Illustrative examples

\[ \mathbf{A} = \begin{bmatrix} 1.3800 & -0.2077 & 6.7150 & -5.6760 \\ -0.5814 & -4.2900 & 0 & 0.6750 \\ 1.0670 & 4.2730 & -6.6540 & 5.8930 \\ 0.0480 & 4.2730 & 1.3430 & -2.1040 \end{bmatrix}, \]
\[ \mathbf{B} = \begin{bmatrix} 0 & 0 \\ 5.6790 & 0 \\ 1.1360 & -3.1460 \\ 1.1360 & 0 \end{bmatrix}. \]
\[ \mathbf{K}_{c} = \begin{bmatrix} 0.1086 & 0.5935 & 0.1185 & 0.0401 \\ -1.0155 & -0.0248 & -0.6106 & 0.4062 \end{bmatrix}. \]
\[ \begin{align*} \phi_{\ell}(t) &= \left(1-\frac{t}{T_{s}}\right) \left(\frac{t}{T_{s}}\right)^{\ell-1},\\ \psi_{\ell}(t) &= \left(\frac{t}{T_{s}}\right)^{\ell-1}, ~~ \ell=1,\ldots,8. \end{align*} \]
\[ \mathbf{K}_{c}= \begin{bmatrix} 0.1085 & 0.5936 & 0.1184 & 0.0402 \\ -1.0155 & -0.0247 & -0.6106 & 0.4062\end{bmatrix}. \]
\[ \mathbf{A}=\begin{bmatrix} 0 & 1 & 0 & 0 \\ -\frac{k_{1}+k_{2}}{m_{1}} & 0 & \frac{k_{2}}{m_{1}} & 0 \\ 0 & 0 & 0 & 1 \\ \frac{k_{2}}{m_{2}} & 0 & -\frac{k_{2}}{m_{2}} & 0\end{bmatrix},~ \mathbf{B}=\begin{bmatrix} 0 \\ \frac{1}{m_{1}} \\ 0 \\ 0\end{bmatrix}. \]
\[ \mathbf{K}_{h}^{[101]} = \begin{bmatrix} 14.7218 & 6.5855 & -2.4442 & 16.4721 \end{bmatrix}, \]
\[ \mathbf{K}_{v}^{[193]} =\begin{bmatrix} 0.7714 & 1.5364 & -0.1604 & 0.9992 \end{bmatrix}, \]
\[ \mathbf{K}_{c} = \begin{bmatrix} -0.0057 & 0.0911 & 0.0041 & 0.0003 \end{bmatrix}. \]

6 Conclusion

References

[1] Frank L Lewis and Draguna Vrabie and Vassilis L Syrmos Optimal Control John Wiley & Sons 2012

[2] David G Hull Optimal Control Theory for Applications Cham, Switzerland: Springer, 2003.

[3] Tong Liu and Miroslav Krstić and Zhong-Ping Jiang Adaptive dynamic programming-regulated extremum seeking for distributed feedback optimization iElectron. Lettėctr. Engėac 2025 70 11 7675–7682 Nov.

[4] David Kleinman On an iterative technique for Riccati equation computations iElectron. Lettėctr. Engėac 1968 13 1 114–115 Feb.

[5] Victor G Lopez and Mohammad Alsalti and Matthias A Müller Efficient off-policy Q-learning for data-based discrete-time LQR problems iElectron. Lettėctr. Engėac 2023 68 5 2922–2933 May

[6] Leilei Cui and Bo Pang and Miroslav Krstić and Zhong-Ping Jiang Learning-based adaptive optimal control of linear time-delay systems: A value iteration approach AutoMechatronicsatica 2025 171 111944

[7] Draguna Vrabie and Octavian Pastravanu and Murad Abu-Khalaf and Frank L Lewis Adaptive optimal control for continuous-time linear systems based on policy iteration AutoMechatronicsatica 2009 45 2 477–484

[8] Yu Jiang and Zhong-Ping Jiang Computational adaptive optimal control for continuous-time linear systems with completely unknown dynamics AutoMechatronicsatica 2012 48 10 2699–2704

[9] Hao Shen and Chuanjun Peng and Huaicheng Yan and Shengyuan Xu Data-driven near optimization for fast sampling singularly perturbed systems iElectron. Lettėctr. Engėac 2024 69 7 4689–4694 Jul.

[10] Biao Luo and Derong Liu and Huai-Ning Wu and Tingwen Huang and Chunhua Yang and Weihua Gui Recent advances on off-policy reinforcement learning for optimization control iElectron. Lettėctr. Engėc in press, DOI: 10.1109/TCYB.2026.3683384

[11] Syed Ali Asad Rizvi and Zongli Lin Output feedback Q-learning for discrete-time linear zero-sum games with application to the H-infinity control AutoMechatronicsatica 2018 95 213–221

[12] Weinan Gao and Zhong-Ping Jiang Adaptive dynamic programming and adaptive optimal output regulation of linear systems iElectron. Lettėctr. Engėac 2016 61 12 4164–4169 Dec.

[13] Bosen Lian and Vrushabh S Donge and Frank L Lewis and Tianyou Chai and Ali Davoudi Data-driven inverse reinforcement learning control for linear multiplayer games iElectron. Lettėctr. EngėNeural Netw.ls 2024 35 2 2028–2041 Feb.

[14] Yi Jiang and Tao Yang and Weinan Gao and Jin Wu and Tianyou Chai and Frank L Lewis Off-Policy Reinforcement Learning for \( H_{\infty}\) Control of Linear Discrete-Time Systems with Network Induced Dropouts iElectron. Lettėctr. Engėac 2025 70 12 8000–8015 Dec.

[15] Corrado Possieri and Mario Sassano Value iteration for continuous-time linear time-invariant systems iElectron. Lettėctr. Engėac 2023 68 5 3070–3077 May

[16] Qinglai Wei and Derong Liu and Hanquan Lin Value iteration adaptive dynamic programming for optimal control of discrete-time nonlinear systems iElectron. Lettėctr. Engėc 2016 46 3 840–853 Mar.

[17] Biao Luo and Yin Yang and Huai-Ning Wu and Tingwen Huang Balancing value iteration and policy iteration for discrete-time control iElectron. Lettėctr. EngėsMechatronicSoft Computṡ 2020 50 11 3948–3958 Nov.

[18] Tao Bian and Zhong-Ping Jiang Value iteration and adaptive dynamic programming for data-driven adaptive optimal control design AutoMechatronicsatica 2016 71 348–360

[19] Bo Pang and Zhong-Ping Jiang Adaptive optimal control of linear periodic systems: An off-policy value iteration approach iElectron. Lettėctr. Engėac 2021 66 2 888–894 Feb.

[20] Weinan Gao and Chao Deng and Yi Jiang and Zhong-Ping Jiang Resilient reinforcement learning and robust output regulation under denial-of-service attacks AutoMechatronicsatica 2022 142 110366

[21] Jiacheng Wu and Bosen Lian and Changyun Wen and Yang Zhu Distributed FilterNet reinforcement learning for achieving output consensus in heterogeneous multiplayer multiagent systems iElectron. Lettėctr. EngėNeural Netw.ls 2026 37 2 575–588 Feb.

[22] Omar Qasem and Hector Gutierrez and Weinan Gao Experimental validation of data-driven adaptive optimal control for continuous-time systems via hybrid iteration: An application to rotary inverted pendulum iElectron. Lettėctr. Engėie 2024 71 6 6210–6220 Jun.

[23] Hao Shen and Yun Wang and Jiacheng Wu and Ju H Park and Jing Wang Secure control for Markov jump cyber-physical systems subject to malicious attacks: A resilient hybrid learning scheme iElectron. Lettėctr. Engėc 2024 54 11 7068–7079 Nov.

[24] Xin Wang and Linghuan Kong and Quanxin Zhu and Ben Niu Secure optimal control of Itô stochastic Markov jump systems subject to DoS attacks: A hybrid learning algorithm AutoMechatronicsatica 2026 183 112681

[25] Shuo-Qiu Zhang and Wei-Wei Che and Zheng-Guang Wu Prescribed-time observer-based HI-RL secure output tracking control for heterogeneous MASs under DoS attacks iElectron. Lettėctr. EngėsMechatronicSoft Computṡ 2026 56 1 709–723 Jan.

[26] Xue Liang and Weinan Gao and Chuan Hu and Tianyou Chai Cooperative adaptive cruise control of connected and autonomous vehicles via hybrid iteration iElectron. Lettėctr. Engėvt 2026 75 6 8793–8804 Jun.

[27] Huaiyuan Jiang and Bin Zhou Bias-policy iteration based adaptive dynamic programming for unknown continuous-time linear systems AutoMechatronicsatica 2022 136 110058

[28] Ci Chen and Frank L Lewis and Bo Li Homotopic policy iteration-based learning design for unknown linear continuous-time systems AutoMechatronicsatica 2022 138 110153

[29] Wenwu Fan and Junlin Xiong A homotopy method for continuous-time model-free LQR control based on policy iteration iElectron. Lettėctr. Engėcaa 2025 12 8 1673–1682 Aug.

[30] Yong-Sheng Ma and Jian Sun and Yong Xu and Shi-Sheng Cui and Zheng-Guang Wu Adaptive dynamic programming for optimal control of unknown LTI system via interval excitation iElectron. Lettėctr. Engėac 2025 70 7 4896–4903 Jul.

[31] Yongliang Yang and Bahare Kiumarsi and Hamidreza Modares and Chengzhong Xu Model-free \( \lambda\) -policy iteration for discrete-time linear quadratic regulation iElectron. Lettėctr. EngėNeural Netw.ls 2023 34 2 635–649 Feb.

[32] Hao Shen and Yun Wang and Huaicheng Yan and Shengyuan Xu Data-driven single-loop policy iteration control of uncertain singularly perturbed systems iElectron. Lettėctr. Engėac 2025 70 12 8314–8320 Dec.

[33] Jianguo Zhao and Chunyu Yang and Weinan Gao and Ju H Park Novel single-loop policy iteration for linear zero-sum games AutoMechatronicsatica 2024 163 111551

[34] Ci Chen and Frank L Lewis and Kan Xie and Shengli Xie Adaptive optimal control of unknown nonlinear systems via homotopy-based policy iteration iElectron. Lettėctr. Engėac 2024 69 5 3396–3403 May

[35] Jiacheng Wu and Yang Zhu and Hongye Su Memory-Efficient Inverse Reinforcement Learning for Multiplayer Differential Games iElectron. Lettėctr. Engėc 2025 55 11 5545–5558 Nov.

[36] Dongdong Li and Jiuxiang Dong Cooperative Optimal Output Tracking for Discrete-Time Multiagent Systems: Stabilizing Policy Iteration Frameworks iElectron. Lettėctr. Engėac 2026 71 4 2746–2753 Apr.

[37] Jing Wang and Zheng Huang and Hao Shen and Ju H Park A Parallel Homotopic Optimized Control Scheme of Uncertain Nonlinear Markov Jump Systems and Its Applications iElectron. Lettėctr. Engėase 2025 22 19403–19414

[38] Gregory C Walsh and Hong Ye and Linda G Bushnell Stability analysis of networked control systems iElectron. Lettėctr. Engėcst 2002 10 3 438–446 May

Discussion: login to participate.