Time-varying policy function

Question

Matheus Silva le 24 Mai 2023

0
Lien

Utiliser le lien direct vers cette question

https://fr.mathworks.com/matlabcentral/answers/1972534-time-varying-policy-function

Commenté : Emmanouil Tzorakoleftherakis le 30 Mai 2023

Hi,

I am wondering if it is possible to have time-varying (non-stationary) policy functions in the reinforcement learning toolbox.

For example, say my episode lasts three periods (t=1,2,3), then I would have the set

where

is some neural network structure indexed by a general vector of parameters ϑ, which will ultimately depend on the time period.

Is that possible to do with the toolbox?

Thank you so much!

0 commentaires
Afficher -2 commentaires plus anciensMasquer -2 commentaires plus anciens

Connectez-vous pour commenter.

Connectez-vous pour répondre à cette question.

Answer 1

Emmanouil Tzorakoleftherakis le 25 Mai 2023

0
Lien

Utiliser le lien direct vers cette réponse

https://fr.mathworks.com/matlabcentral/answers/1972534-time-varying-policy-function#answer_1244654

Why don't you just train 3 separate policies and pick and choose as needed?

4 commentaires
Afficher 2 commentaires plus anciensMasquer 2 commentaires plus anciens

Matheus Silva le 28 Mai 2023

Modifié(e) : Matheus Silva le 28 Mai 2023

My problem is that my periods can be related in some arbitrary way. For example, I am thinking of a model where the state

can vary according to

Where

is a stochastic term and

is some transition function. However, I may want to allow some relation between the stochastic terms in periods 1 and 3. Solving the problem period by period would eliminate that dependence, no?

Emmanouil Tzorakoleftherakis le 30 Mai 2023

Honestly, I think your best bet would be to use the same policy throughout, but maybe use an input signal to the neural net to indicate which period you are in based on your state.

Another option, which is similar to what I mentioned earlier, is to train 3 different policies. To work around the period dependencies, you can place the RL policy block inside a triggered subsystem and only enable the subsystem for training when the system is in the appropriate period. Do that for each policy and then you can switch between the 3 as needed. See here

Connectez-vous pour commenter.

Time-varying policy function

0 commentaires
Afficher -2 commentaires plus anciensMasquer -2 commentaires plus anciens

Réponses (1)

4 commentaires
Afficher 2 commentaires plus anciensMasquer 2 commentaires plus anciens

Voir également

Catégories

Tags

Produits

Version

Community Treasure Hunt

Time-varying policy function

0 commentaires Afficher -2 commentaires plus anciensMasquer -2 commentaires plus anciens

Réponses (1)

4 commentaires Afficher 2 commentaires plus anciensMasquer 2 commentaires plus anciens

Voir également

Catégories

Tags

Produits

Version

Community Treasure Hunt

0 commentaires
Afficher -2 commentaires plus anciensMasquer -2 commentaires plus anciens

4 commentaires
Afficher 2 commentaires plus anciensMasquer 2 commentaires plus anciens