https://arxiv.org/abs/2402.15391v1 world envs… i love google papers they follow no conventions whatsoever lmao lmao the build up

alright how do they do it

how does it work

it’s trained on video only data. they pass the latent actions and a current state to get the next state.

since classic transformers, quadratic transformer attention too heave so they use something spatio-temporal transformers. where they have a spatiotemporal blocks with interleaved spatial and temporal attention llayers followed by a feed forward layer (just one?) spatial layer has tokens for each time step and temporal continues across time steps. they also have a casual mask which makes sense ig. and since it’s a game it would be dependent on the frames and continued history (i’m not sure about time context though) so frames can be directly proportional to the complexity…

this is weird though i don’t get how they are interleaving it and then is it only one layer afterwards?

the flow is like this vid tokenizer and latent action look sane ig we gotta understand the dynamics model on how it predicts the next frame.

latent action model

so obviously since they have video data they don’t know the actions to do the geniuses decided they will just make a model to actually decide the actions… the flow is encoder that converts all previous input frames and the next frame and outputs a set of continuous latent actions decoder then takes all previous frames and latent actions as input and predicts the next frame . since it’s a VQ-VAE it just tries to re construct these actions by itself the code book here are the actual actions.

i don’t understand why train frame to action though? like they won’t use it later they just map the user action to embedding anyways if it’s just going to be the same vector for the user action then why not just use a simple encoder at the beginning why train a VAE for that ?

oh wait fuck me because it’s unsupervised they don’t have labels to compare with so they just fuck all run it with VAEs predict the last action man damn. i will need to come back to this later…

video tokenizer

they use VQ-VAE again take in frames of video as input and generate discrete representations for each frame spatial transformers is just them braking each frame into a patches each patches then for attention layer it attends spatially each patch to the other patch in the same frame and then temporal each patch attending to patch at same position in previous frames. one after other. then they quantize the vector and then pass it as to the decoder to reconstruct it.

dynamics model

standard transformer takes in tokens from the tokenizer for clips takes in actiosn from LAM aims to predict the final frame from the next frame.