the transitional stable difusion model between static and moving ai imagery is the 303 of latent visuals