I’m interested in this answer as well. @sgugger has an excellent post on single precision vs mixed precision. I just posted a question on the Mixed Precision thread asking him what memory usage he’s seen in practice on the mixed precision work he’s done.