Official Project Description
With the explosion of AI-based models and architectures, a ripe opportunity presents itself to use these statistical models to help us bridge the gap large scale simulations and functional insight.
In particular, with the explosion of methods like AlphaFold2, there is a clear potential for these models to potentially predict dynamics, or use them to predict different conformations of a system at extremely large scales for a diverse set of sequences. However, for folks to be able to generate those kinds of models, a broad set of training data is needed that captures dynamics across a variety of different protein topologies.
This project seeks to generate that dataset - capturing dynamics of systems across a variety of different protein sizes and topologies.
17651: Small proteins (20,000 atoms)
17652 and 17654: medium sized proteins (40,000 atoms)
17653 and 17655: Large sized proteins (70,000 atoms).