The team published new datasets pulled from arXiv, the go to repository for unreviewed research papers, and GitHub, the platform coders use to store and share their projects. The goal is to test “lead lag forecasting,” a method that uses early engagement, including views, downloads, and likes, to predict which papers or projects will matter years down the line. Right now, a paper’s importance is judged by how often other researchers cite it. The catch is that citations can...