I had essentially the same problem, which I tasked myself to solve at the start of the year. I recently wrote about some of the nuts-and-bolts things that I did [in this post](https://www.reddit.com/r/DataHoarder/comments/hhou76/critique_my_archivingbackup_strategy/), but I didn't cover any of the curation bits there. I think the best approach is to ask yourself what is the *end state* you want, and then figure out how to get there from where you're starting. Do you want to (*can* you) reasonably create a single, central pool like RoboYoshi suggests? If that's not possible for you, what *is* the way you want subsets of your data organized? By file format? Topic? Timeline? I took an "activity level" approach to sorting my own data. If there's something I haven't done much with in 5 years, it didn't get a lot of attention other than copying for preservation. Recent/current files get grouped by the events/activities they're associated with (e.g., Business Clients, Personal Finances, etc.) As files get used, though, they get pulled out of the archive and I can think about reclassifying them as needed. So, generally, I'd recommend a breadth-first strategy. If you want to get it all sorted immediately, pick those broadest interests that make sense to you and get all the old stuff into the folder/disk they belong. Then drop down a level and figure out if there is any way to segment items into smaller groups. Repeat until you're satisfied that you can now quickly find what you're looking for. If you have multiple copies of files floating around (e.g., backups), you *will* want to figure out what to do with them at some point. That's part of why I wrote my own software; turns out only about a quarter of my storage space was unique files. By using a warehousing approach for my backups/archive, I can minimize the amount of space taken up by unnecessary duplicates. ad7f95b1422d207888a2ebf897ba3bc087d75bd9842e16d93ac2f5a7c46486bd