Maybe we're getting too far into the weeds and/or I don't fully understand the design or use case for your software. It just makes no sense to me to have files represented in a system by directly referring to chunks internal to other compressed files (if that is indeed what you're saying). As I said, these operations are to a certain extent fungible. If two or more files are similar enough to yield a significant savings from chunking, I'd be inclined to avoid all the overhead of rolling my own solution (database backed or not) and simply compress them together into a single common archive file. aacbc168784b42c37675077473cbcc8dafe39890261552332715ff7d32414a1d