I am fairly new in the selfhosting business. I host a few services like Jellyfin or AdGuard, but most of my data just sits on plain SMB shares on my NAS.
I am interested in setting up things like Paperless-ngx, Nextcloud maybe Immich as well. One thing I am struggling with is how to organize all this so I don’t have the data scattered around multiple places or in the worst-case even duplicated. I like the simplicity of network shares and the fact that I am not relying on a third-party application to keep being maintained. I am currently not sure I am willing to give that up. However, I am also intrigued by the features these services provide. I know Nextcloud has the option to mount external storage, but I don’t know which downsides come with that. It gets more complicated with Paperless. As far as I know you typically have a consume directory where you throw in your data and Paperless stores it using its own system in the media directory. This means I either throw in everything or I suddenly have two locations where my documents are stored. If I would mount that media directory in things like Nextcloud I probably wouldn’t be able to find anything because of the different structure and naming scheme.
The idea that I had in mind is that I have a single directory structure of my data that can be used on its own and all these tools are just different frontends and provide different views and information of the same data. Maybe this approach is just something from the past and I should move on.
I should add that I am not planning to expose any of these services to the public. All of this is only accessed from inside my house or using a Wireguard VPN.
- How do you guys handle all this?
- How do you avoid data duplication?
- How do you avoid multiple potential file locations? Is a document in my Nextcloud, in my Paperless or just on the network share?
- Do you prepare in some ways in case an application stops being maintained?
- How has this reliance on multiple services impacted other things e.g. your backups?


Hey, Welcome! You’re having a wonderful and sometimes exhausting journey in front of you :)
I’ll just brain dump based on your questions and my associations. Hope something useful is in between!
First the basic setup options because they tie into how to handle your date flow:
basic Most popular I think is docker compose: here I suggest splitting it into one compose file per service though with one file holding your port config. This prevents you yourself getting confused by your port mappings :)
Second in line is a proxmox setup - similar vein and I lack the hands-on experience to talk about the difference.
Then there’s the “everything native” approach where you don’t rely on containers but manage it yourself or via a dedicated OS that makes life easier (after the learning curve) like nixos.
DATA
All this foundational stuff is important because it changes your approach. In general: don’t fear data duplication. Duplicate it until you learn where you want your data to life and only then define your flow.
Specific example: after I got used to paperless I don’t look into my opencloud anymore, at all. I still duplicate them there but as distributed backup, not for consumption.
If a dataset has a clear place ten it’s easy. If not then your options are different depending on your setup: For the *arr stack the official recommendation is to use one shared folder and mount that into each part for example. I personally don’t like that and have hard links for everything - that’s basically a pointer to the file that looks like the file itself everywhere. As long as one pointer exists the file still stays on your drive but when the last pointer is gone, the file is effectively deleted. On Linux, you can think of every file this way but by default only one pointer exists (which often people test as synonymous to “the file”. Drawdown: this only works really well if you manually keep either track of which tool links where or you don’t containerize everything.
Again a specific example: My downloaded torrents never get moved - instead hard links are created into whichever path and naming scheme I defined for each consumer - this way, out of murdrrbot_07.mp3 a new author/series/booktitle.mp3 was created, both pointing to the same data and seeing it as a proper file.
But then there is one more thing: I suggest you split your thinking into data consumption and manipulation - because for the first, data duplication doesn’t matter. Especially for documents you’re talking about a ridiculous small amount of disk space and if it’s only reading/watching/hearing you as manager have no problem that data might exist multiple times.
If you want to keep it clean by design then you’re leaving the starter mode self holster - welcome to system design and infrastructure architecture! Here your approach could be to define lifecycles for each data type that you have. What a “data type” is in this context btw is a user term, NOT the underlying tech stack. You need to understand and document how an invoice should be treated and consumed by you differently than an invitation or a informal letter. Only then do you map file types, incoming channels, transformation steps, etc etc.
In my opinion: huge overkill to this upfront.
In short: spin everything up, observe how you use it and only then decide where things need to stay unique and cleaned up. Don’t break your head over something that’s actually quite easy to repair!
As I have said in the other comment: I already use Docker and docker-compose for all the services I host. They are also all configured to use bind mounts instead of volume mounts, so that I am in control where everything is located. I heard about Proxmox, but I never really looked into it because never saw the need for something different than a docker container. I also don’t have a dedicated server. All my services are running as Docker containers directly on the NAS.
I am not a fan of data duplication, disk space aside. You are pretty much guaranteed to have diverging file structures sooner or later. I don’t want to look up a file on three different applications just to find the newest version of it. I know you can use rsync and a cron job, but that just adds more complexity to a problem that I don’t want to have in the first place. This might work for something like a read-only backup like i presume you do with Paperless and OpenCloud, but I am not sure how this handles a case where, at least in theory, your files can be changed, renamed etc. in multiple different locations.
How do you handle your Paperless documents? Do you have a local file structure that you manage on your own for these documents or do you shove them all into Paperless and process them entirely in there (naming, tagging, etc.)?
I get your idea of trying things out even though it might result in temporary data duplication to find the way that works best for me. I am just curious how other peoples workflow looks like. Maybe I can also learn from the mistakes other people made in the past :)
Paperless specifically: yes, shoving everything in it.
But I have honestly no idea what kind of data you have that are suitable for paperless AND have different versions.
I just checked mine, for me it’s… Zero. Not a single item, by definition, is in Paperless that can receive an update.
If it’s updatable personally I have everything version controlled - and I mean EVERYTHING, from CV over tutorials to construction ideas - hosted locally on forgejo.