The datasource container starts empty: food search returns nothing until you import at least one catalog. Each catalog is independent and builds into its own database file, so you can import them in any order and the service serves whatever is present. Imports never interrupt the running service: each one builds a new file, verifies it, then swaps it in atomically.
Run all commands from the selfhost/ directory unless otherwise noted.
USDA FoodData Central
wger Exercise Catalog
Open Food Facts
USDA FoodData Central contains generic whole foods with lab-measured nutrition data, all in the public domain. This is the catalog that makes a search for “chicken breast” return an actual ingredient rather than a supermarket SKU. Start here. It is the most useful single catalog for day-to-day logging.The download is approximately 3.1 GB and the import takes a few minutes.Download the USDA dataset
From the selfhost/ directory, download and unzip the USDA CSV export: Copy the files into the container
Clean up the temporary files
The wger catalog provides the exercise database used throughout the training features. Unlike the food catalogs, there is no file to download, because the importer fetches data directly from the wger API.Run the import
From the selfhost/ directory:There is no download step and no cleanup required. The import fetches the catalog over the network and builds the database file in place. Open Food Facts contains packaged products and barcodes from around the world. This is the catalog that makes barcode scanning useful outside the US, and it includes product photos that USDA does not have.The full dump is large and the import is the most time-consuming of the three. Try a partial import first to confirm everything is working before committing to the full run.The gzip file is read as a stream and never fully expanded on disk, so you only need space for the download itself rather than the ~50 GB it would unpack to. Products with no barcode, no name, or no nutrition data are dropped during import.
Download the Open Food Facts dump
From the selfhost/ directory: Copy the file into the container
Run a partial import to verify (recommended first)
A limit of 50,000 products takes minutes rather than hours and is enough to confirm barcode scanning and search are working: Run the full import when ready
Once you are satisfied, re-run without the limit to import the entire catalog: Clean up the temporary file
Checking import stats
To see how many records are in each catalog currently serving requests:
Rolling back an import
Each import keeps the previous database file so you can roll back if something goes wrong. To revert a catalog to its previous state:
The datasource README at apps/datasource/README.md documents how the catalogs are ranked, the de-duplication passes that run during import, and how search merges results across catalogs when more than one is loaded.