From 40ed94590f149cd3656ddd09dfa0141f37686d7c Mon Sep 17 00:00:00 2001 From: tobi Date: Mon, 15 Jun 2026 09:34:28 +0200 Subject: [PATCH] [docs] Update db maintenance docs, add status cleanup docs, update media cleanup docs (#4857) Reviewed-on: https://codeberg.org/superseriousbusiness/gotosocial/pulls/4857 Reviewed-by: kim --- docs/admin/database_maintenance.md | 43 +++++++++++++++++--- docs/admin/media_caching.md | 44 ++++++++++++--------- docs/admin/post_caching.md | 63 ++++++++++++++++++++++++++++++ docs/configuration/statuses.md | 2 +- example/config.yaml | 2 +- mkdocs.yml | 1 + 6 files changed, 129 insertions(+), 26 deletions(-) create mode 100644 docs/admin/post_caching.md diff --git a/docs/admin/database_maintenance.md b/docs/admin/database_maintenance.md index be52e64ec..17e894d87 100644 --- a/docs/admin/database_maintenance.md +++ b/docs/admin/database_maintenance.md @@ -12,7 +12,7 @@ Regardless of whether you choose to run GoToSocial with SQLite or Postgres, you ## SQLite -### Vacuum +### SQLite Vacuum To minimize fragmentation, GoToSocial does not currently enable auto-vacuum for SQLite. To defragment the database file and repack it to an optimal size you may want to run a `VACUUM` command on your SQLite database periodically (eg., every few months). @@ -30,7 +30,7 @@ Once you've met these requirements, do the following: This may take quite a few minutes depending on the size of your database. DO NOT INTERRUPT IT. 3. When the command has finished running, start GoToSocial again. -### Analyze / Optimize +### SQLite Analyze / Optimize GoToSocial runs a [full analyze command](https://sqlite.org/lang_analyze.html) after each set of database migrations (eg., when starting an updated version of GoToSocial), to ensure that any indexes added or removed by migrations are taken into account correctly by the query planner. @@ -40,7 +40,7 @@ Because of the above automated steps, in normal circumstances you should not nee However, if you notice that queries are running very slowly, it could be the case that the index metadata stored in SQLite's internal tables has become out of date, or has been removed or otherwise undesirably altered, leading the query planner to make poor choices. -This is particularly prone to happening if a large cleanup operation has just occured, eg., you've just [cleaned up a lot of old statuses](../configuration/statuses.md). +This is particularly prone to happening if a large cleanup operation has just occured, eg., you've just [cleaned up a lot of old statuses](./post_caching.md). If you notice lots of timeouts, for example when trying to view your timelines or profile page, you can use the GoToSocial binary to manually run a full `analyze`. @@ -72,10 +72,43 @@ The command may take up to 15 minutes to finish running, depending on the size o You will likely notice degraded performance of GoToSocial while the `analyze` is running, this is normal. If you prefer, you can stop GoToSocial before running the command, and start it again after running the command. -### Replication +### SQLite Replication It's a common practice to set up safeguards for your database like replication. SQLite can be replicated using external software. The basic steps are described on the [Replicating SQLite](../advanced/replicating-sqlite.md) page. ## Postgres -TODO: Maintenance recommendations for Postgres. +The commands in this section rely on having the Postgres [`psql` command line tool](https://www.postgresql.org/docs/current/app-psql.html) installed either on the same machine that is running Postgres, or inside the Postgres docker container. + +### Postgres Vacuum + +By default, [autovacuum is enabled for Postgres](https://www.postgresql.org/docs/current/runtime-config-vacuum.html#GUC-AUTOVACUUM). This means that unless you've configured your Postgres deployment differently from the default, you should not need to run a manual vacuum operation on your Postgres GoToSocial database. + +However, if you have autovacuum disabled, or have recently [removed a lot of entries from the database](./post_caching.md), you may wish to run a vacuum manually. + +1. Ensure you have some spare disk space, about 1.5x the size of the GoToSocial database. +2. Stop GoToSocial. +3. With `psql` connected to the GoToSocial database, run the command `VACUUM FULL;`. + This may take quite a few minutes depending on the size of your database. DO NOT INTERRUPT IT. If you want more feedback you can run `VACUUM FULL VERBOSE;` instead. +4. When the command has finished running, start GoToSocial again. + +For more info, see the [Postgres docs for the `VACUUM` command](https://www.postgresql.org/docs/current/sql-vacuum.html). + +### Postgres Analyze + +If you notice degradation in the performance of GoToSocial on your Postgres database, it is possible that you need to use `psql` to force Postgres to rebuild statistics that the query planner uses to make decisions on which indexes to use. + +To do this: + +1. Ensure you have some spare disk space, about 1.5x the size of the GoToSocial database. +2. Stop GoToSocial. +3. With `psql` connected to the GoToSocial database, run the command `VACUUM FULL ANALYZE;`. + This may take quite a few minutes depending on the size of your database. DO NOT INTERRUPT IT. +4. DO NOT SKIP THIS STEP: When the previous command has finished running, run the command `VACUUM (DISABLE_PAGE_SKIPPING ON);` to force Postgres to rebuild its [visibility map](https://www.postgresql.org/docs/current/routine-vacuuming.html#VACUUM-FOR-VISIBILITY-MAP). + This may take quite a few minutes depending on the size of your database. DO NOT INTERRUPT IT. +5. When the second command has finished running, start GoToSocial again. + +For more info on this process, see [this issue about Postgres performance](https://codeberg.org/superseriousbusiness/gotosocial/issues/4757#issuecomment-17348090). + +!!! tip "Verbose" + If you want more feedback when running either of the above commands, you can use [the `VERBOSE` option](https://www.postgresql.org/docs/current/sql-vacuum.html). diff --git a/docs/admin/media_caching.md b/docs/admin/media_caching.md index bbd8b8995..bf0d83af5 100644 --- a/docs/admin/media_caching.md +++ b/docs/admin/media_caching.md @@ -1,4 +1,4 @@ -# Media Caching +# Media Caching and Pruning GoToSocial uses the configured [storage backend](https://docs.gotosocial.org/en/latest/configuration/storage/) in order to store media (images, videos, etc) uploaded to the instance by local users, as well as to cache media attached to posts and profiles federated in from remote instances. @@ -21,37 +21,43 @@ Remote media, on the other hand, is cached only temporarily. After a certain amo Cleanup of the remote media cache occurs as a scheduled background process, and no manual intervention is required by admins. Cleanup takes somewhere between 5-30 minutes depending on the speed of the server, the speed of the configured storage, and the amount of media to work through. -GoToSocial exposes three variables that let you, the admin, tune when and how this work is performed: `media-remote-cache-days`, `media-cleanup-from` and `media-cleanup-every`. +GoToSocial exposes two variables that let you, the admin, tune when and how this work is performed: `media-cleanup-cron` (accepts a cron expression), and `media-remote-cache-duration` (accepts a human-language duration string). + +!!! info "Cron expressions" + A ["cron expression"](https://en.wikipedia.org/wiki/Cron#Cron_expression) is a string that allows a user or computer administrator to specify when a background task should be run. Cron expressions are typically used in programming and system maintenance to schedule jobs using the program ["cron"](https://en.wikipedia.org/wiki/Cron), which is a time-based job scheduler. Because of their ubiquity, however, cron expressions are also accepted by some other programs -- such as GoToSocial! -- to allow users to customize the scheduling of background tasks. + + For more information on cron expressions and for help writing them, see the following resources: + + - [wikipedia page for cron expressions](https://en.wikipedia.org/wiki/Cron#Cron_expression) + - [cron expression helper website](https://crontab.guru) By default, these variables are set to the following values: -| Variable name | Default | Meaning | -|---------------------------|--------------|----------| -| `media-remote-cache-days` | `7` | 7 days | -| `media-cleanup-from` | `"00:00"` | midnight | -| `media-cleanup-every` | `"24h"` | daily | +| Variable name | Default | Meaning | +|-------------------------------|--------------|------------------------------------------------| +| `media-cleanup-cron` | `0 0 * * *` | Cron expression meaning every night @ midnight | +| `media-remote-cache-duration` | `7 days` | 7 days | -In other words, the default settings mean that every night at midnight, remote media older than a week will be uncached and removed from storage. +In other words, the default settings mean that every night at midnight, remote media older than seven days will be uncached and removed from storage. -You can achieve different results by tuning these variables. For example, say you wanted to prune at 4.30am instead of midnight, you could change `media-cleanup-from` to `"04:30"`. +You can achieve different results by tuning these variables. For example, say you wanted to prune at 4.30am instead of midnight, you could change `media-cleanup-cron` to `30 4 * * *`. -If you only want to prune every couple of days instead of every night, you could set `media-cleanup-every` to a higher value, like `"48h"` or `"72h"`. +If you only want to prune every two days instead of every night, you could set `media-cleanup-cron` to something like `0 0 */2 * *` -If you wanted to adopt a more aggressive cleanup strategy to minimize storage usage, you could set the following values: +If you wanted to adopt an aggressive cleanup strategy to minimize storage usage, you could set the following values: -| Variable name | Setting | Meaning | -|---------------------------|--------------|-------------| -| `media-remote-cache-days` | `1` | 1 day | -| `media-cleanup-from` | `"00:00"` | midnight | -| `media-cleanup-every` | `"8h"` | every 8 hrs | +| Variable name | Setting | Meaning | +|-------------------------------|---------------|-------------| +| `media-cleanup-cron` | `0 */8 * * *` | every 8 hrs | +| `media-remote-cache-duration` | `1` | 1 day | -The above settings would mean that every 8 hours starting from midnight, GoToSocial would prune any media older than 1 day (24hrs). The prune jobs would run at 00:00, 08:00, and 16:00, ie., midnight, 8am, and 4pm. With this configuration, the longest amount of time you could possibly keep remote media in your storage would be about 32 hours. +The above settings would mean that every 8 hours, GoToSocial would prune any media older than 1 day (24hrs). With this configuration, the longest amount of time you could possibly keep remote media in your storage would be about 32 hours. !!! tip - Setting `media-remote-cache-days` to 0 or less means that remote media will never be uncached. However, cleanup jobs for orphaned local media and other consistency checks will still be run using the schedule defined by the other variables. + Setting `media-remote-cache-duration` to 0 or less means that remote media will never be uncached. However, cleanup jobs for orphaned local media and other consistency checks will still be run using the schedule defined by the other variables. !!! tip You can also run cleanup manually as a one-off action through the admin panel, if you so wish ([see docs](./settings.md#media)). !!! warning - Setting `media-cleanup-every` to a very small value like `"30m"` or less will probably cause your instance to just constantly iterate through attachments, causing high database use for very little benefit. We don't recommend setting this value to less than about `"8h"` and even that is probably overkill. + Setting `media-cleanup-cron` to a very small value like every hour or less will probably cause your instance to just constantly iterate through attachments, causing high database use for very little benefit. We don't recommend setting this value to less than about every eight hours, and even that is probably overkill. diff --git a/docs/admin/post_caching.md b/docs/admin/post_caching.md new file mode 100644 index 000000000..0fb937c6b --- /dev/null +++ b/docs/admin/post_caching.md @@ -0,0 +1,63 @@ +# Post Caching and Pruning + +GoToSocial stores posts (both local and remote) in whatever [database backend](../configuration/database.md) the instance is configured to use. + +Typically, the `statuses` table that GoToSocial uses to store posts is by far the largest table in the database, and will continue to grow over time the longer you use your GoToSocial instance for, as your instance receives new posts from people you follow, and dereferences boosted posts and replied-to posts, etc. + +To avoid the issue of ever-increasing database sizes, GoToSocial provides a mechanism whereby posts can be regularly pruned from the database by a background job, freeing up space. + +## Which posts will be pruned + +When selecting posts to clean up, GoToSocial considers posts in the context of the thread that they're part of (if applicable). For the purposes of cleanup, single posts are considered to be in a thread of length 1. + +To be eligible for pruning, a thread must meet the following criteria: + +1. **All posts in the thread are remote**, ie., created by a remote account. If a local account created or has participated in a thread, that thread will never be removed using this pruning method. +2. **No posts in the thread have been interacted with by a local account**. If a local user has faved, boosted, or bookmarked any post in a thread, that thread will not be pruned. +3. **All posts in the thread have been created or fetched less recently than the duration `statuses-cleanup-remote-older-than`**. Ie., only threads that haven't been dereferenced or added to since `statuses-cleanup-remote-older-than` will be pruned. (See below for more info on this config setting.) + +If a thread meets the above criteria, then the cleanup process will mark that thread for removal, and all posts in the thread (and any attached media) will be removed from the database. + +!!! info "Posts will be refetched on demand" + Much like with [media caching + cleanup](./media_caching.md), if a post is removed from post pruning, it can always be fetched again by the instance later, for example if a local user looks up one of the posts in the thread by its URI/URL, or if the instance sees a post in that thread due to someone replying to it, etc. + +## How to enable remote post pruning + +By default, the remote post pruning task is not enabled. To enable it, update your config to set `statuses-cleanup-remote-older-than` to a valid duration string, for example "1 year", and restart your instance. Posts last created/fetched/updated before that duration will then become [eligible for cleanup](#which-posts-will-be-pruned). + +## When will post pruning occur + +The scheduling of the post cleanup operation is controlled by the config variable `statuses-cleanup-cron`, which accepts a cron expression as a value. By default, this is `"0 1 * * 0`, which means every Sunday at 1am (ie., weekly). + +!!! info "Cron expressions" + A ["cron expression"](https://en.wikipedia.org/wiki/Cron#Cron_expression) is a string that allows a user or computer administrator to specify when a background task should be run. Cron expressions are typically used in programming and system maintenance to schedule jobs using the program ["cron"](https://en.wikipedia.org/wiki/Cron), which is a time-based job scheduler. Because of their ubiquity, however, cron expressions are also accepted by some other programs -- such as GoToSocial! -- to allow users to customize the scheduling of background tasks. + + For more information on cron expressions and for help writing them, see the following resources: + + - [wikipedia page for cron expressions](https://en.wikipedia.org/wiki/Cron#Cron_expression) + - [cron expression helper website](https://crontab.guru) + +If you wish to this cleanup operation to run more or less frequently, depending on your needs and resources, you can adjust the config value to your liking. For example, to run the job once per month instead of once per week, you could change `statuses-cleanup-cron` to `0 1 1 * *`, which will cause the job to run at 1am on the first day of every month. + +!!! tip "First cleanup can take a long time" + If you've enabled remote post pruning for the first time on a large database (eg., larger than a few GB), be aware that the first cleanup operation can take quite a long time, as there will be a significant number of threads that must be iterated through, depending on how you've set `statuses-cleanup-remote-older-than`. + + For example, on a 17GB SQLite database with `statuses-cleanup-remote-older-than` set to "1 year", on a 1CPU, 1GB ram VPS, the first remote post prune operation was observed to take about 48+ hours in total. + + While the cleanup is running, depending on the specs of the machine that GoToSocial is running on, you may notice high CPU usage and some degradation of performance (eg., home timeline queries taking longer to respond, that sort of thing). This is normal and does not indicate a problem with the operation. + + After you've run a cleanup for the first time, subsequent cleanups should be much faster, as there will be fewer threads remaining that must be iterated through. + +## Considerations for after cleanup + +### SQLite + +With SQLite, when you clean up posts from the database the database file will not shrink in size automatically, as GoToSocial does not enable [autovacuum](https://sqlite.org/pragma.html#pragma_auto_vacuum) in SQLite. Instead, freed space will be held by the database file and written into when new posts arrive. + +If you want to actually shrink the size of your database file after cleanup, you should [run a `vacuum` command against it using the `sqlite3` CLI tool](./database_maintenance.md#sqlite-vacuum). To ensure indexes remain performant after removal of a large amount of statuses, you may also wish to [run `analyze` on your database file using the GoToSocial binary](./database_maintenance.md#sqlite-analyze--optimize) + +### Postgres + +By default, [autovacuum is enabled for Postgres](https://www.postgresql.org/docs/current/runtime-config-vacuum.html#GUC-AUTOVACUUM). This means that any space freed up by the removal of remote posts should be made available to the operating system again once the autovacuum daemon has done its thing, and you should not need to manually intervene in order to shrink your database volume. + +However, it may be that you wish to run vacuum manually, and/or reanalyze the database after the removal of a large amount of posts, to ensure that indexes remain performant. For steps on how to do this, check the [Postgres maintenance docs](./database_maintenance.md#postgres). diff --git a/docs/configuration/statuses.md b/docs/configuration/statuses.md index fa346e4a5..fc9c76fbe 100644 --- a/docs/configuration/statuses.md +++ b/docs/configuration/statuses.md @@ -44,7 +44,7 @@ statuses-cleanup-cron: "0 1 * * 0" # Integer duration. # -# Examples: ["7 days", "1 week", "1 month"] +# Examples: ["6 months", "1 year", "2 years"] # Default: "0" (i.e. disabled) statuses-cleanup-remote-older-than: "0" diff --git a/example/config.yaml b/example/config.yaml index 46a677670..d34d6cb45 100644 --- a/example/config.yaml +++ b/example/config.yaml @@ -951,7 +951,7 @@ statuses-cleanup-cron: "0 1 * * 0" # Integer duration. # -# Examples: ["7 days", "1 week", "1 month"] +# Examples: ["6 months", "1 year", "2 years"] # Default: "0" (i.e. disabled) statuses-cleanup-remote-older-than: "0" diff --git a/mkdocs.yml b/mkdocs.yml index 61fadf3a0..67357a2f0 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -148,6 +148,7 @@ nav: - "admin/cli.md" - "admin/backup_and_restore.md" - "admin/media_caching.md" + - "admin/post_caching.md" - "admin/spam.md" - "admin/database_maintenance.md" - "admin/themes.md"