# Does the Package Server always download the entire repository?

**URL:** <https://racket.discourse.group/t/does-the-package-server-always-download-the-entire-repository/3268>\
**Category:** Questions & Answers\
**Tags:** package-server\
**Created:** [October 26, 2024, 4:11pm UTC](https://racket.discourse.group/t/does-the-package-server-always-download-the-entire-repository/3268 "2024-10-26T16:11:22Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![default.kramer](https://avatars.discourse-cdn.com/v4/letter/d/ea5d25/32.png) [@default.kramer](https://racket.discourse.group/u/default.kramer)\
**Post date:** [October 26, 2024, 4:11pm UTC](https://racket.discourse.group/t/does-the-package-server-always-download-the-entire-repository/3268/1 "2024-10-26T16:11:22Z")

</div>

I have a package configured using the "Path within repository" option. I thought this would cause the package server to download only that part of the repository. The path is [https://github.com/default-kramer/HermitsHeresy.git?path=hermits-heresy#release](https://github.com/default-kramer/HermitsHeresy.git?path=hermits-heresy#release)

But I just noticed it is running out of disk space "while updating the package checksum" and it appears to be downloading the entire repository. (Here it shows that it is downloading the `/test` subdirectory, which I hoped to exclude.)

```scheme
copy-file: error writing destination file
  source path: /var/tmp/git17299539951729953995740/obj166
  destination path: /var/tmp/17299539951729953995740-default-kramer_HermitsHeresy_git_release/test/fresh-topias/4pvf1r91tm1.BIN
  system error: No space left on device; errno=28

```

Is it possible to configure the package so that only one subdirectory will be downloaded? Or do I need to move the desired package content into a separate repository?

---

<div class="post-metadata">

**Author:** ![Zeb](https://avatars.discourse-cdn.com/v4/letter/z/57b2e6/32.png) [@Zeb](https://racket.discourse.group/u/Zeb)\
**Post date:** [October 30, 2024, 6:26pm UTC](https://racket.discourse.group/t/does-the-package-server-always-download-the-entire-repository/3268/2 "2024-10-30T18:26:15Z")

</div>

I think it's only supposed to install where the info.rkt file is but I'm not sure. I reme this article has some relevant information: [How to Organize Your Racket Library – Terminally Undead](https://countvajhula.com/2022/02/22/how-to-organize-your-racket-library/#how-to-adopt-lib-test-doc)

---

<div class="post-metadata">

**Author:** ![greghendershott](https://yyz2.discourse-cdn.com/free1/user_avatar/racket.discourse.group/greghendershott/32/98_2.png) [@greghendershott](https://racket.discourse.group/u/greghendershott)\
**Post date:** [October 30, 2024, 6:57pm UTC](https://racket.discourse.group/t/does-the-package-server-always-download-the-entire-repository/3268/3 "2024-10-30T18:57:25Z")

</div>

AFAIK When a package is backed by a git repo, `raco pkg install` or `update` needs to clone that git repo. AFAIK git doesn't let you clone only a subdirectory. So I don't see how this could work.

For the granularity of "what is downloaded", AFAIK you probably need to use distinct repos. At least, you won't have to wonder, it will be clear that's the granularity you're getting.

For the granularity of "what is installed/built", you can have multiple packages as subdirs within a repo.

* * *

If the issue is mainly that you have some huge example files/data used for test purposes, maybe you can move just those to be in a distinct repo, and the test itself does a `git clone` or `wget` or `curl` or whatever you prefer?

You could even have your test do that only when an env var like "CI" is defined, so this happens for say GitHub Actions. but not when users run tests as a result of building package locally.

---

<div class="post-metadata">

**Author:** ![LiberalArtist](https://yyz2.discourse-cdn.com/free1/user_avatar/racket.discourse.group/liberalartist/32/151_2.png) [@LiberalArtist](https://racket.discourse.group/u/LiberalArtist)\
**Post date:** [October 30, 2024, 8:05pm UTC](https://racket.discourse.group/t/does-the-package-server-always-download-the-entire-repository/3268/4 "2024-10-30T20:05:06Z")

</div>

> [@greghendershott](#):
>
> AFAIK When a package is backed by a git repo, `raco pkg install` or `update` needs to clone that git repo. AFAIK git doesn't let you clone only a subdirectory. So I don't see how this could work.
> 
> For the granularity of "what is downloaded", AFAIK you probably need to use distinct repos. At least, you won't have to wonder, it will be clear that's the granularity you're getting.

It is possible to limit what is downloaded without making a separate repository! When possible, the package manager will make a “shallow clone” (like [`git clone --depth=1`](https://git-scm.com/docs/git-clone/2.47.0#Documentation/git-clone.txt-code--depthcodeemltdepthgtem), but implemented in pure Racket by the `#:depth` argument to [`git-checkout`](https://docs.racket-lang.org/net/git-checkout.html#%28def._%28%28lib._net%2Fgit-checkout..rkt%29._git-checkout%29%29) with only the contents of the desired commit, not the full history. So, you can create a release branch that leaves out the problematic directory. It looks like @default.kramer figured out that workaround in [GitHub - default-kramer/HermitsHeresy at release-for-racket-pkg](https://github.com/default-kramer/HermitsHeresy/tree/release-for-racket-pkg).

It seems like some fairly new features in Git _might_ support automatically downloading only a desired subdirectory, particularly [`git clone --filter=`](https://git-scm.com/docs/git-clone/2.47.0#Documentation/git-clone.txt-code--filtercodeemltfilter-specgtem) with its [_`<filter-spec>`_](https://git-scm.com/docs/git-rev-list/2.45.0#Documentation/git-rev-list.txt---filterltfilter-specgt) options and the experimental [`git sparse-checkout`](https://git-scm.com/docs/git-sparse-checkout) command ([more docs](https://git-scm.com/docs/sparse-checkout)). For Racket's purposes, what matters is not the `git` porcelain per se, but what we can do with the underlying wire protocol.

---

<div class="post-metadata">

**Author:** ![benknoble](https://yyz2.discourse-cdn.com/free1/user_avatar/racket.discourse.group/benknoble/32/16_2.png) [@benknoble](https://racket.discourse.group/u/benknoble)\
**Post date:** [October 30, 2024, 9:00pm UTC](https://racket.discourse.group/t/does-the-package-server-always-download-the-entire-repository/3268/5 "2024-10-30T21:00:56Z")

</div>

I can confirm that you can use Git to clone just a specific directory into a checkout: [clone supports a sparse option](https://git-scm.com/docs/git-clone#Documentation/git-clone.txt-code--sparsecode) that checks out only the top level, which you can then embiggen with sparse-checkout.

This is independent of partial clones using depth or filter.
