cgit - CGI for Git ================== This is an attempt to create a fast web interface for the Git SCM, using a built-in cache to decrease server I/O pressure. Installation ------------ Building cgit involves building a proper version of Git. How to do this depends on how you obtained the cgit sources: a) If you're working in a cloned cgit repository, you first need to initialize and update the Git submodule: $ git submodule init # register the Git submodule in .git/config $ $EDITOR .git/config # if you want to specify a different url for git $ git submodule update # clone/fetch and checkout correct git version b) If you're building from a cgit tarball, you can download a proper git version like this: $ make get-git When either a) or b) has been performed, you can build and install cgit like this: $ make $ sudo make install This will install `cgit.cgi` and `cgit.css` into `/var/www/htdocs/cgit`. You can configure this location (and a few other things) by providing a `cgit.conf` file (see the Makefile for details). Lua is optional and only powers the lua: filter extensions (authentication, email and commit-message filters in custom/extensions/). A plain build auto-detects a Lua through pkg-config, preferring LuaJIT, and falls back to a Lua-less binary when none is found. Acceptable values are generally "luajit", "lua", "lua5.4", "lua5.3", "lua5.2" and "lua5.1". To pin an implementation: $ make LUA_PKGCONFIG=lua5.4 To build without Lua, so the binary needs no Lua at runtime: $ make NO_LUA=1 Previewing locally ------------------ cgit is a CGI program, so trying it out normally means configuring a web server. For development there is a small dependency-free preview server at `tools/serve.py` that runs the built binary and serves the static assets. It needs only Python 3. $ python3 tools/serve.py --config /path/to/cgitrc --port 8080 It is a development aid and is not meant to face the internet. Dependencies ------------ * zlib * optional: luajit or lua, most reliably used when pkg-config is available cgit builds without OpenSSL or libcurl, so no development packages for those are needed. Filter extensions ----------------- The optional Lua filters in `custom/extensions/` need extra Lua modules. Each script's header lists the exact install commands for its own dependencies. * The auth filters (`auth-file.lua`, `auth-inline.lua`) need `luaossl` and `luaposix`. * The email filters (`email-gravatar.lua`, `email-libravatar.lua`) need `luaossl`. * The syntax highlighter (`syntax-highlight.lua`) needs `lpeg` and a Scintillua lexer set. * The about-page renderer (`about-render.lua`) needs `lpeg` for markdown and man pages. Plain text needs only Lua. These filters target Lua 5.1 through 5.4 and LuaJIT. `luaossl` has no Lua 5.5 build, so build cgit against 5.1 to 5.4 if you use the auth or email filters. `link-commits.lua` needs nothing beyond Lua itself. Web server configuration ------------------------ cgit is a CGI program. Complete, commented configurations for nginx, Apache and lighttpd live in `custom/servers/`, each explaining how that server routes requests to the binary and serves the static assets off disk. URLs and query parameters ------------------------- cgit serves every page from one CGI program, so a url names the repository, the page and a path inside that page. There are two forms and they carry the same information. With virtual-root set, a request is a path, written as /// with anything else as a query string. Without it, the whole request lives in the query string as ?url=//, which is the form the test suite uses. In the path form the first extra argument opens the query string with a question mark, and in the url= form it continues the existing one with an ampersand, which is why links in the two forms are spelled differently. The page name is the second element and is one of about, atom, blame, blob, commit, diff, log, patch, plain, rawdiff, refs, snapshot, stats, summary, tag or tree. Leaving it out gives the repository summary, and leaving the repository out as well gives the index of repositories. Everything after the page name is the path argument, so /demo/tree/src/main.c asks for the tree page at src/main.c. A repository whose name contains a character that is special in a url has it percent-encoded, and a plus sign has to be written %2b because a bare plus decodes to a space. These parameters are accepted, all of them optional. url repository, page and path in one value, used instead of the path form h the branch or ref to read, defaulting to the repository default id pin the page to one commit or object, which every page that shows history honours id2 the second object for a diff, so id and id2 name the two sides ofs offset into a paged listing, used by log, refs and stats path restrict the page to one path, equivalent to the trailing path q the search term qt what to search, one of grep, author, committer or range s sort key on the index and refs pages showmsg show full commit messages in a log listing period the statistics window, one of w, m, q or y dt diff type, selecting unified, side by side or raw ss shorthand for the side by side diff all include every ref rather than one branch, used by atom context lines of context in a diff ignorews ignore whitespace when diffing follow follow a single path across renames in a log A few endpoints are not ordinary pages. The snapshot page takes a filename rather than a ref, so /demo/snapshot/demo-1.0.tar.gz names both the ref and the archive format through the suffix, and the formats on offer are set by the snapshots option. The plain page serves a blob as its own bytes under headers that stop a browser treating repository content as markup. The atom page is a feed rather than a page and accepts h, path and all. The clone endpoints under info and objects implement the dumb HTTP protocol and are only present when http clone is enabled. Some worked examples, in the url= form. ?url= the repository index ?url=demo summary for the demo repository ?url=demo/log&h=next log of the next branch ?url=demo/log&qt=author&q=alice commits authored by alice ?url=demo/tree/src&h=v1.0 the src directory at tag v1.0 ?url=demo/commit&id=HEAD~3 one commit, pinned ?url=demo/diff&id=main&id2=next diff between two branches ?url=demo/plain/README.md the raw bytes of one file ?url=demo/atom&h=main the commit feed for a branch Runtime configuration --------------------- The file `/etc/cgitrc` is read by cgit before handling a request. In addition to runtime parameters, this file may also contain a list of repositories displayed by cgit (see `cgitrc.5.txt` for further details). A fully commented starting point with every option at its default is in `custom/cgitrc`. Securing an instance -------------------- A public instance needs a few deliberate choices, all set in cgitrc and documented in `cgitrc.5.txt`. * Keep private repositories out of `scan-path`, or set `strict-export` to a marker filename so only repositories that contain it are published. * Gate the whole instance behind a login with `auth-filter`. Two example filters ship in `custom/extensions/`, `auth-inline.lua` and `auth-file.lua`. * Terminate TLS at the web server in front of cgit. * The example configs in `custom/servers/` set a Content-Security-Policy and related headers at the web server, where they also cover the static assets. * `max-blob-size` bounds how much a single request reads into memory, and defaults to 10 MB. * Leave `enable-cache-list` off, since it exposes the cache path and the URLs other visitors requested. * Build the deployed binary with the hardening flags via `tools/release-build.sh`. The cache --------- When cgit is invoked it looks for a cache file matching the request and returns it to the client. If no such cache file exists (or if it has expired), the content for the request is written into the proper cache file before the file is returned. If the cache file has expired but cgit is unable to obtain a lock for it, the stale cache file is returned to the client. This is done to favour page throughput over page freshness. The generated content contains the complete response to the client, including the HTTP headers.