Introduction
Welcome to Real World Irmin. This book is all about Irmin.
Irmin is an OCaml library for building mergeable, branchable distributed data stores.
- Built-in Snapshotting - backup and restore
- Storage Agnostic - you can use Irmin on top of your own storage layer
- Custom Datatypes - (de)serialization for custom data types, derivable via ppx_irmin
- Highly Portable - runs anywhere from Linux to web browsers and Xen unikernels
- Git Compatibility - irmin-git uses an on-disk format that can be inspected and modified using Git
- Dynamic Behavior - allows the users to define custom merge functions, use in-memory transactions (to keep track of reads as well as writes) and to define event-driven workflows using a notification mechanism
Concepts
What is Irmin?
Key Concepts
Architecture
A Primer on Functors
Programming applications in Irmin requires a familiarity with OCaml's functors. Here are a few resources for learning more about them:
- The OCaml manual: The definitive guide to different aspects of the OCaml language.
What follows is an Irmin-specific introduction to functors.
Generalisation
Functors, in a non-theoretical sense, can be thought of as functions over modules. A functor takes a module and produces another module. For example, we might produce a new module that hashes values provided they can be serialised to a string.
(* A module signature for modules who have a main type called [t]
and that provide a function [serialise] to convert [t] to a [string]. *)
module type Serialisable = sig
type t
val serialise : t -> string
end
We will also need another module signature describing what a digestable type looks like.
module type Digestable = sig
type t
(** The values to digest *)
type hash
(** The type of digests produced *)
val hash_to_string : hash -> string
(** Convert digests to a string *)
val digest : t -> hash
end
We could then provide a SHA256 hashing functor for serialisable types.
module SHA256 (S : Serialisable) : Digestable with type t = S.t and type hash = Digestif.SHA256.t = struct
type t = S.t
type hash = Digestif.SHA256.t
let hash_to_string = Digestif.SHA256.to_raw_string
let digest v = Digestif.SHA256.digest_string (S.serialise v)
end
We can then use our functor.
module Integer = struct
type t = int
let serialise = string_of_int
end
module Digest_int = SHA256(Integer)
And then use it.
# let d = Digest_int.digest 42;;
val d : Digest_int.hash = <abstr>
# Digest_int.hash_to_string d |> Base64.encode_exn;;
- : string = "c0dctApWjo2ooEXO0RATfhWfiQrE2og7axfcZRs6gEk="
Schemas
Every Irmin store has a schema. This specifies the concrete implementations of the various kinds of customisable aspects of the store. For example what type are the keys, the branches and the contents.
In Irmin everything gets specified upfront using functor application. Whereas with something like Hashtbl its type will be inferred but its use.
Most stores take a Schema module to provide the concrete implementation of the various customisable parts of the Irmin store. This can be a little verbose, but most of the time the defaults are just fine.
module Schema : Irmin.Schema.S = struct
open Irmin
module Hash = Hash.BLAKE2B
module Info = Info.Default
module Branch = Branch.String
module Path = Path.String_list
module Metadata = Metadata.None
module Contents = struct
type t = int[@@deriving irmin]
let merge = Irmin.Merge.(option @@ default t)
end
end
module S = Irmin_mem.Make (Schema)
(* A fake clock and info function for our commits *)
let clock = let time = ref 0L in fun () -> time := Int64.add !time 1L; !time
let info () = S.Info.v (clock ())
Here we've defined a schema to use Blake2B hash functions, default commit information, branches are strings, keys (or paths) are string lists, no extra metadata and finally the content of the store is integers.
However, we'll quickly run into problems as soon as we try to use the store.
# let config = Irmin_mem.config () in
let* repo = S.Repo.v config in
let* main = S.main repo in
let* () = S.set_exn ~info main [ "hello" ] 1 in
S.get main [ "hello" ];;
Line 4, characters 34-45:
Error: This expression has type 'a list
but an expression was expected of type S.path
This is because in our definition of the schema we hid the implementation details by adding : Irmin.Schema.S. So either we canleave that part out or we must provide a module type to expose the information (or at the very least the bits we need).
module type String_list_int_schema = Irmin.Schema.S
with type Hash.t = Irmin.Schema.default_hash
and type Branch.t = string
and type Info.t = Irmin.Info.default
and type Metadata.t = unit
and type Path.step = string
and type Path.t = string list
and type Contents.t = int
And then use that to define the store.
module Schema2 : String_list_int_schema = struct
open Irmin
module Hash = Hash.BLAKE2B
module Info = Info.Default
module Branch = Branch.String
module Path = Path.String_list
module Metadata = Metadata.None
module Contents = struct
type t = int[@@deriving irmin]
let merge = Irmin.Merge.(option @@ default t)
end
end
module S = Irmin_mem.Make (Schema2)
let info () = S.Info.v (clock ())
And now the compiler knows what the types are (they haven't been abstracted away).
# let config = Irmin_mem.config () in
let* repo = S.Repo.v config in
let* main = S.main repo in
let* () = S.set_exn ~info main [ "hello" ] 1 in
S.get main [ "hello" ];;
- : int = 1
Runtime Types
One thing that will quickly become very apparent is the need to define an 'a Irmin.Type.t for nearly every type. This is a so-called runtime type. It represents the type 'a at runtime as an OCaml value.
Internally it reuses the mirage/repr library. Some runtime types come predefined, such as that for boolean values.
# Irmin.Type.bool;;
- : bool Repr.ty = <abstr>
This carries information with it about how to interact with actual bool values, for example how to serialise it in different ways.
# Irmin.Type.to_string Irmin.Type.bool true;;
- : string = "true"
# Irmin.Type.to_json_string Irmin.Type.bool true;;
- : string = "true"
ppx_irmin
Because so many modules and functions expect a runtime type, Irmin provides a ppx that can in the majority of cases derive the runtime representation of your type for you.
# type t = { name : string }[@@deriving irmin];;
type t = { name : string; }
val t : t Repr.ty = <abstr>
You can see that this created a value called t (not to be confused with the type itself).
# Irmin.Type.to_string t { name = "Bob" };;
- : string = "{\"name\":\"Bob\"}"
You can also build up representations of more complex types using the built-in combinators.
# let t = Irmin.Type.(list (pair string (option int)));;
val t : (string * int option) list Repr.ty = <abstr>
Uses
Irmin mainly uses the runtime types for serialisation purposes. For example if your Irmin store is using the file system backend (Irmin.FS) then Irmin needs to know how to turn your contents into a string and your key into a file path. Conversely, it will need to know how to turn file paths into your path type and file contents into your content type.
# let s = Irmin.Type.to_string t [ "key1", Some 1; "key2", None ];;
val s : string = "[[\"key1\",{\"some\":1}],[\"key2\",null]]"
# let v = Irmin.Type.of_string t s;;
val v : ((string * int option) list, [ `Msg of string ]) result =
Ok [("key1", Some 1); ("key2", None)]
Backends
An Irmin backend describes how data is persisted. By abstracting over some notion of backend, Irmin can:
- Be very portable. See the portability section for more details.
- Be very efficient for specific datatypes. You can image providing a highly specific backend for use with your specific datatype.
There are lots of pre-existing backends including:
Irmin_mem: the in-memory backend. No data is actually persisted.Irmin_fs: a Unix filesystem implementation where there's a natural mapping from the steps in your path (key) implementation to file paths and your contents are stored in files.Irmin_indexeddb: a slightly more experimental backend that uses the browser's IndexedDB storage API.Irmin_pack: another filesystem-based backend but using pack files.
Making a Backend
MirageOS-style Portability
Contents of the Store
Irmin is a key-value store. The contents of the store references the value part.
This next section will cover basic concepts you might want to impose on your contents such as:
- How to define three-way merge functions for your contents.
- Always serialise the contents to JSON.
- Version your content store so you can update the type in the future.
Mergeable Datatypes
Contents in Irmin are mergeable datatypes (MDT). These are values that have a three-way merge function. If you are familiar with the git version control system then the idea will hopefully be familiar.
Whenever you want to store some new value x in an Irmin store S at key k, there are two other versions of x to consider.
- The current version of
xin storeSat keyklet's call itx'. - The shared lowest-common ancestor (LCA) of
xandx'in storeSat keyklet's call itlca.
A merge function takes these three values and either produces some "merged" value or we have a merge conflict and we return an error. A merge function for values of type 'a are of type 'a Irmin.Merge.f.
# #show_type Irmin.Merge.f;;
type nonrec 'a f =
old:'a Irmin.Merge.promise ->
'a -> 'a -> ('a, Irmin.Merge.conflict) result Lwt.t
Let's take a look at a few examples.
Mergeable Counters
The classic MDT is a counter, an integer value that can be incremented and decremented.
module Counter = struct
type t = int [@@deriving irmin]
let incr t = t + 1
let decr t = t - 1
end
By using ppx_irmin we derive a runtime representation of the type t. This creates a value in the module called t. We're not quite ready to use our new module to instantiate a new, in-memory key-value Irmin store though.
# module Store = Irmin_mem.KV.Make (Counter);;
Line 1, characters 16-43:
Error: Modules do not match:
sig
type t = int
val t : t Repr__Type.t
val incr : t -> t
val decr : t -> t
end
is not included in Irmin__.Contents.S
The value `merge' is required but not provided
File "src/irmin/contents_intf.ml", line 25, characters 2-30:
Expected declaration
We need to provide a merge function! What properties should it have? The main one is probably that simultaneous increments and/or decrements should not be lost in the new merged value.
For example if Alice has a copy of the counter and increments it five times and Bob has a copy and decrements it twice, the final merged counter should be the lowest common ancestor plus five minus two.
module Counter = struct
type t = int [@@deriving irmin]
let incr t = t + 1
let decr t = t - 1
let merge ~old t1 t2 =
let open Irmin.Merge.Infix in
old () >>=* fun old ->
let old = match old with None -> 0 | Some v -> v in
let diff1 = t1 - old in
let diff2 = t2 - old in
Fmt.pr "LCA: %i, Diff1: %i, Diff2: %i%!" old diff1 diff2;
Irmin.Merge.ok (old + diff1 + diff2)
let merge = Irmin.Merge.(option (v t merge))
end
module Store = Irmin_mem.KV.Make (Counter)
let info () = Store.Info.v (Unix.gettimeofday () |> Int64.of_float);;
Note that the merge function here contains a print statement that you might actually debug log (with Logs.debug) to show the two diffs whenever the merge function is called. This is so we can see what is happening later on.
From here we can recreate the scenario between Alice and Bob. We'll use different branches to represent multiple stores.
let alice_action s =
let* v = Store.get s [ "counter" ] in
let c =
Counter.incr v
|> Counter.incr
|> Counter.incr
|> Counter.incr
|> Counter.incr
in
Store.set_exn ~info s [ "counter" ] c
let bob_action s =
let* v = Store.get s [ "counter" ] in
let c =
Counter.decr v
|> Counter.decr
in
Store.set_exn ~info s [ "counter" ] c
Now for the main function which initialises the main store and applies both Alice's and Bob's actions and tries to merge them into the store.
# let config = Irmin_mem.config () in
(* Initialise a new empty store and add counter with value 10 *)
let* repo = Store.Repo.v config in
let* main = Store.main repo in
let* () = Store.set_exn ~info main [ "counter" ] 10 in
(* Create two new branches as clones of the [main] branch *)
let* alice_branch = Store.clone ~src:main ~dst:"alice" in
let* bob_branch = Store.clone ~src:main ~dst:"bob" in
(* Apply the actions *)
let* () = alice_action alice_branch in
let* () = bob_action bob_branch in
(* Merge the results *)
let* () =
let+ merge = Store.merge_into ~into:main ~info alice_branch in
Result.get_ok merge
in
let* () =
let+ merge = Store.merge_into ~into:main ~info bob_branch in
Result.get_ok merge
in
Store.get main [ "counter" ];;
LCA: 10, Diff1: -2, Diff2: 5
- : int = 13
The merge function was only needed once. When alice_action is applied, it is a simple "fast-forward" merge because there is no three-way merge required. However, when bob_action is applied there is now the LCA (the initial 10 value), Bob's new value (8) and Alice's value that has been merged (15).
JSON at Rest
As has been mentioned many times, Irmin is fundamentally a key-value store. Thanks to its portability and flexibility both in storage backend and data format, Irmin is not the only means by which to interact with the data.
Storing JSON values
We can instantiate a simple in-memory Irmin store that stores JSON objects.
module Store = Irmin_mem.KV.Make (Irmin.Contents.Json)
let info () = Store.Info.v (Unix.gettimeofday () |> Int64.of_float)
The type of JSON objects is identical to that of the Ezjsonm library. The objects are association lists (lists of pairs where the first pair is a string, like a dictionary in other programming languages).
# #show Irmin.Contents.Json.t;;
val t : Store.contents Repr.ty
type nonrec t = (string * Irmin.Contents.json) list
This is very convenient, and we can quickly get and set values directly in the store using JSON-like OCaml values. The fact that the Ezjsonm representation of JSON values and the Irmin representation are the same is no coincidence, however, there is no strict dependency between the two so they could change in the future.
# let set_json_string_exn s k v =
match Ezjsonm.value_from_string v with
| `O assoc -> Store.set_exn ~info s k assoc
| _ -> Lwt.fail (Failure "Expected a JSON object as a string");;
val set_json_string_exn : Store.t -> Store.path -> string -> unit Lwt.t =
<fun>
From here we can now add JSON objects directly into the store.
# let config = Irmin_mem.config () in
let* repo = Store.Repo.v config in
let* main = Store.main repo in
let* () = set_json_string_exn main [ "a" ] {|{ "hello": "world" }|} in
let+ s = Store.get main [ "a" ] in
print_endline @@ Ezjsonm.value_to_string (`O s);;
{"hello":"world"}
- : unit = ()
Custom Types Stored as JSON
One problem with using Irmin.Contents.Json.t is that we've lost the richness of the OCaml type system to a certain extent. This means it isn't obvious what are store is actually storing. Is it random JSON objects or a serialisation of a more rich OCaml value? If it is the latter, it probably isn't the interface we want.
For example, consider the following simple message datatype.
module type Message = sig
type t = string [@@deriving irmin]
include Irmin.Contents.S with type t := t
end
module Message : Message = struct
type t = string [@@deriving irmin]
let merge ~old:_ a b =
match String.compare a b with
| 0 ->
if Irmin.Type.(unstage (equal t)) a b then
Irmin.Merge.ok a
else
let msg = "Conflicting entries have the same timestamp but different values" in
Irmin.Merge.conflict "%s" msg
| 1 -> Irmin.Merge.ok a
| _ -> Irmin.Merge.ok b
let merge = Irmin.Merge.(option (v t merge))
end
By default if we create a store with this content type, the data will be stored using the string representation defined in repr. For the most part this is actually quite JSON-like.
# Irmin.Type.to_string Message.t "Hello World";;
- : string = "Hello World"
But there is an actual JSON-backend to the representation.
# Irmin.Type.to_json_string Message.t "Hello World";;
- : string = "\"Hello World\""
In fact for the most part the encoding does use JSON to format the OCaml values. The difference are usually very subtle, for example OCaml strings are just bytes whereas for JSON they must be UTF-8. So we can get very different formats (or as above where the JSON string requires the inverted-commas).
let _no_output_because_utf8 = Irmin.Type.to_string Message.t "\xc3\x28"
Whereas we must convert to a UTF-8 string that we can serialise and deserialise.
# Irmin.Type.to_json_string Message.t "\xc3\x28";;
- : string = "{\"base64\":\"wyg=\"}"
Fortunately, we can override the runtime representation of the type Irmin uses to store the values and keep the richness of the actual type when programming with the Irmin interface, but be serialising the data into JSON values. This is particularly useful, for example, with the Git.FS backend to read and write JSON values in Git stores.
module Message_json : Message = struct
type t = string [@@deriving irmin]
let merge ~old:_ a b =
match String.compare a b with
| 0 ->
if Irmin.Type.(unstage (equal t)) a b then
Irmin.Merge.ok a
else
let msg = "Conflicting entries have the same timestamp but different values" in
Irmin.Merge.conflict "%s" msg
| 1 -> Irmin.Merge.ok a
| _ -> Irmin.Merge.ok b
let t = Irmin.(Type.like ~pp:(Type.pp_json t) ~of_string:(Type.of_json_string t) t)
let merge = Irmin.Merge.(option (v t merge))
end
Care with Serialisation of your Types
One thing to watch out for with serialisation of types is that it won't be applied
recursively to all of your types. Consider the following defintion where we make
an indirection via a type alise and forget we haven't converted the name_t to use
JSON.
module Message_json = struct
type name = string [@@deriving irmin]
type t = name list [@@deriving irmin]
let merge ~old:_ a b =
match List.compare String.compare a b with
| 0 ->
if Irmin.Type.(unstage (equal t)) a b then
Irmin.Merge.ok a
else
let msg = "Conflicting entries have the same timestamp but different values" in
Irmin.Merge.conflict "%s" msg
| 1 -> Irmin.Merge.ok a
| _ -> Irmin.Merge.ok b
let t = Irmin.(Type.like ~pp:(Type.pp_json t) ~of_string:(Type.of_json_string t) t)
let merge = Irmin.Merge.(option (v t merge))
end
Now if we accidently used Message_json.name_t directly it won't be like Message_json.t.
# let msg = [ "\xc3\x28" ] in
let v = Fmt.str "%s" Irmin.Type.(to_string Message_json.t msg) in
let v' = Fmt.str "[%a]" Fmt.(list string) (List.map Irmin.Type.(to_string Message_json.name_t) msg) in
v = v';;
- : bool = false
This example is pretty contrived, but it is meant to just show the problem rather than an exact real-world example.
Versioned Data
Currently Irmin only supports monomorphic operations over stores, meaning there can
only be one return type from functions like Store.get. On the way to heterogeneity, a logical
first stop is being able to version content datatypes.
Versioning in this sense relates to the type and not the value. Irmin let's you version OCaml values, but we want to have different versions of the type of that value whilst still preserving key characteristics of the Irmin store.
Changing the Content Type
Before we do that, let's first look at how it can go wrong. First let's define two types that are meant to represent the same value but just one version is newer and has added new fields.
module C1 = struct
type t = { name : string } [@@deriving irmin]
let merge = Irmin.Merge.(option @@ default t)
end
module C2 = struct
type t = { name : string; age : int }[@@deriving irmin]
let merge = Irmin.Merge.(option @@ default t)
end
The only difference between the two types is the extra age : int field in C2.t. We can instantiate two Irmin
key-value stores that use the filesystem.
module S1 = Irmin_fs_unix.KV.Make (C1)
module S2 = Irmin_fs_unix.KV.Make (C2)
And now we can store a C1.t in the store and try and read it back using the S2 interface.
let main () =
let conf = Irmin_fs.config "./tmp" in
let* repo = S1.Repo.v conf in
let* main = S1.main repo in
let* () = S1.set_exn ~info main [ "a" ] C1.{ name = "Alice" } in
let* repo = S2.Repo.v conf in
let* main = S2.main repo in
let* v = S2.get main [ "a" ] in
Fmt.pr "Name: %s" v.name;
Lwt.return_unit
But this goes wrong if we try running it!
$ ./bad-content-store/main.exe
Fatal error: exception Irmin.Tree.find_all: encountered dangling hash 15942488a5c800c506817379631ace9263149cf42b0c8bc409a5a6e9698d6e5194ff6d178365e1642d9b5b29dee30ed18e5b4605c281a35db7d214428bfff510
[2]
Using Views
The following is courtesy of Thomas Gazagnaire.
Raw views
One way to fix this problem is to abstract the content type behind a view on it. This view let's us hide some details to the end user whilst giving us the power to do more complex manipulations of the data we are storing. To begin with, the view contains information about the version of the data.
type v1 = { age : int } [@@deriving irmin ~pre_hash]
type v2 = { age : int; name : string } [@@deriving irmin ~pre_hash]
type t = V1 of v1 | V2 of v2 [@@deriving irmin]
(* change depending on what you want to index - here we just skip the
version field. Can only hash the age (but that will merge values
with the same hash, so it's not a great index here). *)
let pre_hash = function
| V1 v -> pre_hash_v1 v
| V2 v -> pre_hash_v2 v
let t = Irmin.Type.like ~pre_hash t
This code handles serialising the version information so we can re-use that later without polluting
the content-addressable storage with fields we might not have later. For example, a user might only
habe their age and not be aware of the V1 version number, but they don't need to be aware of it
for us to do content-addressed lookups for the data.
However, we do need to manage smooth upgrades and downgrades from the versions but this is only verbose, not complicated.
let default_name = "Default Name"
let v1 age = { age }
let v2 ?(name = default_name) age = { age; name }
let v1_to_v2 : v1 -> v2 = fun { age } -> v2 age
let v2_to_v1 : v2 -> v1 = fun { age; _ } -> { age }
let to_v2 = function V1 v -> v1_to_v2 v | V2 v -> v
let to_v1 = function V1 v -> v | (V2 _) as v -> v2_to_v1 (to_v2 v)
let merge_v1 = Irmin.Merge.(default v1_t)
let merge_v2 = Irmin.Merge.(default v2_t)
Finally, we need to define a sufficient default merge function over versioned data.
let merge : t Irmin.Merge.t =
let open Lwt_result.Infix in
let promise x = Irmin.Merge.promise x in
let upgrade x = V2 (to_v2 x) in
let wrap_v1 v = V1 v in
let wrap_v2 v = V2 v in
let rec f ~old x y =
old () >>= fun old ->
match (old, x, y) with
| Some (V1 old), V1 x, V1 y ->
Irmin.Merge.f merge_v1 ~old:(promise old) x y >|= wrap_v1
| Some (V2 old), V2 x, V2 y ->
Irmin.Merge.f merge_v2 ~old:(promise old) x y >|= wrap_v2
| _ ->
let old =
match old with
| None -> fun () -> Lwt.return (Ok None)
| Some old -> promise (upgrade old)
in
f ~old (upgrade x) (upgrade y)
in
Irmin.Merge.seq [ Irmin.Merge.default t; Irmin.Merge.v t f ]
Higher-level Abstraction
With our raw views we can now provide a higher-level abstraction to use within our actual Irmin stores.
type 'a view = { payload : 'a; raw : Raw.t }
let version v = match v.raw with V1 _ -> 1 | V2 _ -> 2
module C1 = struct
type t = Raw.v1 view
let of_raw raw =
let payload = Raw.to_v1 raw in
{ payload; raw }
let to_raw t = t.raw
let t = Irmin.Type.map Raw.t of_raw to_raw
let v age =
let v = Raw.v1 age in
{ payload = v; raw = V1 v }
let merge = Irmin.Merge.(option @@ like t Raw.merge to_raw of_raw)
end
To be concise only V1 is shown here but as you can see it reuses all of the Raw values we defined
previously. We can always convert to and from an Raw.t and that is how the runtime value t is defined.
This means if we pull a V2 out of the store but the serialised version is a V1 we can still deserialise it.
We can use these Content modules for stores now.
module S1 = Irmin_fs_unix.KV.Make (C1)
module S2 = Irmin_fs_unix.KV.Make (C2)
And finally write a program where we interleave the stores, reading old and new values as we go!
let main () =
let config = Irmin_fs.config "./tmp2" in
let* repo = S1.Repo.v config in
let* main = S1.main repo in
let* repo2 = S2.Repo.v config in
let* main2 = S2.main repo2 in
let c1 = C1.v 42 in
let _h1 = S1.Contents.hash c1 in
let* () = S1.set_exn ~info:info1 main [ "a" ] c1 in
Fmt.pr "Storing S1 at a: { age = %i }\n%!" c1.payload.age;
let* v1 = S1.get main [ "a" ] in
Fmt.pr "S1 lookup a: %i (version = %d)\n%!" v1.payload.age (version v1);
let* v2 = S2.get main2 [ "a" ] in
Fmt.pr "S2 lookup a: %i (version = %d)\n%!" v2.payload.age (version v2);
let c2 = C2.v 43 ~name:"Alice" in
let* () = S2.set_exn ~info:info2 main2 [ "b" ] c2 in
Fmt.pr "Storing S2 at b: { age = %i; name = %s }\n%!" c2.payload.age c2.payload.name;
let* v = S2.get main2 [ "a" ] in
Fmt.pr "S2 lookup a: %s %i (version = %d)\n%!" v.payload.name v.payload.age
(version v);
let* v = S1.get main [ "b" ] in
Fmt.pr "S1 lookup b for age: %i\n%!" v.payload.age;
Lwt.return_unit
Let's run the program!
$ ./views/main.exe
Storing S1 at a: { age = 42 }
S1 lookup a: 42 (version = 1)
S2 lookup a: 42 (version = 1)
Storing S2 at b: { age = 43; name = Alice }
S2 lookup a: Default Name 42 (version = 1)
S1 lookup b for age: 43