Skip to content

mizu_map() maps a function over a vector or list on the pool and returns the results in input order. It is not a loop over mizu_submit(). The function, its constant arguments, and the data are staged once in shared memory. A handful of chunk tasks divide the elements, and each worker sets up the map at most once. The per-element cost approaches the cost of lapply(), while the work spreads across workers and balances itself through stealing.

library(mizu)

p <- mizu_pool(n_workers = 4L)
mizu_map(p, 1:5, \(i) i * 2L)
#> [[1]]
#> [1] 2
#> 
#> [[2]]
#> [1] 4
#> 
#> [[3]]
#> [1] 6
#> 
#> [[4]]
#> [1] 8
#> 
#> [[5]]
#> [1] 10

Templates

A .template (in the style of the FUN.VALUE argument of vapply()) returns an atomic vector or matrix instead of a list. The workers write the results straight into shared memory: each result crosses the process boundary exactly once, unserialized:

mizu_map(p, seq.int(-5, 5), abs, .template = numeric(1))
#>  [1] 5 4 3 2 1 0 1 2 3 4 5

With .collect = "view", the result is the shared output area itself, wrapped as a read-only ALTREP view: the gather copy is skipped, and the region’s teardown defers to the view’s garbage collection.

An integer64 template (bit64’s layout, constructed without bit64) is also supported. int64 joins no widening lattice, so results must be exact integer64 of the template’s length.

Reproducible randomness

Random numbers drawn inside the function are not reproducible by default, and cost nothing extra. Pass .seed to give every element its own L’Ecuyer-CMRG stream. Then the results are identical for any chunking, worker count, or steal order:

identical(mizu_map(p, 1:4, \(i) rnorm(i), .seed = 123L),
          mizu_map(p, 1:4, \(i) rnorm(i), .seed = 123L, .chunks = 4L))
#> [1] TRUE

Prepared maps

mizu_map_prepare() stages a map — the function, its constant arguments, and the data — without running it: the serialization and the region create are paid once. Each mizu_map_run() then costs only task submission and collection, re-arms in O(1), and reuses the cached map context of each worker. Repeated stochastic simulation is the headline use: .seed rides each run rather than the region, so the random streams vary per run for free:

pm <- mizu_map_prepare(p, 1:1000, \(i, draws) mean(rnorm(draws)) * i, draws = 100L)
runs <- lapply(1:4, \(s) mizu_map_run(pm, .seed = s))
vapply(runs, \(r) r[[1L]], numeric(1))
#> [1] 0.13109678 0.08153219 0.17986914 0.17683176

mizu_map_run(pm, .x = x2) replaces the staged data: an atomic vector of the same type and length swaps in place at memcpy cost; any other replacement restages transparently. A run collected with .collect = "view" hands its region to the view, so the next run restages into a fresh one. The handle pins the staged data and the region for its lifetime; both release at garbage collection.

Nested maps

A map can nest: mizu_map(mizu_current_pool(), ...) inside a task expression uses the evaluating worker’s own handle. The runner submissions push straight onto the worker’s own deque, and the blocked collect executes its own runners while idle peers steal the rest: nested maps never deadlock the pool.

t <- mizu_submit(
  p,
  sum(unlist(mizu_map(mizu_current_pool(), parts, sum))),
  parts = split(as.numeric(1:100000), rep(1:10, each = 1e4))
)
mizu_collect(t)
#> [1] 5000050000

The first nested map of a worker claims a submitter slot, so at the default max_submitters = 8 — one held by the controller — at most 7 workers can nest concurrently. Raise max_submitters for wider nested fan-outs.