stephen wolff

Inverse Foley

working note research

Part of the series Watch instrument

So, I was on a paddleboard the other day, off Studland Beach, Dorset, on the way to Old Harry’s Rocks, and I was wearing a Watch (capitalised by Apple). I often find myself singing a song or whistling while in motion, and this time I had an idea: could I use the data from the Watch in some way with its many sensors to make music?

When we play, the instrument isn’t usually our body; except if it is our voice (and pats on the head/tummy). We usually play an instrument directly with our hands (Bongo, cajón etc) or blow into it (wind, brass etc). Sometimes there is a mediator (usually hands, can be feet) like via a stick, a bow or a pedal.

From my understanding, gestural instruments are usually discussed structurally, ie which sensor drives which parameter, one-to-many or otherwise. I was wondering how that relationship might work meaningfully between the gesture and sound.

Take a stick: it turns a swing into an impulse. A bow turns a pull into sustained friction. Would a listener think “ah, that’s a string”, or would they hear bowing? The mediator transposes the cause, and the ear reads through it.

While working on this I found the term “everyday listening”, coined by William Gaver in the late 1980s. We hear the event, not the acoustics - you know a glass is nearly full when it is struck, without ever listening for the pitch. Gaver sorts these events by how they’re made: struck, blown, poured.

You might say: but isn’t every electronic instrument causal? Sure - pressing a key closes a circuit, code runs (or the modular electronics flow), and a speaker cone moves. So causation was never the distinction. The drum’s sound specifies the strike, written in by physics; the synth’s specifies the algorithm. Every instrument is causal; almost none are causally legible.

Physical causation moves things in bundles - hit harder and loudness, brightness and attack shift together in ratios the material fixes. A mapping moves one parameter, or a designer’s chosen few. Thin bundle, and the ear answers “a synthesiser.”

Film has the same severed chain and doesn’t reconnect it. Foley reconstructs specifications well enough that nobody notices the sounds are made by different objects. Invert it: in film that reconstruction hides a substitution; for the body, in real time, it does the opposite and carries the cause. What crosses the gap is a description of the cause, not a control number. The same three verbs at both ends.

Right now, the prototype is only the first link: a Watch streaming its sensors to an iPhone. Collecting that data, and training a model to recognise the cause, is the work still ahead.

Over the next few posts, I’ll build it in the open: collecting the data, training the encoder, and seeing where it breaks.