Skip to content
All lab notes

Lab note · from Carbon

Per-app power on Apple silicon: Mach ticks and ri_energy_nj

By Mark Santos · · 5 min read

On Apple silicon, the CPU times that proc_pidinfo(PROC_PIDTASKINFO) returns in pti_total_user and pti_total_system are Mach absolute-time ticks, not nanoseconds. Multiply them by mach_timebase_info's numer / denom, which was 125/3 on the M5 Max we tested, or CPU time comes out 41.7 times too small. And for power there is a better input than CPU time × TDP. Since macOS 13, proc_pid_rusage(pid, RUSAGE_INFO_V6, …) returns ri_energy_nj, the kernel's running estimate of a process's CPU energy in nanojoules, and an ordinary user can read it for the same processes proc_pidinfo can.

Symptoms

Carbon is a menu bar app that estimates how many watts each app draws. Every 5 seconds it samples every process's CPU time and multiplies its share of the CPU by a TDP figure for the chip. We measured both ways that goes wrong on an M5 Max running macOS 27.2:

  • Unconverted ticks. A thread that spun for 2 seconds added 47,028,061 to its pti_total_user. Read as nanoseconds that is 0.047 s, about 2.4% of one core. Converted, it is 1.9595 s, the same as getrusage reported.
  • Right units, wrong watts. Carbon's model gives every fully busy core 30 W / 18 cores = 1.67 W on this chip. The kernel's counter gave 2.4 to 2.6 W for a spinning C loop and 3.3 to 3.5 W for a spinning Swift loop, each one thread for 2 seconds. Across processes using more than 5% of a core, energy per CPU-second ranged from 0.66 to 4.87 J over three 10-second windows, a 5× to 7× spread within each. In one 5-second sample, mediaanalysisd at 101% CPU drew 0.81 W while a terminal app at 37% drew 1.76 W. Ranking by CPU share puts them the wrong way round.

Why it happens

Task CPU times use the Mach timebase

<sys/proc_info.h> describes pti_total_user only as "total time". Apple's guide to architectural differences says to "Always apply timebase information to values you receive from mach_absolute_time and never assume that the function returns the number of nanoseconds since boot." Our test shows these fields use the same units. On Intel, xnu's clock_timebase_info sets numer = denom = 1 (osfmk/i386/rtclock.c), so dividing by 1e9 happens to work there. On Apple silicon the tick is 24 MHz (hw.tbfrequency is 24000000), and 1e9 / 24e6 = 125/3.

Rosetta adds a trap. An x86_64 build of the same test reported a timebase of 1/1, yet pti_total_user still grew by 47,014,615 in 1.96 s, which is native ticks. In a translated process mach_timebase_info is the wrong scale for these fields, while hw.tbfrequency read 24000000 in both builds.

CPU time is not energy

A CPU-second costs different amounts of energy depending on the core, its frequency and the instructions running. CPU time multiplied by a constant cannot see any of that. Apple's tools don't report per-process watts either. Activity Monitor's Energy Impact is "A relative measure of the current energy consumption of the app (lower is better)" (Activity Monitor guide). man powermetrics calls its per-process energy impact number "a rough proxy for the total energy the process uses, including CPU, GPU, disk io and networking", and the tool exits with "powermetrics must be invoked as the superuser" unless run as root.

The kernel does keep a per-process CPU energy figure. rusage_info_v6 in <sys/resource.h> has ri_energy_nj and ri_penergy_nj, with no comment on either. In xnu-8792.41.9, the context-switch hook passes a thread's energy_estimate_nj, a field "populated by CLPC and used to update the energy estimate of the thread", to recount_add_energy, which adds it to the thread and its task. The underlying counter is flagged HAS_CPU_DPE_COUNTER, "Has a hardware counter for digital power estimation". Apple's release manifests list xnu-8792.41.9 for macOS 13.0, and the macOS 12.5 kernel, xnu-8020.140.41, has no RUSAGE_INFO_V6.

So ri_energy_nj is an estimate too, and nothing in that path reads the GPU, but it comes from the kernel, per process, off a hardware counter. Carbon requires macOS 14 and does not read it.

The fix

The minimal broken version:

let cpuSeconds = Double(info.pti_total_user + info.pti_total_system) / 1_000_000_000  // 41.7x low on Apple silicon

Carbon's MachTimeConverter.swift gets the units right:

public init() {
    var info = mach_timebase_info_data_t()
    mach_timebase_info(&info)
    numer = UInt64(info.numer)
    denom = UInt64(info.denom)
}

/// Convert Mach absolute time ticks to seconds.
public func seconds(fromTicks ticks: UInt64) -> Double {
    let nanos = (ticks &* numer) / denom
    return Double(nanos) / 1_000_000_000
}

ProcessEnergyTracker feeds it the difference between two samples of pti_total_user + pti_total_system, divides by the elapsed wall time, and turns the share into watts with (cpuPercent / (cores × 100)) × tdpWatts.

To replace the TDP step, read the kernel's counter. This is our test code, not in the repository:

/// Cumulative CPU energy for a process, in nanojoules (macOS 13+).
func energyNanojoules(pid: pid_t) -> UInt64? {
    var info = rusage_info_v6()
    let rc = withUnsafeMutablePointer(to: &info) {
        $0.withMemoryRebound(to: rusage_info_t?.self, capacity: 1) {
            proc_pid_rusage(pid, RUSAGE_INFO_V6, $0)
        }
    }
    return rc == 0 ? info.ri_energy_nj : nil
}

Sample it twice and divide the difference by the elapsed seconds times 1e9 to get watts. It needs no table of chips. Carbon's table gives 10, 20, 30 or 60 W by tier (base, Pro, Max, Ultra), the same for every generation.

How to check you've fixed it

  • sysctl kern.pervasive_energy prints 1 when the kernel keeps this accounting. It did on our M5 Max, both natively and under Rosetta.
  • Spin one thread for 2 seconds and compare your converted pti_total_user difference with getrusage(RUSAGE_SELF)'s ru_utime. They should agree. Ours were both 1.9595 s.
  • Print mach_timebase_info in the build you ship. Our native build printed 125/3 and our x86_64 build printed 1/1.
  • Try proc_pid_rusage on pid 1 as a normal user and you get EPERM. Across the whole process list, proc_pid_rusage and proc_pidinfo succeeded for exactly the same PIDs (1,763 of 2,115 in one pass) and both failed with EPERM for the rest.

Caveats

  • GPU and display. Don't expect ri_energy_nj to include GPU work. Carbon takes GPU load from the IOAccelerator's PerformanceStatistics "Device Utilization %", multiplies it by a tier GPU TDP and splits it across a hard-coded list of 18 GPU-heavy bundle IDs in proportion to their CPU. For the display it reads brightness from IODisplayConnect, but ioreg listed no such service on our test Mac, so Carbon's display estimate stays at its fallback of 0.5 × 7 W = 3.5 W.
  • No battery check. Nothing in the repository reads IOPMPowerSource or battery discharge, so no measured total checks the per-app sum.
  • Carbon intensity. The grid figures are a hard-coded table of 47 countries, commented "approximate national averages", with no source and a fallback of 440 gCO2/kWh. Ember's yearly electricity data would be a sourced replacement. Its data page states "All content is released under a Creative Commons Attribution Licence (CC-BY-4.0)", so bundling it requires attribution.
  • History. The view model writes to SQLite once every 300 seconds, storing that moment's per-app watts with durationSeconds: 5.0. Daily and weekly totals therefore cover about one sixtieth of the time, and display watts are never stored.
  • Platform. Carbon targets macOS 14, its build script compiles arm64 only, and it is not yet released.

More lab notes