Skip to content
All lab notes

Lab note · from SwipeClean

Laplacian sharpness always 0 on iOS: Core Image and Vision traps

By Mark Santos · · 5 min read

SwipeClean's photo quality score was 0 for every photo because all three of its inputs could come back as zero without an error. The Laplacian sharpness pass read its result through CIContext.render(_:toBitmap:rowBytes:bounds:format:colorSpace:) with .RGBA8 and a one-channel grey colour space, a call that wrote nothing at all in our runs (and an 8-bit format would clamp the Laplacian's negative values anyway), while the saliency and face-detection requests threw on the iOS Simulator and the code caught those errors as 0 and "no faces". The fix renders to single-channel float (.Rf) and has each Vision component return nil when it didn't run, so the score is renormalised over what was actually measured.

Symptoms

  • In Similar Photos, the quality badge read "0%" on all four photos in a group, including the one marked Best. The badge is Int(score.qualityScore * 100), which truncates, so anything under 0.01 prints as 0%. "Best" is simply the first photo in the sorted group, so with every score tied it was whichever photo sorted first.
  • Smart Cleanup keeps photos scoring below 0.35 and sorts them worst first. With every score at 0, every photo analysed qualified, in no meaningful order.
  • Nothing crashed and nothing was logged. Every failure path returned a number.

The work was being done on the Simulator; the project had no signing team for a device yet.

Why it happens

The composite was sharpness * 0.3 + saliency * 0.4 + (hasFaces ? 0.3 : 0). All three terms were zero, for three different reasons.

The sharpness read-back was empty

// Before (PhotoAnalyzer.computeSharpness)
var pixelData = [UInt8](repeating: 0, count: height * bytesPerRow)
ciContext.render(scaledImage, toBitmap: &pixelData, rowBytes: bytesPerRow,
                 bounds: scaledExtent, format: .RGBA8,
                 colorSpace: CGColorSpaceCreateDeviceGray())
// ...variance of pixelData[i * 4], divided by 2500

.RGBA8 has four components and a device-grey colour space has one. To see what Core Image does with that, we filled the buffer with the value 7 first, then rendered a mid-grey image with this exact call. The buffer still held 7 afterwards: nothing had been written. The same call with a device-RGB colour space wrote 128s, and .L8 with device grey wrote 127 on the Simulator and 128 on the Mac. This was on the iOS 26.5 Simulator and on macOS 27.2. render(...toBitmap:...) returns nothing and doesn't throw, so the zero-filled array stayed zero, its variance was 0, and so was the sharpness. Apple's documentation for the method doesn't describe this case, so treat it as observed behaviour, not specified behaviour.

Eight bits would have dropped the negative half anyway

The kernel [0, 1, 0, 1, -4, 1, 0, 1, 0] sums to zero, so its output is signed: an edge produces a positive lobe on one side and a negative lobe on the other. Apple documents RGBA8 as a fixed-point format. Rendering a constant −0.25 gave a byte of 0 in .RGBA8 and −0.25 in .Rf, "a 32-bit-per-pixel, floating-point pixel format in which the sole component is a red color value".

We also re-ran the old function with a device-RGB colour space in place of the grey one, changing nothing else. It scored a checkerboard at 0.059 and a flat grey field at 0.164, so the image with no detail at all scored higher. The flat field's score came entirely from the image border, where a 3×3 kernel has no real neighbours.

Vision threw, and the throw became a zero

// Before (PhotoAnalyzer.computeSaliency)
do {
    try handler.perform([request])   // VNGenerateAttentionBasedSaliencyImageRequest
    if let observation = request.results?.first as? VNSaliencyImageObservation,
       let salientObjects = observation.salientObjects, !salientObjects.isEmpty {
        return salientObjects.map(\.confidence).max() ?? 0
    }
} catch {
    // Saliency detection failed.
}
return 0

Re-run for this note on the iOS 26.5 Simulator, VNGenerateAttentionBasedSaliencyImageRequest throws "Failed to create espresso context." and VNDetectFaceRectanglesRequest throws "Could not create inference context". Both functions caught the error and returned 0 or false, which looks exactly like "nothing interesting here" or "no faces here". That zeroed 70% of the score's weight before sharpness was even counted. The project's worklog records the same trap: "Any code path that treats a Vision failure as a zero result will look fine locally and be wrong on device, or vice versa."

Not the cause: salientObjects

An earlier write-up of this bug, including the comment in the fixed code, says an attention-based request never populates salientObjects. Apple's Cropping Images Using Saliency says otherwise: "Attention-based saliency requests return only one bounding box", and object-based requests "return up to three bounding boxes". On macOS 27.2, the attention request returned exactly one box for each of four test images, with confidences from 0.495 to 0.607. On the Simulator, the code never got that far, because perform had already thrown.

The fix

Commit 01c2a18 changes both paths. Sharpness is now read back as one float per pixel, with the border trimmed off:

let scaledExtent = scaledImage.extent.insetBy(dx: 2, dy: 2)
var pixelData = [Float](repeating: 0, count: width * height)
ciContext.render(scaledImage, toBitmap: &pixelData, rowBytes: width * 4,
                 bounds: scaledExtent, format: .Rf, colorSpace: nil)
// variance of pixelData, divided by 0.02 (Core Image works in 0-1, not 0-255)

.Rf matches what the code wants, which is one channel. It keeps the negative lobe and removes the format and colour-space mismatch. The 2-pixel inset drops the ring where the kernel reads edges that don't exist.

The Vision components now say when they didn't run:

func detectFaces(in cgImage: CGImage) -> Bool? {
    let request = VNDetectFaceRectanglesRequest()
    let handler = VNImageRequestHandler(cgImage: cgImage, options: [:])
    do { try handler.perform([request]) } catch { return nil }
    guard let results = request.results else { return nil }
    return !results.isEmpty
}

The composite renormalises over whatever is present:

var quality: Float {
    var weighted = sharpness * Weight.sharpness
    var total = Weight.sharpness
    if let saliency { weighted += saliency * Weight.saliency; total += Weight.saliency }
    if let hasFaces { weighted += (hasFaces ? 1 : 0) * Weight.faces; total += Weight.faces }
    guard total > 0 else { return 0 }
    return min(max(weighted / total, 0), 1)
}

A missing component now leaves the denominator instead of adding a zero to the numerator. Sharpness alone can then span 0 to 1, where before it was capped at 0.3.

Saliency also changed what it measures. It is now the mean of the observation's heat map, not the box's confidence. Apple describes that map as "a CVPixelBuffer in a one-component floating-point pixel format", and the code reads it row by row using CVPixelBufferGetBytesPerRow, because rows can be padded.

How to check you've fixed it

Tests/SwipeCleanTests/RegressionTests.swift holds PhotoAnalyzerScoringTests:

  • A checkerboard scores above 0.01 for sharpness.
  • A blurred copy scores lower than the original.
  • A flat grey image scores below 0.005.
  • Saliency is either nil or within 0–1.
  • A detailed image's composite is above 0.005.
  • The renormalisation cases hold: sharpness 1 with nothing else gives 1, and sharpness 0 with nothing else gives 0.

The saliency test asserts a contract, not a value, because Vision can't run on the Simulator. The commit records 140 tests passing.

To catch this class of bug directly, pre-fill the bitmap with a sentinel value before render and assert that it changed. A render that silently writes nothing then fails loudly instead of looking like a flat image.

Caveats

  • SwipeClean targets iOS 16.0. The saliency requests date from iOS 13, and .Rf from iOS 9.
  • The Simulator and macOS observations here were made with Xcode 27.0. No device was available for this check, so whether the old read-back is also empty on hardware is unverified. If it is, the old score on a device would still have carried no sharpness information, even though Vision would have run.
  • The heat-map mean is a small number. On the four macOS test images it ranged from 0.04 to 0.09. The saliency term therefore adds a narrow band to the score, and its 40% weight hasn't been re-tuned for the new measure.
  • PhotoScore still turns a missing component into 0 or false in its per-component fields (hasFaces: analysis.hasFaces ?? false). Only the composite renormalises.

More lab notes