Darktable uses the à-trous wavelet transform for denoising images converted into the Y0U0V0 format. The Y0U0V0 seems to be just a simple linear manipulation of the raw camera pixel values.
From the source code, it looks like a variance stabilizing transform (VST) is applied to the debayered camera-raw-RGB pixel values, then this is converted to Y0U0V0 and finally the wavelet decomposition begins. This happens in the precondition_Y0U0V0() function in denoiseprofile.c, if I’m not mistaken.
It looks to me like all of this (the VST) is done just so that the BayesShrink algorithm can easily compute valid noise variance of the “detail” wavelet coefficient image. This also seems to be the reason why float sigma = 1.0; in variance_stabilizing_xform() in denoiseprofile.c.
Based on my limited understanding, the VST is using the RAW image’s Poisson noise profile coefficients “a” and “b”, which are applicable only on the linear raw pixel values and perhaps also to their linear transformations (e.g. Y0U0V0)?
I feel like by converting from the raw RGB camera values to the Lab color space, it would no longer be possible to apply VST on the debayered RGB L, a, b channels (because the conversion from RGB to Lab applies non-linear operations on the pixel values), therefore the wavelet transform wouldn’t work as good as it does now with unit variance because the variances of the individual channels would no longer be 1.0?
A tangential question: What would happen if darktable didn’t use VST and kept different pixels having different noise variance? Would there be more noise in the “denoised” image? Would it be possible to get meaningful results when denoising the a and b chroma channels of the Lab color space when not using VST?
Could someone more knowledgable illuminate me and correct my wrong assumptions, please?
in general your observations are all correct. Lab is non-linear, so it’ll not stabilise variance the same way (and yes, noise sigma=1 after the transform). yuv seemed like an improvement, since it separates luminance and chroma. the colour dimensions you can denoise much more aggressively without people perceiving it. there’s also a matter of taste: some people hate colour noise but don’t mind monochrome grain.
applying the gauss/poisson noise model to data that is a linear transform away from the actual sensor data can still be worked out. also note that darktable’s noise profiles are measured within the pipeline, at some convenient spot. if things didn’t change over the years, this used to be after black point subtraction/demosaicing. not great, clear disadvantages.
i did experiment with wavelets in Lab, but besides theoretical concerns it simply didn’t work so well.
leaving away the stabilising transform is not a good idea. also note that even neural network denoisers often use a thing like this. in place of messy/slower batch normalisation, this keeps values in a lukewarm range.
so: yes you can denoise in Lab, i’m going to bet that it’s not better than yuv (which also separates luminance and the rest but doesn’t have problems with nonlinearity and negative values). yes, you can do without stabilising transform, but bayesshrink will be wrong/suboptimal and neural networks also profit from highlight-compressing transforms (so it’s probably not a great idea either).
Thank you for your time in answering my questions, it’s an honor to get a reply from someone with your skills!
What you wrote makes sense to me, but I continued to dig into the source code and couldn’t make sense of some of the choices made there. Would you be willing to point me to the right direction please? No problem if you don’t want to / don’t have time for that, I completely understand :-)!
For the sake of simplification, let’s assume that a camera would capture a RAW image with (1, 1, 1) white balance gains, so the RAW image would contain unscaled pixel values straight from the sensor. In such scenario:
Would it still be necessary to call set_up_conversion_matrices(), which performs white balance-related scaling of the toY0U0V0 matrix? With unit WB gains, wouldn’t it be enough to simply use the normal toY0U0V0 conversion matrix?
I assume that precondition_Y0U0V0() performs something similar to the generalized Anscombe transform (GAT) on the raw RGB camera pixel values, but the formula in the source code is not exactly the same as the GAT formula (3) in this paper. There are the p variables (p for “power”?) and also sqrt is applied on the a noise parameter, which GAT doesn’t do. Maybe the a and b noise parameters used in Darktable have different units / meaning then the a and b noise parameters in the NoiseProfile DNG tag though?
Under the assumption of processing only images captured with unit white balance gains, this is my understanding of the pipeline; do you think it’s correct?
Load RAW pixels, subtract the black level, normalize RGB pixel values to 0-1.
Debayer.
Apply the generalized Anscombe transform using the Poisson/Gaussian a and b parameters.
Convert to Y0V0Y0 using the toY0U0V0 matrix.
Perform wavelet decomposition, filter the wavelet detail coefficients using the BayesShrink soft-threshold algorithm with the assumption of unit noise sigma (thanks to the GAT). Accumulate the filtered wavelet detail coeffs.
Apply the inverse of GAT, and then apply the inverse of toY0U0V0. This will produce denoised camera RGB.
Continue the RAW processing pipeline with denoised RGB values. Maybe apply custom white balance gains?
i’m not familiar with the “recent” changes in dt code, can’t comment on individual functions and what they do.
the noise profile values are very different from the dng spec, yes. applying all this after black point subtraction and debayering is suboptimal.
in your recipe, i’d swap 3 and 4, at least 4 will work as intended then. it’s probably possible to transform the moments defining your variance through 4 and then apply a transformed anscombe there (or profile it there in the first place).
in vkdt i do the denoising in raw and linear, and only use the perceptual yuv for blending in post. i think that’s the safe way of doing it.