feat: tiled grayscale rendering to drop the storeBwBuffer peak (#2106)
Tiled grayscale rendering to drop the storeBwBuffer peak largest contiguous free block (the value that actually drives OOM on the C3) from ~114 KB to ~82-90 KB. This renders each grayscale plane band-by-band into a small (~8 KB) scratch and streams each band straight to controller RAM (community-sdk writeGrayscalePlaneStrip), leaving the BW framebuffer intact. No save, no restore; controller RAM is re-synced for the next differential turn directly from the live framebuffer. Three writers honor the active band target so per-band re-rendering stays cheap and correct: - drawPixel (text) redirects writes to the band scratch and clips to it. - renderCharImpl skips glyphs whose physical y-extent is outside the band before the bitmap decode (glyphIntersectsStrip), so the per-band re-render doesn't pay N x glyph decode. - DirectPixelWriter (images) writes the band scratch via getWriteTarget instead of the framebuffer. Without this, image pixels wrote the live BW frame directly and cleanup re-synced that corruption, leaving thin outlines after navigating away from an image. Controller specifics live in the SDK (X4 setRamArea windowing, X3 PTL); the reader checks supportsStripGrayscale() and is otherwise controller-agnostic. Measured on hardware (X4 and X3, text and images, visually correct): - Grayscale scratch ~8 KB vs ~50 KB save; largest contiguous free block held at full size during grayscale instead of dropping ~25-32 KB. - X4 text page about +25 ms/page; X3 page time is dominated by its intrinsic grayscale refresh, not tiling. Depends on community-sdk #13 (the writeGrayscalePlaneStrip API). The submodule bump here points at that branch, so until #13 merges the submodule won't resolve from upstream and CI will fail there; keeping this a draft until then. Will rebase onto master and re-point the submodule to the merged SDK commit once #13 lands. Did you use AI tools to help write this code? partial
This commit is contained in:
@@ -16,6 +16,12 @@ struct DirectPixelWriter {
|
||||
uint8_t* fb;
|
||||
GfxRenderer::RenderMode mode;
|
||||
uint16_t displayWidthBytes; // Runtime framebuffer stride (X4: 100, X3: 99)
|
||||
// Active write target: for tiled grayscale, fb is the band scratch, originY is
|
||||
// the band's top physical row, and clipRows is the band height. Off-band
|
||||
// pixels are dropped. With no strip active these collapse to the full frame
|
||||
// (originY 0, clipRows panelHeight) so the clip doubles as a bounds guard.
|
||||
int originY;
|
||||
int clipRows;
|
||||
|
||||
// Orientation is collapsed into a linear transform:
|
||||
// phyX = phyXBase + x * phyXStepX + y * phyXStepY
|
||||
@@ -28,7 +34,9 @@ struct DirectPixelWriter {
|
||||
int rowPhyXBase, rowPhyYBase;
|
||||
|
||||
void init(GfxRenderer& renderer) {
|
||||
fb = renderer.getFrameBuffer();
|
||||
fb = renderer.getWriteTarget();
|
||||
originY = renderer.getWriteOriginY();
|
||||
clipRows = renderer.getWriteRows();
|
||||
mode = renderer.getRenderMode();
|
||||
displayWidthBytes = renderer.getDisplayWidthBytes();
|
||||
|
||||
@@ -120,7 +128,12 @@ struct DirectPixelWriter {
|
||||
const int phyX = rowPhyXBase + logicalX * phyXStepX;
|
||||
const int phyY = rowPhyYBase + logicalX * phyYStepX;
|
||||
|
||||
const uint16_t byteIndex = phyY * displayWidthBytes + (phyX >> 3);
|
||||
// Band-local row. The unsigned compare drops both off-band pixels (strip
|
||||
// mode) and any out-of-frame row (full-frame mode) in one branch.
|
||||
const int sy = phyY - originY;
|
||||
if (static_cast<unsigned>(sy) >= static_cast<unsigned>(clipRows)) return;
|
||||
|
||||
const uint16_t byteIndex = static_cast<uint16_t>(sy * displayWidthBytes + (phyX >> 3));
|
||||
const uint8_t bitMask = 1 << (7 - (phyX & 7));
|
||||
|
||||
if (state) {
|
||||
|
||||
Reference in New Issue
Block a user