fix: improve KOSync bidirectional position matching accuracy (#1897)
## Summary **Goal:** Fix bidirectional KOSync position matching between CrossPoint and KOReader so that syncing in either direction lands on the correct page with character-level accuracy. **Changes included:** **Download — `toCrossPoint` (server XPath → CrossPoint page)** - **XPath ancestry mode for structured elements**: The previous `ParagraphStreamer` only tracked `<p>` elements. Replaced with a full ancestor-walking mode that correctly resolves XPaths pointing into `<li>`, `<ul>`, and other structured elements. Char offset within the target element is bounded to the matched element's content only. - **Slash-in-attribute-value corrupts depth tracking**: `processByteInTag()` treated every `/` byte as a self-closing tag marker, including `/` inside quoted attribute values (e.g. `xmlns="http://..."`, `src="Links/image.jpg"`). This drove `htmlDepth` to 0 prematurely, causing the ancestry search to exit far short of the target paragraph. Fixed with `inAttrQuote` tracking. - **Off-by-one in page formula**: `intra * totalPages` rounds up incorrectly for last-page positions. Changed to `intra * (totalPages - 1)` to map the `[0, 1]` intra fraction correctly onto the `[0, totalPages-1]` page range. Example: page 14 of 17 was returned as 15. **Upload — `toKOReader` (CrossPoint page → server XPath)** - **Off-by-one in page-to-intra formula**: Symmetric fix — `pageNumber / totalPages` changed to `pageNumber / (totalPages - 1)`, with the guard updated from `> 0` to `> 1` to avoid division by zero. - **`<li>`-based XPath generation**: When the current page starts on a list item, `findXPathForProgress` now generates `ul[N]/li[M]` XPaths rather than falling back to the preceding `<p>`. Requires the new `listItemIndex` field in `PageLutEntry` (section cache version bumped to 23). - **Text-node precision with correct `text()[N].M` format**: KOReader expects `text()[N].M` where `N` is the 1-based index of the specific text node within the element. The previous attempt generated `text().M` (no brackets), which caused KOReader to jump to the front of the book. Implements a per-element text-node index stack in `XPathProgressResolver` — parallel to the existing element path stack — that correctly tracks text node indices relative to each element. Empty text nodes from bare anchor elements (`<a id="anchor"/>`) are intentionally skipped, matching KOReader's own text node counting behavior. **Reviewer-caught bugs** - **Double `onCloseTag()` on malformed `</br/>`**: Both the `tagIsClose` path and the self-closing `/` check were firing, double-decrementing `htmlDepth`. Fixed with a `!tagIsClose` guard. - **Dangling pointer in `LOG_DBG`**: `std::to_string(*nextParagraphPage).c_str()` passed a pointer to a temporary destroyed before the variadic call. Fixed with `snprintf` into a stack `char[8]` buffer. ## Additional Context - Section cache version bumped from 22 → 23 due to the new `listItemIndex` field in `PageLutEntry`. Users upgrading will see a one-time re-render of all cached sections on first load — no data loss. - The `textNodeIndexStack` in `XPathProgressResolver` is a `std::vector<int>` that mirrors the existing `path` and `parentStates` stacks — same depth, same lifetime. No additional heap pressure beyond what was already present. - All fixes verified on device with *Gentle and Lowly* by Dane C. Ortlund (spine 21, 17 pages). Download syncs land on the correct page; upload syncs land at the correct paragraph with character-level offset. ## Test plan - [ ] Download: sync from KOReader → CrossPoint lands on correct page for `text()[N].M` XPaths - [ ] Download: ancestry correctly resolves `<li>` positions inbound from KOReader - [ ] Upload: sync from CrossPoint → KOReader lands within one page for mid-paragraph positions - [ ] Upload: sync from CrossPoint → KOReader correctly targets `<li>` elements when page starts on a list item - [ ] Upload: `text()[N].M` format XPaths do not cause KOReader to jump to front of book - [ ] Section cache version 23: delete `.crosspoint/` and verify clean re-parse with no crashes --- ### AI Usage Did you use AI tools to help write this code? **YES** — developed with Claude Code (Anthropic). --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Sonnet 4.6
parent
628794f8b8
commit
bf894fd343
@@ -49,13 +49,14 @@ struct PathSegment {
|
||||
int index;
|
||||
};
|
||||
|
||||
std::string buildParagraphXPath(const int spineIndex, const std::vector<PathSegment>& path, const int charOffset) {
|
||||
std::string buildParagraphXPath(const int spineIndex, const std::vector<PathSegment>& path, const int textNodeIndex,
|
||||
const size_t charOffset) {
|
||||
std::string xpath = "/body/DocFragment[" + std::to_string(spineIndex + 1) + "]/body";
|
||||
for (const auto& segment : path) {
|
||||
xpath += "/" + segment.name + "[" + std::to_string(segment.index) + "]";
|
||||
}
|
||||
if (charOffset > 0) {
|
||||
xpath += "/text()." + std::to_string(charOffset);
|
||||
if (textNodeIndex > 0 && charOffset > 0) {
|
||||
xpath += "/text()[" + std::to_string(textNodeIndex) + "]." + std::to_string(charOffset);
|
||||
}
|
||||
return xpath;
|
||||
}
|
||||
@@ -280,7 +281,7 @@ class XPathParagraphResolver final : public Print {
|
||||
if (name == "p") {
|
||||
paragraphCount++;
|
||||
if (paragraphCount == targetParagraph) {
|
||||
xpath = buildParagraphXPath(spineIndex, path, 0);
|
||||
xpath = buildParagraphXPath(spineIndex, path, 0, 0);
|
||||
stopped = true;
|
||||
XML_StopParser(parser, XML_FALSE);
|
||||
}
|
||||
@@ -410,10 +411,14 @@ class XPathProgressResolver final : public Print {
|
||||
const int siblingIndex = parentStates.back().nextIndex(name);
|
||||
path.push_back({name, siblingIndex});
|
||||
parentStates.emplace_back();
|
||||
textNodeIndexStack.push_back(0);
|
||||
pendingTextNode = true;
|
||||
|
||||
if (name == "p") {
|
||||
paragraphDepth++;
|
||||
paragraphVisibleChars = 0;
|
||||
}
|
||||
if (name == "li") {
|
||||
liDepth++;
|
||||
}
|
||||
|
||||
depth++;
|
||||
@@ -431,14 +436,23 @@ class XPathProgressResolver final : public Print {
|
||||
insideBody = false;
|
||||
parentStates.clear();
|
||||
path.clear();
|
||||
textNodeIndexStack.clear();
|
||||
return;
|
||||
}
|
||||
|
||||
if (name == "p" && paragraphDepth > 0) {
|
||||
paragraphDepth--;
|
||||
paragraphVisibleChars = 0;
|
||||
}
|
||||
if (name == "li" && liDepth > 0) {
|
||||
liDepth--;
|
||||
}
|
||||
|
||||
if (!textNodeIndexStack.empty()) {
|
||||
textNodeIndexStack.pop_back();
|
||||
}
|
||||
if (paragraphDepth > 0 || liDepth > 0) {
|
||||
pendingTextNode = true;
|
||||
}
|
||||
if (!path.empty()) {
|
||||
path.pop_back();
|
||||
}
|
||||
@@ -448,23 +462,38 @@ class XPathProgressResolver final : public Print {
|
||||
}
|
||||
|
||||
void onCharacterData(const XML_Char* data, const int len) {
|
||||
if (!insideBody || paragraphDepth <= 0 || len <= 0 || stopped) {
|
||||
if (!insideBody || (paragraphDepth <= 0 && liDepth <= 0) || len <= 0 || stopped) {
|
||||
return;
|
||||
}
|
||||
|
||||
const size_t codepointCount = countUtf8Codepoints(data, len);
|
||||
if (codepointCount == 0) {
|
||||
return;
|
||||
}
|
||||
|
||||
// Start a new text node on first non-empty content after any element boundary.
|
||||
// Only counting non-empty nodes matches KOReader's text()[N] indexing behavior,
|
||||
// which skips empty text nodes created by bare <a id="anchor"/> anchors.
|
||||
if (pendingTextNode) {
|
||||
if (!textNodeIndexStack.empty()) {
|
||||
textNodeIndexStack.back()++;
|
||||
}
|
||||
textNodeStartChars = visibleChars;
|
||||
pendingTextNode = false;
|
||||
}
|
||||
|
||||
const size_t nextVisibleChars = visibleChars + codepointCount;
|
||||
if (targetVisibleChar <= nextVisibleChars) {
|
||||
const size_t delta = targetVisibleChar - visibleChars;
|
||||
const int charOffset = static_cast<int>(paragraphVisibleChars + delta);
|
||||
xpath = buildParagraphXPath(spineIndex, path, std::max(1, charOffset));
|
||||
const int texNode = textNodeIndexStack.empty() ? 0 : textNodeIndexStack.back();
|
||||
const size_t charOff = visibleChars - textNodeStartChars + delta;
|
||||
xpath = buildParagraphXPath(spineIndex, path, texNode, charOff);
|
||||
stopped = true;
|
||||
XML_StopParser(parser, XML_FALSE);
|
||||
return;
|
||||
}
|
||||
|
||||
visibleChars = nextVisibleChars;
|
||||
paragraphVisibleChars += codepointCount;
|
||||
}
|
||||
|
||||
XML_Parser parser = nullptr;
|
||||
@@ -472,11 +501,14 @@ class XPathProgressResolver final : public Print {
|
||||
bool parseOk = true;
|
||||
bool insideBody = false;
|
||||
bool stopped = false;
|
||||
bool pendingTextNode = true;
|
||||
int depth = 0;
|
||||
int bodyDepth = -1;
|
||||
int paragraphDepth = 0;
|
||||
int liDepth = 0;
|
||||
size_t visibleChars = 0;
|
||||
size_t paragraphVisibleChars = 0;
|
||||
size_t textNodeStartChars = 0;
|
||||
std::vector<int> textNodeIndexStack;
|
||||
std::vector<ParentState> parentStates;
|
||||
std::vector<PathSegment> path;
|
||||
std::string xpath;
|
||||
|
||||
Reference in New Issue
Block a user