Rebuilds the PDF's text as a .docx — reading the real text layer where there is one, and falling back to OCR on scanned pages. Note: plain paragraphs only; layout is not preserved.