fix(archive): sanitize non-UTF-8 entry names during extraction (#3225)
ZIP archives with GBK-encoded names inserted raw bytes into the database, failing the whole task with 'pq: invalid byte sequence for encoding UTF8'. Entry names now pass through ToValidUTF8 at both the listing and extraction layers, so masked extraction stays consistent and undecodable names degrade to replacement chars instead of a crash. The existing encoding picker (gbk/gb18030/big5/shiftjis/...) remains the way to recover proper names. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>pull/3582/head
parent
01e65f4e69
commit
4b5433179e
Loading…
Reference in new issue