ESTD Year: 2017 | Impact Factor: 6.9
DOI Prefix: 10.47001/IRJIET
Vol 10 No 7 (2026): Volume 10, Issue 7, July 2026 | Pages: 92-100
International Research Journal of Innovations in Engineering and Technology
OPEN ACCESS | Research Article | Published Date: 30-07-2026
This paper analyzes and compares the effectiveness of two representative lossless compression algorithms, Huffman coding and LZW, on Vietnamese text data encoded in UTF-8. A web-based text compression application was developed in JavaScript, implementing byte-level canonical Huffman coding and LZW with variable code width from 9 to 16 bits. The two algorithms were evaluated on a corpus of approximately 360 KB covering four text categories (encyclopedic, literary, legal–administrative, and scientific–technological) collected from Vietnamese Wikipedia and Wikisource, through three experiments: by text category, by data size, and by the influence of Vietnamese diacritics. The results show that LZW achieves compression ratios of 29.3–48.2% of the original size and becomes increasingly effective as data grows, clearly outperforming Huffman coding (64.3–68.6%), which is lower-bounded by the symbol entropy of the data (about 5.2 bits/byte). In contrast, Huffman compresses roughly four times faster (about 80 MB/s versus 20 MB/s). The experiments also indicate that the multi-byte UTF-8 encoding of Vietnamese script significantly reduces the effectiveness of byte-level Huffman coding while having almost no effect on LZW. Overall, these findings indicate that dictionary-based coding is better suited to Vietnamese UTF-8 text, whereas byte-level statistical coding is best reserved for speed-critical settings or as an entropy-coding stage after another transform; the compression tool and the four-category corpus built here are offered as reusable resources.
Huffman, LZW, lossless data compression, Vietnamese text, UTF-8, entropy.
Nguyen Minh Hoang, & Trinh Van Hung. (2026). Analysis of the Efficiency of Huffman and LZW Compression Algorithms on Vietnamese Text Data. International Research Journal of Innovations in Engineering and Technology - IRJIET, 10(7), 92-100. Article DOI https://doi.org/10.47001/IRJIET/2026.107010
This work is licensed under Creative common Attribution Non Commercial 4.0 Internation Licence
T. A. Welch, “A Technique for High-Performance Data Compression,” Computer, vol. 17, no. 6, pp. 8–19, 1984, doi: 10.1109/MC.1984.1659158.
J. Ziv and A. Lempel, “Compression of Individual Sequences via Variable-Rate Coding,” IEEE Transactions on Information Theory, vol. 24, no. 5, pp. 530–536, 1978.
C. E. Shannon, “A Mathematical Theory of Communication,” Bell System Technical Journal, vol. 27, no. 3, pp. 379–423, 1948.
The Unicode Consortium. The Unicode Standard, Version 15.0. Mountain View, CA: The Unicode Consortium, 2022.
D. Salomon, Data Compression: The Complete Reference, 4th ed., London: Springer, 2007.
T. M. Cover and J. A. Thomas, Elements of Information Theory, 2nd ed., Hoboken, NJ: Wiley-Interscience, 2006.
R. Arnold and T. Bell, "A Corpus for the Evaluation of Lossless Compression Algorithms," Proceedings DCC, 1997.
F. Yergeau, “UTF-8, a transformation format of ISO 10646,” IETF RFC 3629, Nov. 2003.
V. H. Nguyen, H. T. Nguyen, H. N. Duong, and V. Snasel, “A syllable-based method for Vietnamese text compression,” in Proc. 10th Int. Conf. Ubiquitous Information Management and Communication (IMCOM), Danang, Vietnam, 2016.
V. H. Nguyen, H. T. Nguyen, H. N. Duong, and V. Snasel, “Trigram-Based Vietnamese Text Compression,” in Recent Developments in Intelligent Information and Database Systems, Cham: Springer, 2016, pp. 297–307.