Joint Chinese Word Segmentation and POS Tagging Using an Error-Driven Word-Character Hybrid Model

Canasai KRUENGKRAI
Kiyotaka UCHIMOTO
Jun'ichi KAZAMA
Yiou WANG
Kentaro TORISAWA
Hitoshi ISAHARA

Publication
IEICE TRANSACTIONS on Information and Systems   Vol.E92-D    No.12    pp.2298-2305
Publication Date: 2009/12/01
Online ISSN: 1745-1361
DOI: 10.1587/transinf.E92.D.2298
Print ISSN: 0916-8532
Type of Manuscript: Special Section PAPER (Special Section on Natural Language Processing and its Applications)
Category: Morphological/Syntactic Analysis
Keyword: 
word segmentation,  POS tagging,  error-driven,  word-character hybrid model,  

Full Text: PDF(337.7KB)>>
Buy this Article



Summary: 
In this paper, we present a discriminative word-character hybrid model for joint Chinese word segmentation and POS tagging. Our word-character hybrid model offers high performance since it can handle both known and unknown words. We describe our strategies that yield good balance for learning the characteristics of known and unknown words and propose an error-driven policy that delivers such balance by acquiring examples of unknown words from particular errors in a training corpus. We describe an efficient framework for training our model based on the Margin Infused Relaxed Algorithm (MIRA), evaluate our approach on the Penn Chinese Treebank, and show that it achieves superior performance compared to the state-of-the-art approaches reported in the literature.