UM  > Faculty of Science and Technology
Residential Collegefalse
Status已發表Published
Word Segmentation by Separation Inference for East Asian Languages
Tong, Yu; Guo, Jingzhi; Zhou, Jizhe; Chen, Ge; Zhen, Guokai
2022-05
Conference NameACL 2022
VolumeFindings of the Association for Computational Linguistics: ACL 2022
Pages3924–3934
Conference DateMay 2021
Conference PlaceACL | Findings, Dublin
CountryIreland
PublisherACL Anthology
Abstract

Chinese Word Segmentation (CWS) intends to divide a raw sentence into words through sequence labeling. Thinking in reverse, CWS can also be viewed as a process of grouping a sequence of characters into a sequence of words. In such a way, CWS is reformed as a separation inference task in every adjacent character pair. Since every character is either connected or not connected to the others, the tagging schema is simplified as two tags “Connection” (C) or “NoConnection” (NC). Therefore, bigram is specially tailored for “C-NC” to model the separation state of every two consecutive characters. Our Separation Inference (SpIn) framework is evaluated on five public datasets, is demonstrated to work for machine learning and deep learning models, and outperforms state-of-the-art performance for CWS in all experiments. Performance boosts on Japanese Word Segmentation (JWS) and Korean Word Segmentation (KWS) further prove the framework is universal and effective for East Asian Languages.

DOI10.18653/v1/2022.findings-acl.309
URLView the original
Indexed ByEI
WOS IDWOS:000828767404002
The Source to Articlehttps://aclanthology.org/2022.findings-acl.309.pdf
Scopus ID2-s2.0-85149131122
Fulltext Access
Citation statistics
Document TypeConference paper
CollectionFaculty of Science and Technology
DEPARTMENT OF COMPUTER AND INFORMATION SCIENCE
AffiliationUniversity of Macau
First Author AffilicationUniversity of Macau
Recommended Citation
GB/T 7714
Tong, Yu,Guo, Jingzhi,Zhou, Jizhe,et al. Word Segmentation by Separation Inference for East Asian Languages[C]:ACL Anthology, 2022, 3924–3934.
APA Tong, Yu., Guo, Jingzhi., Zhou, Jizhe., Chen, Ge., & Zhen, Guokai (2022). Word Segmentation by Separation Inference for East Asian Languages. , Findings of the Association for Computational Linguistics: ACL 2022, 3924–3934.
Files in This Item:
There are no files associated with this item.
Related Services
Recommend this item
Bookmark
Usage statistics
Export to Endnote
Google Scholar
Similar articles in Google Scholar
[Tong, Yu]'s Articles
[Guo, Jingzhi]'s Articles
[Zhou, Jizhe]'s Articles
Baidu academic
Similar articles in Baidu academic
[Tong, Yu]'s Articles
[Guo, Jingzhi]'s Articles
[Zhou, Jizhe]'s Articles
Bing Scholar
Similar articles in Bing Scholar
[Tong, Yu]'s Articles
[Guo, Jingzhi]'s Articles
[Zhou, Jizhe]'s Articles
Terms of Use
No data!
Social Bookmark/Share
All comments (0)
No comment.
 

Items in the repository are protected by copyright, with all rights reserved, unless otherwise indicated.