dc.contributor.author | Cruz Díaz, Noa Patricia | |
dc.contributor.author | Maña López, Manuel Jesús | |
dc.date.accessioned | 2016-04-20T07:49:45Z | |
dc.date.available | 2016-04-20T07:49:45Z | |
dc.date.issued | 2015 | |
dc.identifier.citation | Cruz Díaz, N.P., Maña López, M.J.: "An analysis of biomedical tokenization : problems and strategies". En: Sixth International Workshop on Health Text Mining and Information Analysis (Louhi), pages 40–49,. Lisbon, Portugal, 17 September 2015 | en_US |
dc.identifier.uri | http://hdl.handle.net/10272/11951 | |
dc.description.abstract | Choosing the right tokenizer is a non-trivial task, especially in the biomedical domain, where it poses additional challenges, which if not resolved means the propagation of errors in successive Natural Language Processing analysis pipeline. This paper aims to identify these problematic cases and analyze the out-put that, a representative and widely used set of tokenizers, shows on them. This work will aid the decision making process of choosing the right strategy according to the down-stream application. In addition, it will help developers to create accurate tokenization tools or improve the existing ones. A total of 14 problematic cases were described, show-ing biomedical samples for each of them. The outputs of 12 tokenizers were provided and discussed in relation to the level of agreement among tools | |
dc.language.iso | eng | en_US |
dc.rights | Atribución-NoComercial-SinDerivadas 3.0 España | * |
dc.rights.uri | http://creativecommons.org/licenses/by-nc-nd/3.0/es/ | * |
dc.title | An analysis of biomedical tokenization : problems and strategies | en_US |
dc.type | info:eu-repo/semantics/conferenceObject | en_US |
dc.rights.accessRights | info:eu-repo/semantics/openAccess | en_US |