Training and intrinsic evaluation of lightweight word embeddings for the clinical domain in Spanish

Chiu, Carolina; Villena, Fabián; Martin, Kinan; Núñez, Fredy R.; Besa Correa, Cecilia; Dunstan, Jocelyn

Training and intrinsic evaluation of lightweight word embeddings for the clinical domain in Spanish

dc.contributor.author	Chiu, Carolina
dc.contributor.author	Villena, Fabián
dc.contributor.author	Martin, Kinan
dc.contributor.author	Núñez, Fredy R.
dc.contributor.author	Besa Correa, Cecilia
dc.contributor.author	Dunstan, Jocelyn
dc.date.accessioned	2022-12-28T15:35:05Z
dc.date.available	2022-12-28T15:35:05Z
dc.date.issued	2022
dc.description.abstract	Resources for Natural Language Processing (NLP) are less numerous for languages different from English. In the clinical domain, where these resources are vital for obtaining new knowledge about human health and diseases, creating new resources for the Spanish language is imperative. One of the most common approaches in NLP is word embeddings, which are dense vector representations of a word, considering the word's context. This vector representation is usually the first step in various NLP tasks, such as text classification or information extraction. Therefore, in order to enrich Spanish language NLP tools, we built a Spanish clinical corpus from waiting list diagnostic suspicions, a biomedical corpus from medical journals, and term sequences sampled from the Unified Medical Language System (UMLS). These three corpora can be used to compute word embeddings models from scratch using Word2vec and fastText algorithms. Furthermore, to validate the quality of the calculated embeddings, we adapted several evaluation datasets in English, including some tests that have not been used in Spanish to the best of our knowledge. These translations were validated by two bilingual clinicians following an ad hoc validation standard for the translation. Even though contextualized word embeddings nowadays receive enormous attention, their calculation and deployment require specialized hardware and giant training corpora. Our static embeddings can be used in clinical applications with limited computational resources. The validation of the intrinsic test we present here can help groups working on static and contextualized word embeddings. We are releasing the training corpus and the embeddings within this publication.
dc.fechaingreso.objetodigital	2022-12-28
dc.fuente.origen	SIPA
dc.identifier.doi	10.3389/frai.2022.970517
dc.identifier.uri	https://www.frontiersin.org/articles/10.3389/frai.2022.970517/full
dc.identifier.uri	https://repositorio.uc.cl/handle/11534/66147
dc.identifier.wosid	WOS:000950590400001
dc.information.autoruc	Facultad de letras; Núñez, Fredy R.; 0000-0002-0643-6628; 157277
dc.information.autoruc	Escuela de medicina; Besa Correa, Cecilia; 0000-0002-0015-0434; 167343
dc.language.iso	en
dc.nota.acceso	Contenido completo
dc.revista	Frontiers in Artificial Intelligence
dc.rights	acceso restringido
dc.subject	Natural language processing
dc.subject	Spanish language
dc.subject	Word embeddings
dc.subject	Medical informatics
dc.subject	Neural networks
dc.subject	Intrinsic evaluation
dc.subject	Semantic evaluation
dc.subject.ods	03 Good health and well-being
dc.subject.odspa	03 Salud y bienestar
dc.title	Training and intrinsic evaluation of lightweight word embeddings for the clinical domain in Spanish
dc.type	artículo
dc.volumen	5
sipa.codpersvinculados	157277
sipa.codpersvinculados	167343
sipa.index	Scopus

Files

Original bundle

Now showing 1 - 1 of 1

Name:: frai-05-970517.pdf
Size:: 333.38 KB
Format:: Adobe Portable Document Format
Description:

Download

Collections

Artículos de revistas