Zaman, Farooq, Shardlow, Matthew ORCID: https://orcid.org/0000-0003-1129-2750, Hassan, Saeed-Ul, Aljohani, Naif Radi and Nawaz, Raheel ORCID: https://orcid.org/0000-0001-9588-0052 (2020) HTSS: A novel hybrid text summarisation and simplification architecture. Information Processing & Management, 57 (6). 102351. ISSN 0306-4573
|
Accepted Version
Available under License In Copyright. Download (2MB) | Preview |
Abstract
Text simplification and text summarisation are related, but different sub-tasks in Natural Language Generation. Whereas summarisation attempts to reduce the length of a document, whilst keeping the original meaning, simplification attempts to reduce the complexity of a document. In this work, we combine both tasks of summarisation and simplification using a novel hybrid architecture of abstractive and extractive summarisation called HTSS. We extend the well-known pointer generator model for the combined task of summarisation and simplification. We have collected our parallel corpus from the simplified summaries written by domain experts published on the science news website EurekaAlert (www.eurekalert.org). Our results show that our proposed HTSS model outperforms neural text simplification (NTS) on SARI score and abstractive text summarisation (ATS) on the ROUGE score. We further introduce a new metric (CSS1) which combines SARI and Rouge and demonstrates that our proposed HTSS model outperforms NTS and ATS on the joint task of simplification and summarisation by 38.94% and 53.40%, respectively. We provide all code, models and corpora to the scientific community for future research at the following URL: https://github.com/slab-itu/HTSS/.
Impact and Reach
Statistics
Additional statistics for this dataset are available via IRStats2.