Distributional Discrepancy: A Metric for Unconditional Text Generation

Cai, Ping; Chen, Xingyuan; Jin, Peng; Wang, Hongjun; Li, Tianrui

Computer Science > Computation and Language

arXiv:2005.01282 (cs)

[Submitted on 4 May 2020 (v1), last revised 2 Jul 2020 (this version, v2)]

Title:Distributional Discrepancy: A Metric for Unconditional Text Generation

Authors:Ping Cai, Xingyuan Chen, Peng Jin, Hongjun Wang, Tianrui Li

View PDF

Abstract:The purpose of unconditional text generation is to train a model with real sentences, then generate novel sentences of the same quality and diversity as the training data. However, when different metrics are used for comparing the methods of unconditional text generation, contradictory conclusions are drawn. The difficulty is that both the diversity and quality of the sample should be considered simultaneously when the models are evaluated. To solve this problem, a novel metric of distributional discrepancy (DD) is designed to evaluate generators based on the discrepancy between the generated and real training sentences. However, it cannot compute the DD directly because the distribution of real sentences is unavailable. Thus, we propose a method for estimating the DD by training a neural-network-based text classifier. For comparison, three existing metrics, bi-lingual evaluation understudy (BLEU) versus self-BLEU, language model score versus reverse language model score, and Fr�chet embedding distance, along with the proposed DD, are used to evaluate two popular generative models of long short-term memory and generative pretrained transformer 2 on both syntactic and real data. Experimental results show that DD is significantly better than the three existing metrics for ranking these generative models.

Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2005.01282 [cs.CL]
	(or arXiv:2005.01282v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2005.01282

Submission history

From: Peng Jin [view email]
[v1] Mon, 4 May 2020 05:53:34 UTC (1,957 KB)
[v2] Thu, 2 Jul 2020 15:40:14 UTC (1,958 KB)

Computer Science > Computation and Language

Title:Distributional Discrepancy: A Metric for Unconditional Text Generation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Distributional Discrepancy: A Metric for Unconditional Text Generation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators