Predicting syntactic equivalence between source and target sentences

Authors

  • Bram Vanroy
  • Arda Tezcan
  • Lieve Macken

Abstract

The translation difficulty of a text is influenced by many different factors. Some of these are specific to the source text and related to readability while others more directly involve translation and the relation between the source and the target text. One such factor is syntactic equivalence, which can be calculated on the basis of a source sentence and its translation. When the expected syntactic form of the target sentence is dissimilar to its source, translating said source sentence proves more difficult for a translator. The degree of syntactic equivalence between a word-aligned source and target sentence can be derived from the crossing alignment links, averaged by the number of alignments, either at word or at sequence level. However, when predicting the translatability of a source sentence, its translation is not available. Therefore, we train machine learning systems on a parallel English-Dutch corpus to predict the expected syntactic equivalence of an English source sentence without having access to its Dutch translation. We use traditional machine learning systems (Random Forest Regression and Support Vector Regression) combined with syntactic sentence-level features as well as recurrent neural networks that utilise word embeddings and accurate morpho-syntactic features.

Author Biographies

  • Bram Vanroy

    University of Ghent

  • Arda Tezcan

    University of Ghent

  • Lieve Macken

    University of Ghent

Downloads

Published

2019-12-16

Issue

Section

Articles

How to Cite

Predicting syntactic equivalence between source and target sentences. (2019). Computational Linguistics in the Netherlands Journal, 9, 101-116. https://clinjournal.org/clinj/article/view/95

Most read articles by the same author(s)