PILOT EDITION 
THE PALACE OF SCIENCE BELGRADE
25-26 September 2025

Instructions for reviewers

Each submitted paper will be reviewed by at least two members of the Scientific committee following a double-blind policy. Reviewers will assess the quality of the papers focusing specifically on the points listed below.

Contribution

Paper type

Performance improvement in NLP and Language services

  • Is the NLP task / Language service clearly defined?
  • What is the baseline? Is it appropriate?
  • Is the improvement consistent across multiple tests?

Evaluation in NLP and Language services

  • Is the NLP task / Language service clearly defined?
  • Is the selection of evaluated models appropriate?
  • Are the evaluation results coherent across multiple tests?

Language structure

  • Is the addressed structural element clearly defined in terms of general linguistics?
  • Is the problem caused by this element realistic?
  • Is the proposed solution feasible?

Language theory and Machine learning

  • Is the theoretical problem relevant to language processing?
  • Is the proposed solution compared to competing ideas?
  • Is the proposed solution feasible?

Survey

  • Is the need for the survey well justified?
  • Is the survey structured as a coherent single argument?
  • Is the relevant domain exhaustively covered?

Methodology

Empirical contribution

  • Can the stated conclusions be drawn from the outcomes of the experiments?
  • Are the data sets used in the experiments publicly available?
  • Are the methods reproducible?
  • Are the data / methods selection criteria appropriate?

Theoretical contribution

  • Is the argument clearly spelled out?
  • Is the inference correct?
  • Are all relevant premises taken into account?

Each paper is scored according to its contribution type by assigning a number to each of the bullet points:

1

There is a problem

2

All is fine

3

Exceptionally well done

The final score for the paper is the sum of all individual scores. The minimal score is 7 for an empirical paper (3 on the contribution plus 4 on the methodology) and 6 for a theoretical paper (3 + 3). The maximal scores are 21 for an empirical paper and 18 for a theoretical paper. The scores will be scaled for comparability. Reviewers can add an optional comment to each score. If there is a strong disagreement between the reviewers, the Organising committee will rely on these comments to decide which review to take as more reliable.