Moses-support Digest, Vol 87, Issue 29

Send Moses-support mailing list submissions to
moses-support@mit.edu

To subscribe or unsubscribe via the World Wide Web, visit
http://mailman.mit.edu/mailman/listinfo/moses-support
or, via email, send a message with subject or body 'help' to
moses-support-request@mit.edu

You can reach the person managing the list at
moses-support-owner@mit.edu

When replying, please edit your Subject line so it is more specific
than "Re: Contents of Moses-support digest..."

Today's Topics:

1. Re: When using moses, cpu is are not stable. (Tom Hoar)
2. 2nd CfP: LREC 2014 Workshop on Building and Using Comparable
Corpora (7th BUCC) (Reinhard Rapp)

----------------------------------------------------------------------

Message: 1
Date: Wed, 15 Jan 2014 00:29:14 +0700
From: Tom Hoar <tahoar@precisiontranslationtools.com>
Subject: Re: [Moses-support] When using moses, cpu is are not stable.
To: moses-support@mit.edu
Message-ID: <52D573EA.4010400@precisiontranslationtools.com>
Content-Type: text/plain; charset="iso-8859-1"

To change the threads, run with '-threads x' added to your command line:

moses -f /path/to/moses.ini -threads 20

On 01/14/2014 11:41 PM, Philipp Koehn wrote:
> Hi,
>
> you can try specifying more threads than the number of CPUs.
> I have no idea if this speeds things up or causes blockage in
> your setup. Just try it out and see what happens.
>
> -phi
>
>
> On Tue, Jan 14, 2014 at 4:31 AM, Hansung Cho <gyuber@gmail.com
> <mailto:gyuber@gmail.com>> wrote:
>
> My server configuration are multiple CPUs 16 core and running
> thread 16.
>
> How do I adjust the moses running thread?
>
> I need your help...
>
> Thank you
>
>
> 2014/1/14 Tom Hoar <tahoar@precisiontranslationtools.com
> <mailto:tahoar@precisiontranslationtools.com>>
>
> With a system as complex as Moses, why would you think the CPU
> usage would be steady? Each sentence task goes through many
> stages with different CPU demands which probably map to the
> variation in each cycle.
>
> You can also look first at your threading configuration. If
> your system has multiple CPUs (cores) and Moses is running
> multi-threaded, I suspect the cycles' frequencies are related
> to the activities on each thread.
>
>
>
> On 01/14/2014 01:09 PM, Hansung Cho wrote:
>> hello..
>>
>> I'm using mosesdecoding
>>
>> When monitoring a physical server, Server cpu is not very stable.
>> However, the state is constantly moving cpu
>>
>> Look at the picture below
>> ?? ??? 1
>>
>> Is it because the moses system cache??
>> Should any part option of the adjustment?
>>
>> My moses commit a1584c608f8bbd83921896bac9eac97d9effc0c8
>> Options related to the cache are "clean-lm-cache" ( default ==1 )
>>
>> Will be waiting for your input.
>>
>> Thank you
>> Hansung cho
>>
>>
>> _______________________________________________
>> Moses-support mailing list
>> Moses-support@mit.edu <mailto:Moses-support@mit.edu>
>> http://mailman.mit.edu/mailman/listinfo/moses-support
>
>
> _______________________________________________
> Moses-support mailing list
> Moses-support@mit.edu <mailto:Moses-support@mit.edu>
> http://mailman.mit.edu/mailman/listinfo/moses-support
>
>
>
> _______________________________________________
> Moses-support mailing list
> Moses-support@mit.edu <mailto:Moses-support@mit.edu>
> http://mailman.mit.edu/mailman/listinfo/moses-support
>
>
>
>
> _______________________________________________
> Moses-support mailing list
> Moses-support@mit.edu
> http://mailman.mit.edu/mailman/listinfo/moses-support

-------------- next part --------------
An HTML attachment was scrubbed...
URL: http://mailman.mit.edu/mailman/private/moses-support/attachments/20140115/287a8777/attachment-0001.htm
-------------- next part --------------
A non-text attachment was scrubbed...
Name: not available
Type: image/png
Size: 4887 bytes
Desc: not available
Url : http://mailman.mit.edu/mailman/private/moses-support/attachments/20140115/287a8777/attachment-0001.png

------------------------------

Message: 2
Date: Tue, 14 Jan 2014 20:40:53 +0100
From: "Reinhard Rapp" <reinhardrapp@gmx.de>
Subject: [Moses-support] 2nd CfP: LREC 2014 Workshop on Building and
Using Comparable Corpora (7th BUCC)
To: <IRList@lists.shef.ac.uk>, <listmaster@loria.fr>, <ln@cines.fr>,
<lr_egroup@mail.iiit.ac.in>, <moses-support@mit.edu>,
<news@multilingual.com>
Message-ID: <044E9B37C13044689CEBA1E4BDDE5CED@ASUSPC>
Content-Type: text/plain; charset="windows-1252"

We apologize for multiple postings
Please distribute to interested colleagues

============================================================

2nd Call for Papers

7th WORKSHOP ON BUILDING AND USING COMPARABLE CORPORA

Building Resources for Machine Translation Research

http://comparable.limsi.fr/bucc2014/

May 27, 2014
Co-located with LREC 2014
Harpa Conference Centre, Reykjavik (Iceland)

DEADLINE FOR PAPERS: February 10, 2014
https://www.softconf.com/lrec2014/BUCC2014/

*** INVITED SPEAKER ***

Chris Callison-Burch (University of Pennsylvania)

============================================================

MOTIVATION

In the language engineering and the linguistics communities, research
in comparable corpora has been motivated by two main reasons. In
language engineering, on the one hand, it is chiefly motivated by the
need to use comparable corpora as training data for statistical
Natural Language Processing applications such as statistical machine
translation or cross-lingual retrieval. In linguistics, on the other
hand, comparable corpora are of interest in themselves by making
possible inter-linguistic discoveries and comparisons. It is generally
accepted in both communities that comparable corpora are documents in
one or several languages that are comparable in content and form in
various degrees and dimensions. We believe that the linguistic
definitions and observations related to comparable corpora can improve
methods to mine such corpora for applications of statistical NLP. As
such, it is of great interest to bring together builders and users of
such corpora.

The scarcity of parallel corpora has motivated research concerning
the use of comparable corpora: pairs of monolingual corpora selected
according to the same set of criteria, but in different languages
or language varieties. Non-parallel yet comparable corpora overcome
the two limitations of parallel corpora, since sources for original,
monolingual texts are much more abundant than translated texts.
However, because of their nature, mining translations in comparable
corpora is much more challenging than in parallel corpora. What
constitutes a good comparable corpus, for a given task or per se,
also requires specific attention: while the definition of a parallel
corpus is fairly straightforward, building a non-parallel corpus
requires control over the selection of source texts in both languages.

Parallel corpora are a key resource as training data for statistical
machine translation, and for building or extending bilingual lexicons
and terminologies. However, beyond a few language pairs such as
English- French or English-Chinese and a few contexts such as
parliamentary debates or legal texts, they remain a scarce resource,
despite the creation of automated methods to collect parallel corpora
from the Web. To exemplify such issues in a practical setting, this
year's special focus will be on

Building Resources for Machine Translation Research

This special topic aims to address the need for:
(1) Machine Translation training and testing data such as spoken or
written monolingual, comparable or parallel data collections, and
(2) methods and tools used for collecting, annotating, and verifying
MT data such as Web crawling, crowdsourcing, tools for language
experts and for finding MT data in comparable corpora.

TOPICS

We solicit contributions including but not limited to the following topics:

Topics related to the special theme:
* Methods and tools for collecting and processing MT data,
including crowdsourcing
* Methods and tools for quality control
* Tools for efficient annotation
* Bilingual term and named entity collections
* Multilingual treebanks, wordnets, propbanks, etc.
* Comparable corpora with parallel units annotated
* Comparable corpora for under-resourced languages and specific domains
* Multilingual corpora with rich annotations:
POS tags, NEs, dependencies, semantic roles, etc.
* Data for special applications: patent translation, movie
subtitles, MOOCs, meetings, chat-rooms, social media, etc.
* Legal issues with collecting and redistributing data
and generating derivatives

Building comparable corpora:
* Human translations
* Automatic and semi-automatic methods
* Methods to mine parallel and non-parallel corpora from the Web
* Tools and criteria to evaluate the comparability of corpora
* Parallel vs non-parallel corpora, monolingual corpora
* Rare and minority languages, across language families
* Multi-media/multi-modal comparable corpora

Applications of comparable corpora:
* Human translations
* Language learning
* Cross-language information retrieval & document categorization
* Bilingual projections
* Machine translation
* Writing assistance

Mining from comparable corpora:
* Extraction of parallel segments or paraphrases from comparable corpora
* Extraction of bilingual and multilingual translations of single words
and multi-word expressions; proper names, named entities, etc.

IMPORTANT DATES

February 10, 2014 Deadline for submission of full papers
March 10, 2014 Notification of acceptance
March 27, 2014 Camera-ready papers due
May 27, 2014 Workshop date

SUBMISSION INFORMATION

Papers should follow the LREC main conference formatting details (to be
announced on the conference website http://lrec2014.lrec-conf.org/en/ )
and should be submitted as a PDF-file via the START workshop manager at
https://www.softconf.com/lrec2014/BUCC2014/

Contributions can be short or long papers. Short paper submission must
describe original and unpublished work without exceeding six (6)
pages. Characteristics of short papers include: a small, focused
contribution; work in progress; a negative result; an opinion piece;
an interesting application nugget. Long paper submissions must
describe substantial, original, completed and unpublished work without
exceeding ten (10) pages.

Reviewing will be double blind, so the papers should not reveal the
authors' identity. Accepted papers will be published in the workshop
proceedings.

Double submission policy: Parallel submission to other meetings or
publications is possible but must be immediately notified to the
workshop organizers.

When submitting a paper from the START page, authors will be asked to
provide essential information about resources (in a broad sense,
i.e. also technologies, standards, evaluation kits, etc.) that have
been used for the work described in the paper or are a new result of
your research. Moreover, ELRA encourages all LREC authors to share
the described LRs (data, tools, services, etc.), to enable their
reuse, replicability of experiments, including evaluation ones, etc.

For further information, please contact
Pierre Zweigenbaum pz (at) limsi (dot) fr

ORGANISERS

Pierre Zweigenbaum, LIMSI, CNRS, Orsay (France)
Ahmet Aker, University of Sheffield (UK)
Serge Sharoff, University of Leeds (UK)
Stephan Vogel, QCRI (Qatar)
Reinhard Rapp, Universities of Mainz (Germany) and Aix-Marseille (France)

SCIENTIFIC COMMITTEE

* Ahmet Aker, University of Sheffield (UK)
* Srinivas Bangalore (AT&T Labs, US)
* Caroline Barri?re (CRIM, Montr?al, Canada)
* Chris Biemann (TU Darmstadt, Germany)
* Herv? D?jean (Xerox Research Centre Europe, Grenoble, France)
* Kurt Eberle (Lingenio, Heidelberg, Germany)
* Andreas Eisele (European Commission, Luxembourg)
* ?ric Gaussier (Universit? Joseph Fourier, Grenoble, France)
* Gregory Grefenstette (INRIA, Saclay, France)
* Silvia Hansen-Schirra (University of Mainz, Germany)
* Hitoshi Isahara (Toyohashi University of Technology)
* Kyo Kageura (University of Tokyo, Japan)
* Adam Kilgarriff (Lexical Computing Ltd, UK)
* Natalie K?bler (Universit? Paris Diderot, France)
* Philippe Langlais (Universit? de Montr?al, Canada)
* Michael Mohler (Language Computer Corp., US)
* Emmanuel Morin (Universit? de Nantes, France)
* Dragos Stefan Munteanu (Language Weaver, Inc., US)
* Lene Offersgaard (University of Copenhagen, Denmark)
* Ted Pedersen (University of Minnesota, Duluth, US)
* Reinhard Rapp (Universit? Aix-Marseille, France)
* Sujith Ravi (Google, Mountain View, US)
* Serge Sharoff (University of Leeds, UK)
* Michel Simard (National Research Council Canada)
* Richard Sproat (OGI School of Science & Technology, US)
* Tim Van de Cruys (IRIT-CNRS, Toulouse, France)
* Stephan Vogel (QCRI, Qatar)
* Guillaume Wisniewski (Universit? Paris Sud & LIMSI-CNRS, Orsay, France)
* Pierre Zweigenbaum (LIMSI-CNRS, Orsay, France)

-------------- next part --------------
An HTML attachment was scrubbed...
URL: http://mailman.mit.edu/mailman/private/moses-support/attachments/20140114/d1436d65/attachment.htm

------------------------------

_______________________________________________
Moses-support mailing list
Moses-support@mit.edu
http://mailman.mit.edu/mailman/listinfo/moses-support

End of Moses-support Digest, Vol 87, Issue 29
*********************************************

Moses-support Digest, Vol 87, Issue 29

0 Response to "Moses-support Digest, Vol 87, Issue 29"

Post a Comment