Provided is a
machine translation
processing system that enables highly accurate
machine translation of source sentences including markup language tags while retaining information about the markup language tags, without the need to prepare a large amount of
parallel translation data including tags. As described above, in the
machine translation
processing system 1000, the training data generation device 1 performs training data generation
processing to detect the start / end correspondence codes and replace the detected start / end correspondence codes with alternative codes in
parallel translation sentences that do not include markup language tags, thereby allowing for easily generating a large amount of data equivalent to
parallel translation data into which markup language tags have been inserted. Using the parallel translation data obtained through the training data generation processing by the training data generation device 1 in the
machine translation processing
system 1000 as training data for learning processing of the
machine translation model achieves the same advantageous effect as when the
machine translation model learning processing is performed using the parallel translation sentences with markup language tags as training data.