Statistical text language recognition with the use of n-gram frequency
Abstract:
Statistical properties of European language texts are investigated with the use of recognition procedure for n-gram distribution patterns. The numerical algorithm is constructed for analysis Hurst exponent for letter distance distributions of the text fragment. The accuracy of binary recognition is estimated as 0,99.
Keywords:
text language recognition, n-gram frequency
Publication language:russian, pages:21
Research direction:
Mathematical modelling in actual problems of science and technics