Ecosyste.ms: Repos

An open API service providing repository metadata for many open source software ecosystems.

GitHub topics: tokeniser

andreihar/taibun.js

Taiwanese Hokkien Transliterator and Tokeniser

Language: JavaScript - Size: 1.63 MB - Last synced: about 8 hours ago - Pushed: 1 day ago - Stars: 1 - Forks: 0

LanguageMachines/ucto

Unicode tokeniser. Ucto tokenizes text files: it separates words from punctuation, and splits sentences. It offers several other basic preprocessing steps such as changing case that you can all use to make your text suited for further processing such as indexing, part-of-speech tagging, or machine translation. Ucto comes with tokenisation rules for several languages and can be easily extended to suit other languages. It has been incorporated for tokenizing Dutch text in Frog, our Dutch morpho-syntactic processor. http://ilk.uvt.nl/ucto --

Language: C++ - Size: 6.04 MB - Last synced: 1 day ago - Pushed: 6 days ago - Stars: 63 - Forks: 13

andreihar/taibun

Taiwanese Hokkien Transliterator and Tokeniser

Language: Python - Size: 6.43 MB - Last synced: 2 days ago - Pushed: 3 days ago - Stars: 10 - Forks: 0

dragonofmercy/Tokenize2 📦

Tokenize2 is a plugin which allows your users to select multiple items from a predefined list or ajax, using autocompletion as they type to find each item. You may have seen a similar type of text entry when filling in the recipients field sending messages on facebook or tags on tumblr.

Language: JavaScript - Size: 322 KB - Last synced: 15 days ago - Pushed: over 1 year ago - Stars: 82 - Forks: 26

kuhumcst/rtfreader

Text segmenter and tokeniser for Danish, English and other languages. Reads an RTF or flat text file and outputs the text, one line per sentence & optionally tokenized.

Language: C++ - Size: 375 KB - Last synced: about 2 months ago - Pushed: over 1 year ago - Stars: 6 - Forks: 4

ztjhz/word-piece-tokenizer

A Lightweight Word Piece Tokenizer

Language: Python - Size: 121 KB - Last synced: 6 days ago - Pushed: over 1 year ago - Stars: 5 - Forks: 0

phughesmcr/happynodetokenizer

Javascript port of HappyFunTokenizer.py by Christopher Potts and HappierFunTokenizing.py by H. Andrew Schwartz

Language: TypeScript - Size: 1.64 MB - Last synced: 22 days ago - Pushed: 3 months ago - Stars: 5 - Forks: 0

adamscybot/tc-message-toolkit

📃 🧰 🔧 [IN ACTIVE DEV August 2023] A well-typed TypeScript toolkit to work with TeamCity service messages that runs in Node and the browser with a fluent API. Parse message logs (with streaming!) and build up a tree, edit messages, validate messages syntactically and semantically, add your own message types, build test reporters and more!

Language: TypeScript - Size: 5.64 MB - Last synced: 8 months ago - Pushed: 8 months ago - Stars: 1 - Forks: 0

jonsafari/tok-tok

A fast, simple, multilingual tokenizer

Language: Python - Size: 13.7 KB - Last synced: 2 months ago - Pushed: almost 7 years ago - Stars: 28 - Forks: 3

ben-sb/jisu

JavaScript Parser

Language: TypeScript - Size: 456 KB - Last synced: about 1 year ago - Pushed: over 1 year ago - Stars: 8 - Forks: 0