bleepit
Keep profanity out of your product. A fast, free profanity checker that plugs into anywhere JavaScript runs — Node, browsers, Deno, Bun, and workers — with zero dependencies and wordlists you control.
- ~2.7 kB gzip
- 0 dependencies
- one O(n) scan
- ~17 MB/s
- en · es · fr · de starter lists
- leet + elongation + diacritic tolerant
| Package | What it does |
|---|---|
@bleepit/core |
Profanity checking for text. The engine everything else builds on. |
@bleepit/ocr |
Profanity checking for images: OCR the page, map matches back to bounding boxes. |
npm i @bleepit/core
pnpm add @bleepit/core
bun add @bleepit/core
deno add npm:@bleepit/core
Live demo
Runs 100% in your browser using the real library bundle — nothing is sent anywhere.
| word | start | end |
|---|
Installation
npm i @bleepit/core
pnpm add @bleepit/core
bun add @bleepit/core
deno add npm:@bleepit/core
Ships tree-shakeable ESM + CJS with types. Import only what you need:
import { ProfanityChecker } from "@bleepit/core";
import { es } from "@bleepit/core/lists"; // individual lists for tiny bundles
Quickstart
import { ProfanityChecker } from "@bleepit/core";
const checker = new ProfanityChecker({ languages: ["en", "es"] });
checker.isProfane("What an a.r.s.e?!"); // true
checker.isProfane("arrrse"); // true (elongation)
checker.isProfane("@rse"); // true (leet-speak)
checker.isProfane("the class"); // false (word boundaries)
checker.censor("That was a d0uche move"); // "That was a ****** move"
API reference
ProfanityChecker
| Method | Signature | Description |
|---|---|---|
isProfane | (text) => boolean | Fast boolean check, early-exits on first valid match. |
find | (text, limit?) => Match[] | All matches as { word, start, end } with original-string offsets (UTF-16). |
censor | (text, mask="*") => string | Masks profane spans; whitespace inside a span is preserved. |
addWords | (words) => void | Adds words (any language/script) and rebuilds the automaton. |
removeWords | (words) => void | Removes words and rebuilds the automaton. |
size | number | Count of normalized patterns loaded. |
Factories & singleton
import { createChecker, isProfane, find, censor } from "@bleepit/core";
const mine = createChecker({ languages: ["fr"] }); // isolated instance
isProfane("bollocks"); // uses the shared default (English) instance
Options
| Option | Type | Default | Description |
|---|---|---|---|
languages | string[] | ["en"] | Built-in lists to load (en, es, fr, de). |
customWords | string[] | [] | Extra words in any language or script. |
whitelist | string[] | [] | Words that must never be flagged (matched against the enclosing word). |
wholeWord | boolean | true | Match on word boundaries only. Set false for aggressive filtering or spaceless scripts (e.g. CJK). |
leet | boolean | true | Map leet-speak (@→a, 0→o, $→s…). |
stripDiacritics | boolean | true | Fold diacritics (é→e) and ligatures (ß→ss). Disable for max speed on plain ASCII. |
Languages
Four starter lists ship with the package. Anything else is customWords — the engine itself is script-agnostic, so the words you add can be in any language or writing system.
| Code | Language | Words |
|---|---|---|
en | English (default) | 66 |
es | Spanish | 24 |
fr | French | 17 |
de | German | 16 |
import { createChecker } from "@bleepit/core";
// Not built in — bring your own words, in any script:
const c = createChecker({ languages: [], customWords: ["чёрт", "クソ"] });
c.isProfane("вот ЧЁРТ"); // true (case-insensitive Cyrillic)
// Spaceless scripts have no word boundaries — use substring mode:
const zh = createChecker({ languages: [], customWords: ["笨蛋"], wholeWord: false });
Guides
Word boundaries (the Scunthorpe problem)
With wholeWord: true (default), matches must sit on word boundaries checked against the original text, so separators still split words while a.r.s.e is caught:
checker.isProfane("the class"); // false
checker.isProfane("Scunthorpe"); // false
checker.isProfane("arse about"); // true
checker.isProfane("a.r.s.e about"); // true
Whitelisting
const c = createChecker({ whitelist: ["arsenal"] });
c.isProfane("arsenal"); // false
Runtime updates
checker.addWords(["heck"]); // community-driven lists
checker.removeWords(["bollocks"]); // tune false positives away
Censoring
checker.censor("a d0uche move", "*"); // "a ****** move"
checker.censor("oh bollocks", "#"); // "oh ########"
Integrations
React (comment box)
import { useMemo, useState } from "react";
import { createChecker } from "@bleepit/core";
const checker = createChecker({ languages: ["en"] });
function CommentBox() {
const [value, setValue] = useState("");
const bad = useMemo(() => checker.isProfane(value), [value]);
return (
<>
<textarea value={value} onChange={(e) => setValue(e.target.value)} />
{bad && <p>Please keep it clean.</p>}
</>
);
}
Express middleware
import { createChecker } from "@bleepit/core";
const checker = createChecker();
app.post("/comments", (req, res, next) => {
if (checker.isProfane(String(req.body.text ?? ""))) {
return res.status(422).json({ error: "Inappropriate content." });
}
next();
});
Deno
import { isProfane } from "npm:@bleepit/core";
Deno.serve((req) => {
const url = new URL(req.url);
const clean = !isProfane(url.searchParams.get("q") ?? "");
return Response.json({ clean });
});
@bleepit/ocr — images
Profanity in an image is profanity. @bleepit/ocr runs your OCR
engine over a page, scans the recognized text with @bleepit/core, and hands back each
match with the bounding boxes it came from — so you can redact, blur, or route for review.
Zero runtime dependencies, same as core. It ships no OCR engine: you bring one, and it can be local WASM or a cloud API.
npm i @bleepit/ocr @bleepit/core
pnpm add @bleepit/ocr @bleepit/core
bun add @bleepit/ocr @bleepit/core
deno add npm:@bleepit/ocr npm:@bleepit/core
import { createChecker } from "@bleepit/core";
import { createImageChecker } from "@bleepit/ocr";
const ic = createImageChecker({
engine: myOcrEngine,
checker: createChecker({ languages: ["en", "es"] }),
});
await ic.isProfane(image); // boolean
await ic.find(image); // ImageMatch[] — word, text, boxes, confidence
await ic.redact(image); // BBox[] — one merged box per match
Image API
ImageProfanityChecker
| Method | Signature | Description |
|---|---|---|
isProfane | (image) => Promise<boolean> | Boolean check on a recognized page. |
find | (image) => Promise<ImageMatch[]> | Each match with boxes, words, text and confidence. |
redact | (image) => Promise<BBox[]> | One merged box per match — geometry only, nothing is drawn. |
findInWords | (words) => ImageMatch[] | Synchronous, for when OCR already ran elsewhere in your pipeline. |
Options
| Option | Type | Default | Description |
|---|---|---|---|
engine | OcrEngine | — | Required. The OCR backend; this package ships none. |
checker | ProfanityCheckerLike | English ProfanityChecker | Pass your own to configure languages or wordlists. |
minConfidence | number | 60 | Drop OCR words below this confidence (0–100) before matching. |
crossWord | boolean | false | Allow a match to span more than one OCR word. See the tradeoff below. |
Bring your own engine
An engine is any object with a recognize method returning positioned
words. That is the whole contract — adapting a cloud OCR API means mapping its response into this shape.
interface OcrEngine {
recognize(image: ImageInput): Promise<{ words: OcrWord[] }>;
terminate?(): Promise<void>;
}
interface OcrWord {
text: string;
bbox: { x0: number; y0: number; x1: number; y1: number };
confidence: number; // 0–100
}
This package never decodes, resizes, or preprocesses images — the input goes straight to the
engine. It also does not draw redactions: redact() returns geometry, because
rasterizing needs a canvas or an image library, and that dependency is yours to choose.
Image accuracy
Word fragments and the crossWord tradeoff
bleepit strips non-alphanumerics before matching — that is what catches
a.r.s.e. The consequence for OCR is that any gap between two recognized words
disappears, so "ar" and "se" in adjacent
boxes would scan as one word. No separator character avoids this; only a letter would, and injecting letters corrupts
offsets.
So each OCR word is scanned on its own by default. Neighbouring words cannot collide, at the cost of missing profanity that OCR split across two boxes. Flip it when a split word is the likelier failure:
createImageChecker({ engine, crossWord: true });
Expect more false positives in that mode — it is the right choice for noisy scans of stylized type, and the wrong one for dense screenshots of prose.
OCR noise raises false positives
bleepit normalizes leet-speak, mapping 1→i and
0→o. OCR makes the same confusions, so the normalizer silently repairs a lot of
recognition error — pr1ck still matches. The cost is that OCR garbage also normalizes
toward dictionary words, and flags that would never fire on typed input.
minConfidence (default 60)
is the main lever. Raise it for photographs and scans; lower it for clean screenshots where a miss costs more than a
false positive. Tune it against images that look like yours — the right value depends on your engine and your inputs,
not on a general default.
- Text only. This finds profane words in an image. It does not detect offensive imagery, gestures, or symbols — that is an image classifier, and a different problem with different failure modes.
- A miss is not a guarantee. Stylized type, low contrast, curved text, and handwriting all defeat OCR. Treat a clean result as “nothing recognized,” not “nothing there.”
How it works
- Normalize — one pass folds case, diacritics (NFKD), ligatures (
ß→ss), and leet-speak, drops separators, and records original offsets. - Scan — an Aho-Corasick automaton matches all patterns in a single
O(n)pass, independent of dictionary size. - Repeat tolerance — a repeated char takes a direct trie transition when one exists (keeps
as ≠ ass) and is otherwise skipped (catchesbollllocks). - Validate — boundaries and whitelist are checked against the original text, so
a.r.s.eflags whileclassdoes not.
Performance
Measured on the minified bundle (Node 22, 123 patterns): ~5.5 ms per 97 KB (~17 MB/s). Tips:
- Reuse one instance — construction builds the automaton once; queries share it.
- Use
isProfanefor boolean gates — it early-exits on the first valid match. - Disable
stripDiacriticsif your input is plain ASCII. - Adding languages or thousands of custom words costs nothing at query time.
Limitations
Single-character masking (a*se, b*llocks) is intentionally out of scope: a * can stand for any letter, so reliable matching needs edit-distance search, which would break the lightweight / O(n) guarantees. Workaround: add the variants you care about via customWords.
Coverage is the lists you load. There is no language detection — a word that is not in a loaded list or in customWords is never flagged. Four starter lists ship (en, es, fr, de), and everything else needs adding explicitly. That includes close cousins of a language you have loaded: Portuguese merda is not matched by Spanish mierda, and Swedish skit is not matched by English shit.
No filter is perfect — pair automated checks with reporting and human review for high-stakes moderation, especially for communities whose languages you do not speak.
FAQ
Does it work in the browser / workers / edge runtimes?
Yes. Zero dependencies and no Node APIs — it runs anywhere JavaScript runs, including Cloudflare Workers and Deno Deploy.
Why is “Scunthorpe” clean but “a.r.s.e” flagged?
Boundaries are checked against the original text. Letters block a match (class), while separators are obfuscation and get skipped.
Can I add my own language?
Yes — customWords / addWords() accept any unicode script, and non-Latin text is matched case-insensitively. You do have to add the words: only en, es, fr and de ship as built-in lists, and there is no detection for anything else.
What about Chinese / Japanese / Thai (no spaces)?
Use wholeWord: false for those instances, since spaceless scripts have no word boundaries to check.
Does censoring preserve positions?
find() returns original-string offsets so you can highlight, redact, or audit matches yourself; censor() masks in place and preserves whitespace.
How big is it?
Core bundle is ~6 kB minified (~2.7 kB gzip). Extra language lists are separate entrypoints (@bleepit/core/lists) so you only ship what you load.