Skip to content

bleepit

Keep profanity out of your product. A fast, free profanity checker that plugs into anywhere JavaScript runs — Node, browsers, Deno, Bun, and workers — with zero dependencies and wordlists you control.

PackageWhat it does
@bleepit/core Profanity checking for text. The engine everything else builds on.
@bleepit/ocr Profanity checking for images: OCR the page, map matches back to bounding boxes.
npm i @bleepit/core
pnpm add @bleepit/core
bun add @bleepit/core
deno add npm:@bleepit/core

Live demo

Runs 100% in your browser using the real library bundle — nothing is sent anywhere.

Checker options
CLEAN
wordstartend

Installation

npm i @bleepit/core
pnpm add @bleepit/core
bun add @bleepit/core
deno add npm:@bleepit/core

Ships tree-shakeable ESM + CJS with types. Import only what you need:

import { ProfanityChecker } from "@bleepit/core";
import { es } from "@bleepit/core/lists"; // individual lists for tiny bundles

Quickstart

import { ProfanityChecker } from "@bleepit/core";

const checker = new ProfanityChecker({ languages: ["en", "es"] });

checker.isProfane("What an a.r.s.e?!");   // true
checker.isProfane("arrrse");              // true (elongation)
checker.isProfane("@rse");                // true (leet-speak)
checker.isProfane("the class");           // false (word boundaries)
checker.censor("That was a d0uche move"); // "That was a ****** move"

API reference

ProfanityChecker

MethodSignatureDescription
isProfane(text) => booleanFast boolean check, early-exits on first valid match.
find(text, limit?) => Match[]All matches as { word, start, end } with original-string offsets (UTF-16).
censor(text, mask="*") => stringMasks profane spans; whitespace inside a span is preserved.
addWords(words) => voidAdds words (any language/script) and rebuilds the automaton.
removeWords(words) => voidRemoves words and rebuilds the automaton.
sizenumberCount of normalized patterns loaded.

Factories & singleton

import { createChecker, isProfane, find, censor } from "@bleepit/core";

const mine = createChecker({ languages: ["fr"] }); // isolated instance
isProfane("bollocks"); // uses the shared default (English) instance

Options

OptionTypeDefaultDescription
languagesstring[]["en"]Built-in lists to load (en, es, fr, de).
customWordsstring[][]Extra words in any language or script.
whiteliststring[][]Words that must never be flagged (matched against the enclosing word).
wholeWordbooleantrueMatch on word boundaries only. Set false for aggressive filtering or spaceless scripts (e.g. CJK).
leetbooleantrueMap leet-speak (@→a, 0→o, $→s…).
stripDiacriticsbooleantrueFold diacritics (é→e) and ligatures (ß→ss). Disable for max speed on plain ASCII.

Languages

Four starter lists ship with the package. Anything else is customWords — the engine itself is script-agnostic, so the words you add can be in any language or writing system.

CodeLanguageWords
enEnglish (default)66
esSpanish24
frFrench17
deGerman16
import { createChecker } from "@bleepit/core";

// Not built in — bring your own words, in any script:
const c = createChecker({ languages: [], customWords: ["чёрт", "クソ"] });
c.isProfane("вот ЧЁРТ"); // true (case-insensitive Cyrillic)

// Spaceless scripts have no word boundaries — use substring mode:
const zh = createChecker({ languages: [], customWords: ["笨蛋"], wholeWord: false });

Guides

Word boundaries (the Scunthorpe problem)

With wholeWord: true (default), matches must sit on word boundaries checked against the original text, so separators still split words while a.r.s.e is caught:

checker.isProfane("the class");      // false
checker.isProfane("Scunthorpe");     // false
checker.isProfane("arse about");     // true
checker.isProfane("a.r.s.e about");  // true

Whitelisting

const c = createChecker({ whitelist: ["arsenal"] });
c.isProfane("arsenal"); // false

Runtime updates

checker.addWords(["heck"]);       // community-driven lists
checker.removeWords(["bollocks"]); // tune false positives away

Censoring

checker.censor("a d0uche move", "*"); // "a ****** move"
checker.censor("oh bollocks", "#");   // "oh ########"

Integrations

React (comment box)

import { useMemo, useState } from "react";
import { createChecker } from "@bleepit/core";

const checker = createChecker({ languages: ["en"] });

function CommentBox() {
  const [value, setValue] = useState("");
  const bad = useMemo(() => checker.isProfane(value), [value]);
  return (
    <>
      <textarea value={value} onChange={(e) => setValue(e.target.value)} />
      {bad && <p>Please keep it clean.</p>}
    </>
  );
}

Express middleware

import { createChecker } from "@bleepit/core";

const checker = createChecker();
app.post("/comments", (req, res, next) => {
  if (checker.isProfane(String(req.body.text ?? ""))) {
    return res.status(422).json({ error: "Inappropriate content." });
  }
  next();
});

Deno

import { isProfane } from "npm:@bleepit/core";

Deno.serve((req) => {
  const url = new URL(req.url);
  const clean = !isProfane(url.searchParams.get("q") ?? "");
  return Response.json({ clean });
});

@bleepit/ocr — images

Profanity in an image is profanity. @bleepit/ocr runs your OCR engine over a page, scans the recognized text with @bleepit/core, and hands back each match with the bounding boxes it came from — so you can redact, blur, or route for review.

Zero runtime dependencies, same as core. It ships no OCR engine: you bring one, and it can be local WASM or a cloud API.

npm i @bleepit/ocr @bleepit/core
pnpm add @bleepit/ocr @bleepit/core
bun add @bleepit/ocr @bleepit/core
deno add npm:@bleepit/ocr npm:@bleepit/core
import { createChecker } from "@bleepit/core";
import { createImageChecker } from "@bleepit/ocr";

const ic = createImageChecker({
  engine: myOcrEngine,
  checker: createChecker({ languages: ["en", "es"] }),
});

await ic.isProfane(image); // boolean
await ic.find(image);      // ImageMatch[] — word, text, boxes, confidence
await ic.redact(image);    // BBox[] — one merged box per match

Image API

ImageProfanityChecker

MethodSignatureDescription
isProfane(image) => Promise<boolean>Boolean check on a recognized page.
find(image) => Promise<ImageMatch[]>Each match with boxes, words, text and confidence.
redact(image) => Promise<BBox[]>One merged box per match — geometry only, nothing is drawn.
findInWords(words) => ImageMatch[]Synchronous, for when OCR already ran elsewhere in your pipeline.

Options

OptionTypeDefaultDescription
engineOcrEngine—Required. The OCR backend; this package ships none.
checkerProfanityCheckerLikeEnglish ProfanityCheckerPass your own to configure languages or wordlists.
minConfidencenumber60Drop OCR words below this confidence (0–100) before matching.
crossWordbooleanfalseAllow a match to span more than one OCR word. See the tradeoff below.

Bring your own engine

An engine is any object with a recognize method returning positioned words. That is the whole contract — adapting a cloud OCR API means mapping its response into this shape.

interface OcrEngine {
  recognize(image: ImageInput): Promise<{ words: OcrWord[] }>;
  terminate?(): Promise<void>;
}

interface OcrWord {
  text: string;
  bbox: { x0: number; y0: number; x1: number; y1: number };
  confidence: number; // 0–100
}

This package never decodes, resizes, or preprocesses images — the input goes straight to the engine. It also does not draw redactions: redact() returns geometry, because rasterizing needs a canvas or an image library, and that dependency is yours to choose.

Image accuracy

Word fragments and the crossWord tradeoff

bleepit strips non-alphanumerics before matching — that is what catches a.r.s.e. The consequence for OCR is that any gap between two recognized words disappears, so "ar" and "se" in adjacent boxes would scan as one word. No separator character avoids this; only a letter would, and injecting letters corrupts offsets.

So each OCR word is scanned on its own by default. Neighbouring words cannot collide, at the cost of missing profanity that OCR split across two boxes. Flip it when a split word is the likelier failure:

createImageChecker({ engine, crossWord: true });

Expect more false positives in that mode — it is the right choice for noisy scans of stylized type, and the wrong one for dense screenshots of prose.

OCR noise raises false positives

bleepit normalizes leet-speak, mapping 1→i and 0→o. OCR makes the same confusions, so the normalizer silently repairs a lot of recognition error — pr1ck still matches. The cost is that OCR garbage also normalizes toward dictionary words, and flags that would never fire on typed input.

minConfidence (default 60) is the main lever. Raise it for photographs and scans; lower it for clean screenshots where a miss costs more than a false positive. Tune it against images that look like yours — the right value depends on your engine and your inputs, not on a general default.

  • Text only. This finds profane words in an image. It does not detect offensive imagery, gestures, or symbols — that is an image classifier, and a different problem with different failure modes.
  • A miss is not a guarantee. Stylized type, low contrast, curved text, and handwriting all defeat OCR. Treat a clean result as “nothing recognized,” not “nothing there.”

How it works

  1. Normalize — one pass folds case, diacritics (NFKD), ligatures (ß→ss), and leet-speak, drops separators, and records original offsets.
  2. Scan — an Aho-Corasick automaton matches all patterns in a single O(n) pass, independent of dictionary size.
  3. Repeat tolerance — a repeated char takes a direct trie transition when one exists (keeps as ≠ ass) and is otherwise skipped (catches bollllocks).
  4. Validate — boundaries and whitelist are checked against the original text, so a.r.s.e flags while class does not.

Performance

Measured on the minified bundle (Node 22, 123 patterns): ~5.5 ms per 97 KB (~17 MB/s). Tips:

  • Reuse one instance — construction builds the automaton once; queries share it.
  • Use isProfane for boolean gates — it early-exits on the first valid match.
  • Disable stripDiacritics if your input is plain ASCII.
  • Adding languages or thousands of custom words costs nothing at query time.

Limitations

Single-character masking (a*se, b*llocks) is intentionally out of scope: a * can stand for any letter, so reliable matching needs edit-distance search, which would break the lightweight / O(n) guarantees. Workaround: add the variants you care about via customWords.

Coverage is the lists you load. There is no language detection — a word that is not in a loaded list or in customWords is never flagged. Four starter lists ship (en, es, fr, de), and everything else needs adding explicitly. That includes close cousins of a language you have loaded: Portuguese merda is not matched by Spanish mierda, and Swedish skit is not matched by English shit.

No filter is perfect — pair automated checks with reporting and human review for high-stakes moderation, especially for communities whose languages you do not speak.

FAQ

Does it work in the browser / workers / edge runtimes?

Yes. Zero dependencies and no Node APIs — it runs anywhere JavaScript runs, including Cloudflare Workers and Deno Deploy.

Why is “Scunthorpe” clean but “a.r.s.e” flagged?

Boundaries are checked against the original text. Letters block a match (class), while separators are obfuscation and get skipped.

Can I add my own language?

Yes — customWords / addWords() accept any unicode script, and non-Latin text is matched case-insensitively. You do have to add the words: only en, es, fr and de ship as built-in lists, and there is no detection for anything else.

What about Chinese / Japanese / Thai (no spaces)?

Use wholeWord: false for those instances, since spaceless scripts have no word boundaries to check.

Does censoring preserve positions?

find() returns original-string offsets so you can highlight, redact, or audit matches yourself; censor() masks in place and preserves whitespace.

How big is it?

Core bundle is ~6 kB minified (~2.7 kB gzip). Extra language lists are separate entrypoints (@bleepit/core/lists) so you only ship what you load.