# `Unicode.CharacterName`
[🔗](https://github.com/elixir-unicode/unicode/blob/v2.2.0/lib/unicode/character_name.ex#L1)

Resolves Unicode character names to their codepoint, and back again.

Names are taken from the `Name` field of the Unicode Character Database
(`UnicodeData.txt`). `to_codepoint/1` matches them loosely: case, whitespace,
`_` and `-` are ignored (as in `\N{...}` name lookups). Codepoints whose name
is a bracketed label such as `<control>` have no `Name` property, so `to_name/1`
does not resolve them, but the names they are known by are recorded as aliases
and `to_codepoint/1` does.

### Name aliases

`NameAliases.txt` gives additional names for the characters that need them,
because a `Name` can never change once published: the control characters, which
have no `Name`, and the characters whose published name contains an error.
`to_codepoint/1` resolves all five alias types, so `NULL`, `LINE FEED`, `LF`,
`BYTE ORDER MARK` and `LATIN CAPITAL LETTER GHA` each name their character.

Aliases are consulted only after the `Name` property and the derivation rules,
so a name always wins over an alias spelled the same way. `aliases/1` returns
the aliases of a codepoint with their types.

The names are prefix-compressed (front-coded) into a single sorted binary blob
with block restart points, and looked up with a binary search over the restart
names followed by a scan-decode within one block. This keeps the table compact
without materialising a large map.

`to_name/1` is the reverse lookup. Each name is stored **once**, in normalized
form: rather than keeping a second copy of the original, the original is
reconstructed by upper-casing and reinserting separators recorded at about 3
bytes per name. That is exact because the `Name` property contains no lower
case. Reverse lookup therefore costs a codepoint index and those separators,
not a second name table, and leaves `to_codepoint/1` comparing whole binaries
as before.

### Derived names

CJK ideographs, Tangut ideographs, Hangul syllables and the Seal and Jurchen
characters are recorded in `UnicodeData.txt` as `<..., First>`/`<..., Last>`
range pairs with no per-character name, and their names are instead derived by
the rules in [UAX #44](https://www.unicode.org/reports/tr44/#Name). These
resolve, but are not part of the name table: they are computed from 15 range
tuples and the Hangul jamo short names, under 4KB in total. Materialising them
would add more than 131,000 names — over three times the size of the whole
listed table — for characters whose names follow a rule.

Derivation is attempted only when the table lookup misses, so it costs nothing
on the common path.

# `aliases`
*since 2.2.0* 

```elixir
@spec aliases(non_neg_integer()) :: [{atom(), String.t()}]
```

Returns the `Name_Alias` values of a codepoint.

A character's `Name` can never change once published, so `NameAliases.txt` carries the additional
names a character needs: the control characters, which have no `Name` at all, and the characters
whose published name contains an error.

### Arguments

* `codepoint` is a codepoint in the range `0..0x10FFFF`.

### Returns

* A list of `{type, name}` tuples in the order the Unicode Character Database lists them, where
  `type` is one of `:correction`, `:control`, `:alternate`, `:figment` or `:abbreviation`.

* An empty list if the codepoint has no aliases.

### Examples

    iex> Unicode.CharacterName.aliases(0x0000)
    [control: "NULL", abbreviation: "NUL"]

    iex> Unicode.CharacterName.aliases(0x01A2)
    [correction: "LATIN CAPITAL LETTER GHA"]

    iex> Unicode.CharacterName.aliases(0xFEFF)
    [alternate: "BYTE ORDER MARK", abbreviation: "BOM", abbreviation: "ZWNBSP"]

    iex> Unicode.CharacterName.aliases(?A)
    []

# `count`

```elixir
@spec count() :: non_neg_integer()
```

Returns the number of names in the table.

# `to_codepoint`

```elixir
@spec to_codepoint(String.t(), Keyword.t()) :: {:ok, pos_integer()} | :error
```

Returns the codepoint for a Unicode character name.

### Arguments

* `name` is a Unicode character name as a string, matched loosely.

* `options` is a keyword list of options.

### Options

* `:fuzzy` enables approximate matching when the name is not found exactly. The value is either
  `true`, meaning use the default Jaro distance of `0.8`, or a number
  between `0.0` and `1.0` giving the minimum distance to accept. The default is `false`, meaning
  exact matching only.

### Returns

* `{:ok, codepoint}` or

* `:error` if the name is not known, if a fuzzy search found no single best match, or if the
  `:fuzzy` option is not one of the forms above.

### Fuzzy matching

A fuzzy search succeeds only when it resolves to one name: the closest name by
`String.jaro_distance/2` must be at least as close as the threshold *and* strictly closer than
every other name. A tie is treated as unresolved and returns `:error`, so an ambiguous query
never silently picks one of several candidates.

Because the threshold is a floor rather than a filter, it does not need to exclude the many names
that are similar to any given query — of the roughly forty thousand names, `LATIN SMALL LETTER B`
is close to `LATIN SMALL LETTER A` but is not the closest.

Matching is against the listed names only; algorithmically derived names such as
`CJK UNIFIED IDEOGRAPH-4E00` are not fuzzy-matched, since a near miss on the hexadecimal part
would name a different character. Name aliases are matched exactly for the same reason: many are
abbreviations of two or three letters, where a single character difference is another abbreviation
rather than a typo.

Fuzzy matching scans every name and is several thousand times slower than an exact lookup. It is
only attempted after an exact lookup has failed, so supplying the option costs nothing when the
name is correct.

### Examples

    iex> Unicode.CharacterName.to_codepoint("LATIN SMALL LETTER A")
    {:ok, 97}

    iex> Unicode.CharacterName.to_codepoint("bullet")
    {:ok, 8226}

    iex> Unicode.CharacterName.to_codepoint("Not A Real Name")
    :error

    iex> Unicode.CharacterName.to_codepoint("NULL")
    {:ok, 0}

    iex> Unicode.CharacterName.to_codepoint("LF")
    {:ok, 10}

    iex> Unicode.CharacterName.to_codepoint("LATIN SMALL LETER A", fuzzy: true)
    {:ok, 97}

    iex> Unicode.CharacterName.to_codepoint("GRINING FACE", fuzzy: 0.9)
    {:ok, 128512}

    iex> Unicode.CharacterName.to_codepoint("LATIN SMALL LETER A")
    :error

# `to_name`
*since 2.1.0* 

```elixir
@spec to_name(non_neg_integer()) :: {:ok, String.t()} | :error
```

Returns the Unicode character name for a codepoint.

### Arguments

* `codepoint` is a codepoint in the range `0..0x10FFFF`.

### Returns

* `{:ok, name}` where `name` is the `Name` property of the codepoint, or

* `:error` if the codepoint has no name. That includes control characters, surrogates, private
  use characters and unassigned codepoints, whose `Name` property is empty. `aliases/1` returns
  the names a control character is known by, which is what `to_codepoint/1` resolves.

### Notes

Reverse lookup shares the name table with `to_codepoint/1` rather than keeping its own copy of
every name. It adds a 6 byte per name codepoint index and about 3 bytes per name of separator
positions. Names that follow a derivation rule are computed instead of stored and cost nothing.

Where two names differ only by a hyphen that loose matching removes, both resolve here to their
own name even though only one of them is reachable through `to_codepoint/1`.

### Examples

    iex> Unicode.CharacterName.to_name(0x4E00)
    {:ok, "CJK UNIFIED IDEOGRAPH-4E00"}

    iex> Unicode.CharacterName.to_name(0xAC00)
    {:ok, "HANGUL SYLLABLE GA"}

    iex> Unicode.CharacterName.to_name(0x0000)
    :error

---

*Consult [api-reference.md](api-reference.md) for complete listing*
