# Curious behaviour in Char

**URL:** <https://discourse.elm-lang.org/t/curious-behaviour-in-char/2496>\
**Category:** Learn\
**Created:** [November 11, 2018, 1:30pm UTC](https://discourse.elm-lang.org/t/curious-behaviour-in-char/2496 "2018-11-11T13:30:28Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![David\_Legard](https://avatars.discourse-cdn.com/v4/letter/d/e79b87/32.png) [@David\_Legard](https://discourse.elm-lang.org/u/David_Legard)\
**Post date:** [November 11, 2018, 1:30pm UTC](https://discourse.elm-lang.org/t/curious-behaviour-in-char/2496/1 "2018-11-11T13:30:28Z")

</div>

I am using some characters outside the main ASCII set, such as Ñ and ñ, and need to differentiate between the lower and upper-case variants.

The documentation says that Char.isUpper only works on ASCII characters, so sure enough, elm -repl gives me:

```
Char.isUpper 'Ñ'
False : Bool

```

however, if I upcase ñ, we get

```
Char.toUpper 'ñ'
'Ñ' : Char.Char

```

That seems inconsistent. It is equivalent to saying:

`Char.isUpper (Char.toUpper somechar) = false`

Surely, both should work or neither.

Any thoughts?

---

<div class="post-metadata">

**Author:** ![glennsl](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.elm-lang.org/glennsl/32/1496_2.png) [@glennsl](https://discourse.elm-lang.org/u/glennsl)\
**Post date:** [November 11, 2018, 2:07pm UTC](https://discourse.elm-lang.org/t/curious-behaviour-in-char/2496/2 "2018-11-11T14:07:19Z")

</div>

> [@David\_Legard](#):
>
> That seems inconsistent. It is equivalent to saying:
> 
> `Char.isUpper (Char.toUpper somechar) = false`

It isn’t though, since `Char.isUpper` is defined to only return true for uppercase ASCII characters. You could instead argue that it isn’t accurately named, which I would agree with.

Unicode case mapping is a non-trivial problem because there isn’t a one-to-one mapping between lowercase and uppercase character. You can [see some of the issues with it here](https://www.unicode.org/faq/casemap_charprop.html). And to add to that, [a unicode “character” might not even be what you expect it to be](https://stackoverflow.com/questions/27331819/whats-the-difference-between-a-character-a-code-point-a-glyph-and-a-grapheme).

But even if someone does implement proper unicode support for `Char.isUpper` wouldn’t it be inconsistent that most other functions don’t fully or properly handle unicode? If consistency is the goal, you might be looking at a very big task then.

I do hope there will be proper unicode support sometime in the future, since it’s really nice to not have to worry about unicode issues and having a good static type system might help a lot in designing a good API for it. But for the above-mentioned issues I don’t expect to see it in the near future.

---

<div class="post-metadata">

**Author:** ![jwoLondon](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.elm-lang.org/jwolondon/32/605_2.png) [@jwoLondon](https://discourse.elm-lang.org/u/jwoLondon)\
**Post date:** [November 11, 2018, 2:14pm UTC](https://discourse.elm-lang.org/t/curious-behaviour-in-char/2496/3 "2018-11-11T14:14:37Z")

</div>

Is there any reason not to use a homegrown version?

```elm
isUpper : Char -> Bool
isUpper c =
    c == Char.toUpper c

```

Presumably there must be a reason why this isn’t how `Char.isUpper` works, but it would seem to do the job, at least for a wider range of characters than ASCII.

---

<div class="post-metadata">

**Author:** ![glennsl](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.elm-lang.org/glennsl/32/1496_2.png) [@glennsl](https://discourse.elm-lang.org/u/glennsl)\
**Post date:** [November 11, 2018, 2:32pm UTC](https://discourse.elm-lang.org/t/curious-behaviour-in-char/2496/4 "2018-11-11T14:32:27Z")

</div>

Since, as I said, there isn’t a one-to-one mapping, this wouldn’t work for all unicode characters. It would as you say work for a wider range of characters, but that range wouldn’t be well-defined. And I for one would rather know for certain when it doesn’t work, than having it suddenly not work as expected. I can always just implement this hack myself, and would by doing so hopefully have a better understanding (or care less) of when it works.

---

<div class="post-metadata">

**Author:** ![David\_Legard](https://avatars.discourse-cdn.com/v4/letter/d/e79b87/32.png) [@David\_Legard](https://discourse.elm-lang.org/u/David_Legard)\
**Post date:** [November 12, 2018, 1:59am UTC](https://discourse.elm-lang.org/t/curious-behaviour-in-char/2496/5 "2018-11-12T01:59:11Z")

</div>

I reckon the name is profoundly misleading. The idea that `Char.toUpper` does not produce an uppercase character as defined by `Char.isUpper` is going to catch out other people, for sure.

It turns out to be simple to implement this for extended ASCII, together with a custom ordering (A,B,C,D,Ð,E… M,N,Ñ,O …) .

Thanks for the responses.

---

<div class="post-metadata">

**Author:** ![malaire](https://avatars.discourse-cdn.com/v4/letter/m/b782af/32.png) [@malaire](https://discourse.elm-lang.org/u/malaire)\
**Post date:** [November 12, 2018, 2:35am UTC](https://discourse.elm-lang.org/t/curious-behaviour-in-char/2496/6 "2018-11-12T02:35:37Z")

</div>

> [@David\_Legard](#):
>
> It turns out to be simple to implement this for extended ASCII

There is no “extended ASCII” in Unicode. If you mean Latin Extended, then which extensions you mean, as there are quite many? Supporting Cyrillic would also be simple, so why support only Latin?

I think it only makes sense to support either only ASCII or full Unicode officially. When trying to support something in between, it’s going to be quite difficult to decide what exactly will be supported.

---

<div class="post-metadata">

**Author:** ![David\_Legard](https://avatars.discourse-cdn.com/v4/letter/d/e79b87/32.png) [@David\_Legard](https://discourse.elm-lang.org/u/David_Legard)\
**Post date:** [November 12, 2018, 4:27am UTC](https://discourse.elm-lang.org/t/curious-behaviour-in-char/2496/7 "2018-11-12T04:27:14Z")

</div>

Latin Extended, then.

I don’t know its formal name, but it’s the one that comes up in Windows’ Character Map when you first open it.

I’m sure you’re right about the difficulties of supporting Unicode.

My point is, that defining `Char.toUpper` in such a way that does not produce an uppercase character as defined by `Char.isUpper` for many characters seems misleading and is probably going to catch out other people.

EDIT: It’s called ISO/IEC 8859-1

---

<div class="post-metadata">

**Author:** ![Qqwy](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.elm-lang.org/qqwy/32/1062_2.png) [@Qqwy](https://discourse.elm-lang.org/u/Qqwy)\
**Post date:** [November 12, 2018, 8:01am UTC](https://discourse.elm-lang.org/t/curious-behaviour-in-char/2496/8 "2018-11-12T08:01:50Z")

</div>

I think it would be nice to have a library that, for each Unicode grapheme, is able to return if it is in a given [Unicode category](https://www.fileformat.info/info/unicode/category/index.htm) (As well as the reverse: looking up all categories for the grapheme). These categories include things like ‘lowercase’, ‘uppercase’, ‘titlecase’, etc.

Such a library would probably best be created in a metaprogramming fashion , similar [as to what Elixir does](https://github.com/elixir-lang/elixir/blob/master/lib/elixir/unicode/unicode.ex). (it is unfortunate that this cannot be written in Elm itself 😅, but writing something like it in either JS or Haskell should not be too much trouble.)

---

<div class="post-metadata">

**Author:** ![malaire](https://avatars.discourse-cdn.com/v4/letter/m/b782af/32.png) [@malaire](https://discourse.elm-lang.org/u/malaire)\
**Post date:** [November 12, 2018, 11:08am UTC](https://discourse.elm-lang.org/t/curious-behaviour-in-char/2496/9 "2018-11-12T11:08:38Z")

</div>

> [@David\_Legard](#):
>
> My point is, that defining `Char.toUpper` in such a way that does not produce an uppercase character as defined by `Char.isUpper` for many characters seems misleading and is probably going to catch out other people.

I think that best fix for now would be to just rename ASCII functions like `Char.isUpper` to `Char.isUpperAscii` so there is no confusion.

---

<div class="post-metadata">

**Author:** ![malaire](https://avatars.discourse-cdn.com/v4/letter/m/b782af/32.png) [@malaire](https://discourse.elm-lang.org/u/malaire)\
**Post date:** [November 12, 2018, 11:12am UTC](https://discourse.elm-lang.org/t/curious-behaviour-in-char/2496/10 "2018-11-12T11:12:28Z")

</div>

> [@Qqwy](#):
>
> I think it would be nice to have a library that, for each Unicode grapheme, is able to return if it is in a given [Unicode category](https://www.fileformat.info/info/unicode/category/index.htm) (As well as the reverse: looking up all categories for the grapheme). These categories include things like ‘lowercase’, ‘uppercase’, ‘titlecase’, etc.

I just wonder how large such a library would be, containing data for all 137439 characters (as of Unicode 11.0). But yes proper library is much better solution that trying to fix current functions one-by-one. And that also leaves faster ASCII versions available for those who need them.

---

<div class="post-metadata">

**Author:** ![Qqwy](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.elm-lang.org/qqwy/32/1062_2.png) [@Qqwy](https://discourse.elm-lang.org/u/Qqwy)\
**Post date:** [November 12, 2018, 12:29pm UTC](https://discourse.elm-lang.org/t/curious-behaviour-in-char/2496/11 "2018-11-12T12:29:05Z")

</div>

> [@malaire](#):
>
> I just wonder how large such a library would be, containing data for all 137439 characters (as of Unicode 11.0)

If you only check the basic Unicode categories, you do not need a function head per character, since many characters are grouped by category (so you can check if the input codepoint is in a given range). This kind of ‘size reduction’ is already apparent in the [Unicode PropList.txt file](https://github.com/elixir-lang/elixir/blob/master/lib/elixir/unicode/PropList.txt).

---

<div class="post-metadata">

**Author:** ![malaire](https://avatars.discourse-cdn.com/v4/letter/m/b782af/32.png) [@malaire](https://discourse.elm-lang.org/u/malaire)\
**Post date:** [November 12, 2018, 12:57pm UTC](https://discourse.elm-lang.org/t/curious-behaviour-in-char/2496/12 "2018-11-12T12:57:59Z")

</div>

Unicode is hard… I was reading a bit about Unicode case mappings ([section 5.18 here](https://www.unicode.org/versions/Unicode11.0.0/ch05.pdf) if interested), and I noticed a bug in `Char.toUpper` where `Char.toUpper('ß')` returns two characters as single `Char`: [https://github.com/elm/core/issues/1001](https://github.com/elm/core/issues/1001)

---

<div class="post-metadata">

**Author:** ![malaire](https://avatars.discourse-cdn.com/v4/letter/m/b782af/32.png) [@malaire](https://discourse.elm-lang.org/u/malaire)\
**Post date:** [November 12, 2018, 4:10pm UTC](https://discourse.elm-lang.org/t/curious-behaviour-in-char/2496/13 "2018-11-12T16:10:25Z")

</div>

> [@David\_Legard](#):
>
> My point is, that defining `Char.toUpper` in such a way that does not produce an uppercase character as defined by `Char.isUpper` for many characters seems misleading and is probably going to catch out other people.
> 
> EDIT: It’s called ISO/IEC 8859-1

I just realized that the problematic character `ß` I mentioned above is in ISO/IEC 8859-1, so even supporting that set of characters is not trivial.

Haskell has solved this problem so that `Data.Char.toUpper` of type `Char -> Char` just returns `ß` unchanged, and other function `Data.Text.toUpper` of type `Text -> Text` does proper conversion to `SS`.

But this does mean that in Haskell

```
Data.Char.isUpper(Data.Char.toUpper('ß')) == False

```

even though that `Data.Char.isUpper` is for all Unicode characters and not just ASCII.

---

<div class="post-metadata">

**Author:** ![system](https://canada1.discourse-cdn.com/flex035/uploads/elm_lang/original/1X/50a05e53677a2c3b47776d7abd0f113eb50193a1.png) [@system](https://discourse.elm-lang.org/u/system)\
**Post date:** [November 22, 2018, 4:10pm UTC](https://discourse.elm-lang.org/t/curious-behaviour-in-char/2496/14 "2018-11-22T16:10:33Z")

</div>

This topic was automatically closed 10 days after the last reply. New replies are no longer allowed.
