ARIB STD B24 character set
Volume 1 of the Association of Radio Industries and Businesses (ARIB) STD-B24 standard for Broadcast Markup Language[2] specifies, amongst other details, a character encoding for use in Japanese-language broadcasting. It was introduced on 1999-10-26.[2] The latest revision is version 6.3 as of 2016-07-06.
| Standard | ARIB STB-B24 Volume 1 |
|---|---|
| Classification | ISO 2022 profile/extension |
| Transforms / Encodes | ARIB STB-B24 Kanji, Kana and mosaic sets, JIS X 0201 |
![]() Weather symbols: a few of the extended symbols included. | |
| Language(s) | Japanese, English, Russian Partial support: Greek, Chinese |
|---|---|
| Standard | ARIB STB-B24 Volume 1 |
| Classification | ISO-2022-structured CJK DBCS |
| Extends | JIS X 0208 |
| Encoding formats | |
It includes a number of ARIB extended characters (ARIB外字, ARIB gaiji) not found in the base standards (JIS X 0208 and JIS X 0201). It was the source standard for many symbol characters which were added to Unicode, including portions of the Miscellaneous Symbols, Enclosed Alphanumeric Supplement and Enclosed Ideographic Supplement blocks.[3] Its contributions partially overlap the Unicode emoji, but were added a year earlier, in Unicode 5.2.[4]
Fascicle 1 of the ARIB STD-B62 standard, published in 2014, defines Unicode mappings for a selection of the B24 extended characters (excluding, for example, those duplicated by JIS X 0213), as well as a few extended Kanji.[5] It also includes a mapping of utilised characters outside the Basic Multilingual Plane to the BMP's private use area.
Sets and codes
The ARIB STD B24 standard defines multiple character sets and a method of switching between them. These include a Kanji set (an extension of JIS X 0208), an Alphanumeric set, a Hiragana set, Katakana sets of two distinct layouts and four mosaic sets.[6] The sets are selected using ISO 2022 mechanisms for 94-sets, using the following codes (proportional sets use the same layout as the corresponding non-proportional ones):[7]
| Set | Type | Code (column/line) | Code (hexadecimal) | Code (ASCII character) | Comments |
|---|---|---|---|---|---|
| Kanji | 2-byte | 4/2 | 42 | B | The escape code B used for the ARIB Kanji set[7] is used for the 1983 version of JIS C 6226 (JIS X 0208, of which the ARIB Kanji set is an extension) in ISO-2022-JP.[8][9] |
| Alphanumeric | 1-byte | 4/10 | 4A | J | JIS_C6220-ro (ISO646-JP, JIS X 0201 Roman set). Similar to ASCII, with two assignments differing. Escape code J matches usage in ISO-2022-JP.[9] |
| Proportional alphanumeric | 1-byte | 3/6 | 36 | 6 | |
| Hiragana | 1-byte | 3/0 | 30 | 0 | Hiragana themselves follow the same layout as row 4 of JIS X 0208, but without a lead byte. Also adds several additional assignments for punctuation. |
| Proportional Hiragana | 1-byte | 3/7 | 37 | 7 | |
| Katakana | 1-byte | 3/1 | 31 | 1 | Katakana themselves follow the same layout as row 5 of JIS X 0208, but without a lead byte. Also adds several additional assignments for punctuation. |
| Proportional Katakana | 1-byte | 3/8 | 38 | 8 | |
| JIS X 0201 Katakana | 1-byte | 4/9 | 49 | I | JIS_C6220-jp (JIS X 0201 Kana set). Escape code matches usage in ISO-2022-JP-3. |
| Mosaic A | 1-byte | 3/2 | 32 | 2 | Pseudographics |
| Mosaic B | 1-byte | 3/3 | 33 | 3 | |
| Mosaic C | 1-byte | 3/4 | 34 | 4 | Non-spacing pseudographics |
| Mosaic D | 1-byte | 3/5 | 35 | 5 |
Code charts
Kanji (double-byte) set
This is a double-byte character set extending JIS X 0208.
Lead byte
The encoding bytes correspond to the row or cell number plus 0x20, or 32 in decimal (see below). Hence, the code set starting with 0x21 has a row number of 1, and its cell 1 has a continuation byte of 0x21 (or 33), and so forth. Most of the code corresponds to JIS X 0208.
| ARIB STD-B24 Kanji (double-byte) set (lead bytes) | ||||||||||||||||
| 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | A | B | C | D | E | F | |
| 2x | SP | 1-_ | 2-_ | 3-_ | 4-_ | 5-_ | 6-_ | 7-_ | 8-_ | 9-_ | 10-_ | 11-_ | 12-_ | 13-_ | 14-_ | 15-_ |
| 3x | 16-_ | 17-_ | 18-_ | 19-_ | 20-_ | 21-_ | 22-_ | 23-_ | 24-_ | 25-_ | 26-_ | 27-_ | 28-_ | 29-_ | 30-_ | 31-_ |
| 4x | 32-_ | 33-_ | 34-_ | 35-_ | 36-_ | 37-_ | 38-_ | 39-_ | 40-_ | 41-_ | 42-_ | 43-_ | 44-_ | 45-_ | 46-_ | 47-_ |
| 5x | 48-_ | 49-_ | 50-_ | 51-_ | 52-_ | 53-_ | 54-_ | 55-_ | 56-_ | 57-_ | 58-_ | 59-_ | 60-_ | 61-_ | 62-_ | 63-_ |
| 6x | 64-_ | 65-_ | 66-_ | 67-_ | 68-_ | 69-_ | 70-_ | 71-_ | 72-_ | 73-_ | 74-_ | 75-_ | 76-_ | 77-_ | 78-_ | 79-_ |
| 7x | 80-_ | 81-_ | 82-_ | 83-_ | 84-_ | 85-_ | 86-_ | 87-_ | 88-_ | 89-_ | 90-_ | 91-_ | 92-_ | 93-_ | 94-_ | DEL |
Unused lead byte
Lead byte
Differences from JIS X 0208 | ||||||||||||||||
Character sets 0x21-0x74 (row numbers 1-84: punctuation, alphabets, numbers, Kana, Kanji)
Character set 0x7A (row number 90, traffic symbols)
Characters 90-45 through 90-63 and 90-66 through 90-84 (shown below shaded) are listed in the B24 standard only in table 7-10 (the list of extension characters), and are also the only characters in rows 90 through 91 which are not transport-related symbols; this is noted in the B24 standard in an endnote to table 7-10.[10] The remainder of the extensions are listed in both table 7-4 (the double-byte code chart) and table 7-10.[10]
| ARIB STD-B24 Kanji (double-byte) set (prefixed with 0x7A)[5][11] | ||||||||||||||||
| 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | A | B | C | D | E | F | |
| 2x | ⛌ | ⛍ | ❗︎ | ⛏ | ⛐ | ⛑ | ⛒ | ⛕ | ⛓ | ⛔︎ | ||||||
| 3x | 🅿 | 🆊 | ⛖ | ⛗ | ⛘ | ⛙ | ⛚ | ⛛ | ⛜ | ⛝ | ⛞ | ⛟ | ⛠ | ⛡ | ||
| 4x | ⭕︎ | ㉈ | ㉉ | ㉊ | ㉋ | ㉌ | ㉍ | ㉎ | ㉏ | ⒑ | ⒒ | ⒓ | ||||
| 5x | 🅊 | 🅌 | 🄿 | 🅆 | 🅋 | 🈐 | 🈑 | 🈒 | 🈓 | 🅂 | 🈔 | 🈕 | 🈖 | 🅍 | 🄱 | 🄽 |
| 6x | ⬛︎ | ⬤ | 🈗 | 🈘 | 🈙 | 🈚︎ | 🈛 | ⚿ | 🈜 | 🈝 | 🈞 | 🈟 | 🈠 | 🈡 | 🈢 | 🈣 |
| 7x | 🈤 | 🈥 | 🅎 | ㊙ | 🈀 | |||||||||||
Additions from table 7-10 not in table 7-4.
| ||||||||||||||||
Character set 0x7B (row number 91, map symbols)
Characters from ARIB STD-B24 which were not retained in ARIB STD-B62 are shown shaded.
| ARIB STD-B24 Kanji (double-byte) set (prefixed with 0x7B)[5][11][12] | ||||||||||||||||
| 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | A | B | C | D | E | F | |
| 2x | ⛣ | ⭖ | ⭗ | ⭘ | ⭙ | ☓ | ㊋ | 〒 | ⛨ | ㉆ | ㉅ | ⛩ | ࿖[lower-alpha 1] | ⛪︎ | ⛫ | |
| 3x | ⛬ | ♨ | ⛭ | ⛮ | ⛯ | ⚓︎ | ✈ | ⛰ | ⛱ | ⛲︎ | ⛳︎ | ⛴ | ⛵︎ | 🅗 | Ⓓ | Ⓢ |
| 4x | ⛶ | 🅟 | 🆋 | 🆍 | 🆌 | 🅹 | ⛷ | ⛸ | ⛹ | ⛺︎ | 🅻 | ☎ | ⛻ | ⛼ | ⛽︎ | ⛾ |
| 5x | 🅼 | ⛿ | ||||||||||||||
| 6x | ||||||||||||||||
| 7x | ||||||||||||||||
Not in ARIB STD-B62
| ||||||||||||||||
Character set 0x7C (row number 92, units, enclosed forms, list markers, arrows)
Characters from ARIB STD-B24 which were not retained in ARIB STD-B62 are shown shaded.
| ARIB STD-B24 Kanji (double-byte) set (prefixed with 0x7C)[5][11][12] | ||||||||||||||||
| 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | A | B | C | D | E | F | |
| 2x | ➡ | ⬅ | ⬆ | ⬇ | ⬯ | ⬮ | 年 | 月 | 日 | 円 | ㎡ | ㎥ | ㎝ | ㎠ | ㎤ | |
| 3x | 🄀 | ⒈ | ⒉ | ⒊ | ⒋ | ⒌ | ⒍ | ⒎ | ⒏ | ⒐ | 氏[lower-alpha 2] | 副[lower-alpha 2] | 元[lower-alpha 2] | 故[lower-alpha 2] | 前[lower-alpha 2] | 新[lower-alpha 2] |
| 4x | 🄁 | 🄂 | 🄃 | 🄄 | 🄅 | 🄆 | 🄇 | 🄈 | 🄉 | 🄊 | ㈳ | ㈶ | ㈲ | ㈱ | ㈹ | ㉄ |
| 5x | ▶ | ◀ | 〖 | 〗 | ⟐ | ² | ³ | 🄭 | (vn)[lower-alpha 3] | (ob)[lower-alpha 3] | (cb)[lower-alpha 3] | (ce[lower-alpha 3] | mb)[lower-alpha 3] | (hp)[lower-alpha 3] | (br)[lower-alpha 3] | (p)[lower-alpha 3] |
| 6x | (s)[lower-alpha 3] | (ms)[lower-alpha 3] | (t)[lower-alpha 3] | (bs)[lower-alpha 3] | (b)[lower-alpha 3] | (tb)[lower-alpha 3] | (tp)[lower-alpha 3] | (ds)[lower-alpha 3] | (ag)[lower-alpha 3] | (eg)[lower-alpha 3] | (vo)[lower-alpha 3] | (fl)[lower-alpha 3] | (ke[lower-alpha 3] | y)[lower-alpha 3] | (sa[lower-alpha 3] | x)[lower-alpha 3] |
| 7x | (sy[lower-alpha 3] | n)[lower-alpha 3] | (or[lower-alpha 3] | g)[lower-alpha 3] | (pe[lower-alpha 3] | r)[lower-alpha 3] | 🄬 | 🄫 | ㉇ | 🆐 | 🈦 | ℻ | ||||
Not in ARIB STD-B62
| ||||||||||||||||
Character set 0x7D (row number 93, game and weather symbols, fractions, units, enclosed forms)
Characters from ARIB STD-B24 which were not retained in ARIB STD-B62 are shown shaded.
| ARIB STD-B24 Kanji (double-byte) set (prefixed with 0x7D)[5][11][12] | ||||||||||||||||
| 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | A | B | C | D | E | F | |
| 2x | ㈪ | ㈫ | ㈬ | ㈭ | ㈮ | ㈯ | ㈰ | ㈷ | ㍾ | ㍽ | ㍼ | ㍻ | № | ℡ | 〶 | |
| 3x | ⚾︎ | 🉀 | 🉁 | 🉂 | 🉃 | 🉄 | 🉅 | 🉆 | 🉇 | 🉈 | 🄪 | 🈧 | 🈨 | 🈩 | 🈔 | 🈪 |
| 4x | 🈫 | 🈬 | 🈭 | 🈮 | 🈯︎ | 🈰 | 🈱 | ℓ | ㎏ | ㎐ | ㏊ | ㎞ | ㎢ | ㍱ | ||
| 5x | ½ | ↉ | ⅓ | ⅔ | ¼ | ¾ | ⅕ | ⅖ | ⅗ | ⅘ | ⅙ | ⅚ | ⅐ | ⅛ | ⅑ | ⅒ |
| 6x | ☀ | ☁ | ☂ | ⛄︎ | ☖ | ☗ | ⛉ | ⛊ | ♦ | ♥ | ♣ | ♠ | ⛋ | ⨀ | ‼ | ⁉ |
| 7x | ⛅︎ | ☔︎ | ⛆ | ☃ | ⛇ | ⚡︎ | ⛈ | ⚞ | ⚟ | ♬ | ☎ | |||||
Not in ARIB STD-B62
| ||||||||||||||||
Character set 0x7E (row number 94, list markers)
Characters from ARIB STD-B24 which were not retained in ARIB STD-B62 are shown shaded.
| ARIB STD-B24 Kanji (double-byte) set (prefixed with 0x7E)[5][11][12] | ||||||||||||||||
| 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | A | B | C | D | E | F | |
| 2x | Ⅰ | Ⅱ | Ⅲ | Ⅳ | Ⅴ | Ⅵ | Ⅶ | Ⅷ | Ⅸ | Ⅹ | Ⅺ | Ⅻ | ⑰ | ⑱ | ⑲ | |
| 3x | ⑳ | ⑴ | ⑵ | ⑶ | ⑷ | ⑸ | ⑹ | ⑺ | ⑻ | ⑼ | ⑽ | ⑾ | ⑿ | ㉑ | ㉒ | ㉓ |
| 4x | ㉔ | 🄐 | 🄑 | 🄒 | 🄓 | 🄔 | 🄕 | 🄖 | 🄗 | 🄘 | 🄙 | 🄚 | 🄛 | 🄜 | 🄝 | 🄞 |
| 5x | 🄟 | 🄠 | 🄡 | 🄢 | 🄣 | 🄤 | 🄥 | 🄦 | 🄧 | 🄨 | 🄩 | ㉕ | ㉖ | ㉗ | ㉘ | ㉙ |
| 6x | ㉚ | ① | ② | ③ | ④ | ⑤ | ⑥ | ⑦ | ⑧ | ⑨ | ⑩ | ⑪ | ⑫ | ⑬ | ⑭ | ⑮ |
| 7x | ⑯ | ❶ | ❷ | ❸ | ❹ | ❺ | ❻ | ❼ | ❽ | ❾ | ❿ | ⓫ | ⓬ | ㉛ | ||
Not in ARIB STD-B62
| ||||||||||||||||
Alphanumeric set
| ARIB STD-B24 Alphanumeric set[14] | ||||||||||||||||
| 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | A | B | C | D | E | F | |
| 2x | ! | " | # | $ | % | & | ' | ( | ) | * | + | , | - | . | / | |
| 3x | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | : | ; | < | = | > | ? |
| 4x | @ | A | B | C | D | E | F | G | H | I | J | K | L | M | N | O |
| 5x | P | Q | R | S | T | U | V | W | X | Y | Z | [ | ¥ | ] | ^ | _ |
| 6x | ` | a | b | c | d | e | f | g | h | i | j | k | l | m | n | o |
| 7x | p | q | r | s | t | u | v | w | x | y | z | { | | | } | ‾ | |
Hiragana set
| ARIB STD-B24 Hiragana set[15] | ||||||||||||||||
| 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | A | B | C | D | E | F | |
| 2x | ぁ | あ | ぃ | い | ぅ | う | ぇ | え | ぉ | お | か | が | き | ぎ | く | |
| 3x | ぐ | け | げ | こ | ご | さ | ざ | し | じ | す | ず | せ | ぜ | そ | ぞ | た |
| 4x | だ | ち | ぢ | っ | つ | づ | て | で | と | ど | な | に | ぬ | ね | の | は |
| 5x | ば | ぱ | ひ | び | ぴ | ふ | ぶ | ぷ | へ | べ | ぺ | ほ | ぼ | ぽ | ま | み |
| 6x | む | め | も | ゃ | や | ゅ | ゆ | ょ | よ | ら | り | る | れ | ろ | ゎ | わ |
| 7x | ゐ | ゑ | を | ん | ゝ | ゞ | ー | 。 | 「 | 」 | 、 | ・ | ||||
Character allocations not following row 4 of JIS X 0208
| ||||||||||||||||
Katakana set
| ARIB STD-B24 Katakana set[16] | ||||||||||||||||
| 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | A | B | C | D | E | F | |
| 2x | ァ | ア | ィ | イ | ゥ | ウ | ェ | エ | ォ | オ | カ | ガ | キ | ギ | ク | |
| 3x | グ | ケ | ゲ | コ | ゴ | サ | ザ | シ | ジ | ス | ズ | セ | ゼ | ソ | ゾ | タ |
| 4x | ダ | チ | ヂ | ッ | ツ | ヅ | テ | デ | ト | ド | ナ | ニ | ヌ | ネ | ノ | ハ |
| 5x | バ | パ | ヒ | ビ | ピ | フ | ブ | プ | ヘ | ベ | ペ | ホ | ボ | ポ | マ | ミ |
| 6x | ム | メ | モ | ャ | ヤ | ュ | ユ | ョ | ヨ | ラ | リ | ル | レ | ロ | ヮ | ワ |
| 7x | ヰ | ヱ | ヲ | ン | ヴ | ヵ | ヶ | ヽ | ヾ | ー | 。 | 「 | 」 | 、 | ・ | |
Character allocations not following row 5 of JIS X 0208
| ||||||||||||||||
Shift_JIS variant
In addition to the modified ISO 2022 encoding, the B24 standard also specifies a Shift JIS encoding following JIS X 0208:1997, but with the addition of the extended characters in the kanji set.[1]
|
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
See also
| Wikimedia Commons has media related to ARIB Extended Font. |
Footnotes
- Glossed as "temple" (i.e. Buddhist temple) in B24 table 7-10 (the list of extension characters).
- Small form (70% size per code chart / table 7-10) of a kanji character. Shown here simulated. Private Use Area code points shown are those used by the Nishiki-teki font.[13]
- Musical abbreviation (or half thereof) not present in Unicode, simulated here with multiple characters. Private Use Area code points shown are those used by the Nishiki-teki font.
References
- ARIB (2008), p. 105, part 2, section 7.3
- ARIB (2008)
- Suignard, Michel (2008-03-11). "ISO/IEC JTC1/SC2/WG2 N 3397: Japanese TV Symbols" (PDF).
- "Unicode 5.2 Emoji List". Emojipedia.
- ARIB (2014), pp. 33–50, part 2, Table 5-2
- ARIB (2008), pp. 48–52
- ARIB (2008), p. 39, part 2, Table 7-3
- "ISO-IR-087" (PDF). Information Technology Standards Commission of Japan (IPSJ/ITSCJ).
- RFC 1468 (IETF)
- ARIB (2008), p. 72
- ARIB (2008), pp. 54–72, part 2, Table 7-10
- ARIB (2008), pp. 46–47, part 2, Table 7-4
- "Nishiki-teki Version 3.82b (2021-07-23) - 6,416 characters in the Private Use Areas" (PDF).
- ARIB (2008), p. 48, part 2, Table 7-5
- ARIB (2008), p. 50, part 2, Table 7-7
- ARIB (2008), p. 49, part 2, Table 7-6
- ARIB (2008), p. 52, part 2, Table 7-9
- Data Coding and Transmission Specification for Digital Broadcasting (PDF) (ARIB Standard). 5.2-E1. 1. Association of Radio Industries and Businesses (ARIB). 2008-06-06 [1999-10-26]. ARIB STD-B24. Archived (PDF) from the original on 2017-07-10. Retrieved 2017-07-10.
- Multimedia Coding Specification for Digital Broadcasting (Second Generation) (PDF) (ARIB Standard). 1.0-E1. 1. Association of Radio Industries and Businesses (ARIB). 2014-07-31. ARIB STD-B62. Retrieved 2019-02-11.
Further reading
- Lunde, Ken Roger (December 2008). CJKV Information Processing (2 ed.). O'Reilly. ISBN 978-0-596-51447-1.
- Lunde, Ken Roger (December 1998). CJKV Information Processing (1 ed.). O'Reilly. ISBN 1-56592-224-7. (NB. Translated into Japanese and Chinese in 2002.)
_ja.svg.png.webp)