Many hyperlinks are disabled.
Use anonymous login
to enable hyperlinks.
Changes In Branch tcl9-design Excluding Merge-Ins
This is equivalent to a diff from 41f20298cd to e2fd275d2f
|
2022-03-11
| ||
| 18:23 | WIP Leaf check-in: e2fd275d2f user: dgp tags: tcl9-design | |
| 17:17 | WIP check-in: 15d2ddfe4d user: dgp tags: tcl9-design | |
| 12:10 | Merge 8.6 check-in: acf3149e65 user: jan.nijtmans tags: core-8-branch | |
|
2022-03-10
| ||
| 16:56 | merge 8.7 check-in: b0463de589 user: dgp tags: tcl9-design | |
| 12:17 | Merge 8.7 check-in: ccad489003 user: jan.nijtmans tags: trunk, main | |
| 11:43 | Add ::tcl::test::build-info command to tcl::test package, so we can find out which compiler/options ... check-in: 41f20298cd user: jan.nijtmans tags: core-8-branch | |
| 10:37 | Eliminate nmake build warning check-in: 61e186fad0 user: jan.nijtmans tags: core-8-branch | |
Added doc/dev/strings.md.
> > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 | # Design of Tcl string values for Tcl 9 ## DRAFT WORK IN PROGRESS. NOT (YET) NORMATIVE All Tcl values are strings, but what is a string? The aim of this document is to spell out what the answer to that question has been, what it should be, and how we get from one to the other. ## Fundamentals A ***string*** is a sequence of zero or more symbols, each ***symbol*** a member of a symbol set known as an ***alphabet***. For example, the Tcl string **cat** is the sequence of three symbols **c**, **a**, **t**. Each symbol in the Tcl alphabet is associated with a non-negative integer, known as the ***code*** of that symbol. Distinct symbols have distinct codes. Said another way, two symbols associated with the same code are equivalent. Schemes that associate symbols with code values are known as ***character sets***. For example, in Tcl's character set, the symbols **c**, **a**, and **t** are associated with code values 99, 97, and 116 respectively. There is an ordering imposed on the alphabet by the numeric order of the associated code values. This ordering allows symbols to be compared and sorted. The natural extension of that ordering to sequences establishes one well-defined way to compare, order and sort string values. The number of symbols in a string's symbol sequence is the ***length*** of that string. A symbol within a string can be identified by its place in the sequence, counting with an integer ***index*** that starts at 0. A string of length *N* has a symbol at each of the index values 0 through *N*-1. We can represent the symbol found at index *i* within string *s* with the notation *s*[*i*]. Given a string *s* of length *N*, every pair of an index value *i*, 0 <= *i* <= *N*, and a length *L*, 0 <= *L* <= *N* - *i*, defines a substring of *s* of length *L* built from the sequence of symbols *s*[*i*], ..., *s*[*i* + *L* - 1]. These substrings are also rooted to a particular location in the original string. Given a string *s* of length *M* and a string *t* of length *N*, the concatenation of *s* and *t* is the string *u* of length *L* = *M* + *N*, with *u*[0] = *s*[0], ... *u*[*M* - 1] = *s*[*M* - 1], *u*[*M*] = *t*[0], ... *u*[*L* - 1] = *t*[*N* - 1]. In all of these descriptions, the string values act in ways with a direct analog to C arrays of integer values. This is a familiar collection of behaviors that a broad collection of programmers can successfully reason about without extensive training in the complexities of character sets and encodings. The basics are accessible to even novice programmers. *Make the easy things easy; make the harder things possible.* To be able to create and store an arbitrary string, a Tcl interpreter has the minimum need for a set of commands or substitutions with the primitive abilities to create an empty string, to create each string of length 1 (one for each symbol in the alphabet), and the ability to concatenate arbitrary strings. Other important primitives are the ability to report the length of a string, index into a string, take a substring from a string, and compare symbols and strings. ## Tcl strings as the universal value set We most frequently think of strings as a data type with the purpose of entering, creating, storing, manipulating, processing and producing text. In Tcl, though, string values also serve as the universal value set. If a value cannot be expressed as a Tcl string, it is not a Tcl value. The representations of other value sets in Tcl is in large part an exercise in creating schemes to encode the values of those other value sets in the form of Tcl strings. Because of that, we want the set of Tcl strings to permit encodings of other value sets that are clear, simple, efficient and convenient. The history of growing Tcl's alphabet has been driven by the desire to better provide for the encoding of another value set that has not easily been accommodated by the legacy alphabet. ## Implementations The fundamentals above describe Tcl's value strings at an abstract level, but to make Tcl interpreters and Tcl libraries we have to create programs that exhibit the described abstract behviors. We will have much more detailed things to say about string representations later, but for now suffice it to say that a proper implementation of the string abstraction needs to reproduce all the fundamentals faithfully. If a representation is used that allows multiple representations of a single symbol, or multiple representations of a single string, this can be a matter of some difficulty. Representations that include the possibility of states that are not valid string values at all are also cases that need careful consideration. ## Tcl alphabet versions This section looks into Tcl's history, which can be tricky. History gets messy. It is full of steps and mis-steps and it can be difficult to get agreement looking back about which were which. In Tcl 7, the alphabet for Tcl strings was a set of symbols associated with code values 1 through 255. This is exactly the set of **NUL**-terminated C strings. The symbols with code values 1 through 127 are defined to follow the ASCII character set. This is important to the definition of Tcl because all symbols with syntactic meaning in Tcl scripts are in the ASCII set. The symbols with code values 128 through 255 were less stringently specified, but such symbols could nevertheless be reliably created, stored, processed and produced by Tcl programs. This set of string values continues to have relevance in Tcl today, because it is exactly this set that can be passed as arguments to Tcl commands defined via **Tcl_CreateCommand**. Even though no Tcl 7 value could contain a symbol with code 0, Tcl still offered a substitution suggesting it was possible. ``` % string length <\x01> 3 % string length <\x00> 1 ``` This seems to be a mis-step, where the \\x00 substitution should have raised an error. The fact that everything created by it broke expectations in some way supports that judgment. But it's not beyond imagination for someone to take another view. The value set of arbitrary binary data, in the form of byte sequences, is not well served by the Tcl 7 alphabet. No encoding into Tcl 7 strings can be both simple and efficient. Either it must be variable-width, or it must use a fixed width of at least two symbols per byte. There are certainly ways to encode arbitrary binary data using only the Tcl 7 alphabet, but the commands of Tcl 7 never chose one for the core commands of the language to use. The problem was left unsolved so that a [**read**] from a binary channel returned a value that the rest of Tcl simply truncated at the first **NUL**. In Tcl 8.0, the Tcl alphabet added a symbol with code value 0. This allowed the direct encoding of a byte sequences of length *N* by a string of length *N* with each symbol determined by the code given by the byte value. Aribtrary binary data could be stored in Tcl variables, and processed by any commands created by the new **Tcl_CreateObjCommand**. Legacy commands still created by **Tcl_CreateCommand** remained "binary unsafe". Note that the alphabet strictly grew between Tcl 7 and Tcl 8.0. All string values representable in Tcl 7 remained representable in Tcl 8.0. Internally there was reform in representation (counted strings replaced terminated strings), but the concept of the Tcl string value accessible to scripts changed along an upward compatible path. Tcl 8.0 string values suffered from two deficits. First, The internals still used two representations that were not reconciled to provide the same functionality. Second, the international character sets of increasing importance could not be encoded into Tcl string values in ways that were both simple and efficient. International character set text became the next value set prompting an expansion of Tcl's alphabet. In Tcl 8.1 (released April, 1999), the Tcl alphabet expanded to a set of symbols with code values ranging from 0 to 65,535 (0x0000 to 0xFFFF). Where a code value had an assigned symbol in Unicode 2.0, the symbol of Tcl's alphabet agreed with the Unicode character set. This implied continued symbol agreement with ASCII. All assigned codepoints of Unicode 2.0 fit in this alphabet. Tcl could store all Unicode text values in a simple and efficient way in its string values. It could perhaps be said that Tcl 8.1 strings *were* Unicode 2.0 text values. The Tcl documentation is certainly phrased in those terms. In hindsight, given the later divergence in the development of Tcl and Unicode, that's not the most useful perspective. It is more useful to think of the set of Tcl 8.1 strings as a superset of the set of Unicode 2.0 text values. This set of string values came to be known as UCS-2 to distinguish it from later versions of Unicode. Tcl 8.1 added a new backslash substitution syntax, **\\u**___HHHH___, capable of producing every symbol in the Tcl alphabet. Built-in Tcl commands such as **format** and **scan** were extended to produce and accept symbols and codes in the extended alphabet. Once again the alphabet strictly grew, preserving upward compatibility in the set of string values available to Tcl programs. The representations of strings inside the Tcl 8.1 library were substantially reformed to support the extended alphabet. These representations appeared in parts of the C programming interface of the Tcl library, so extensions written for Tcl 7 or Tcl 8 had to be adapted to the reforms. Tcl 8.1 string values did bring with them a new mis-step. A new command **encoding** provided scripts with the ability to transform Unicode text values into a variety of encoded forms. This **encoding** command included support for an encoding called **identity** that empowers scripts to store an arbitrary byte sequence inside an internal representation for Tcl strings. The consequence is that scripts can create multiple representations for the same string that are not consistently treated as the same string, contrary to our fundamental expectations. ``` % info patch 8.1.1 % set s \u0080 % set t [encoding convertfrom identity \x80] ?? % string length $s 1 % string length $t 1 % scan [string index $s 0] %c code; set code 128 % scan [string index $t 0] %c code; set code 128 % string equal $s $t 0 ``` While scripts can avoid use of the **identity** encoding (perhaps treating any use of it as introducing *undefined* behavior), the underlying implementation flaws that allow for its mischief are available to all Tcl extensions, so the failure of Tcl 8.1 strings to fully conform to fundamental expectations still lurks. More on this when we examine representations below. Tcl 8.2 (released August, 1999) kept the same string alphabet, but revised all relevant commands to process symbols according to the Unicode 2.1 standard. Notably this included the assignment of code point **U+20AC** to the symbol EURO SIGN. The Tcl 8.3.* series of releases (February 2000 - October 2002) continued that same standard of Unicode support. From Chapter 1 of *The Unicode Standard, Version 2.0*, "**The Unicode Standard is a fixed-width, uniform encoding scheme for written characters and text.**" Chapter 2 presents a set of desgin principles. The first principle states, in part, "**Plain Unicode text consists of pure 16-bit Unicode character sequences.**" The second principle states "**The full 16-bit codespace... is available to represent characters.**" A plain reading of those designs and assurances suggests that sequences of symbols from a 16-bit alphabet will be a suitable fixed-width representation for Unicode text. The second design principle, though, goes on to describe a "*surrogate extension mechanism*" which uses a pair of 16-bit code values to represent a single character. This mechanism, of course, makes the specified Unicode definition a variable-width encoding of abstract characters, not a fixed-width encoding at all. That said, Unicode 2.0, does specify surrogate pairs as the way to represent a full collection of 1,114,112 distinguishable symbols in the potential Unicode collection. It also reserves the surrogate code points to be used properly in Unicode text only in the formation of such pairs. It did not assign any characters to those pairs, instead describing the mechanism as one "for encoding extremely rare characters". Tcl string values processed symbols corresponding to Unicode surrogates no differently from any other symbols in the Tcl alphabet. In this state of things, it is clear that the set of Tcl strings is strictly a superset of the set of well-formed Unicode text. This is not only because of unassigned codepoints awaiting their symbols, but includes an ability to store symbol sequences in Tcl string values that can never become well-formed Unicode text at any point in the future. Unicode 3.1.0 was released March 2001. This was the first version of the Unicode Standard that assigned characters to surrogate code pairs in the 16-bit encoding, which in Unicode 3 came to have the name UTF-16. In Unicode 3 it was acknowledged that UTF-16 is a variable-length encoding. Tcl's source code was updated to support this specification of Unicode in May 2001, but support for surrogate pairs to represent Unicode characters with codes greater than **U+FFFF** was not implemented. This was when Tcl Unicode support broke away from conformance to the Unicode Standard. Many programming assumptions rooted in the existence of a fixed-width, 16-bit encoding for every Tcl string value had become deeply embedded in both Tcl's implementation and in some of its interfaces. The assigned Unicode characters to be gained outside the Basic Multilingual Place remained those of "extremely rare" interest. This level of partial Unicode 3.1.0 support was first released in Tcl 8.4.0 in September 2002. And then Tcl's Unicode support fell into a deep sleep. While Tcl's support of Unicode slept, Unicode itself evolved and had the language of its conformance standards tightened and refined. The conception of Unicode defined as seqeunces of 16-bit code units faded away, and the 16-bit representation, UTF-16, became just one of several encodings to be used. Fundamentally each distinguishable Unicode text that exists *or that ever will exist under future revisions of Unicode* became defined as a sequence of zero or more symbols from the alphabet of ***unicode scalar values***. A unicode scalar value is associated in the Unicode character set with a code value less than 1,114,112 and also constrained to exclude the code values associated with the surrogate extension mechanism. The code values of unicode scalar values are in the integer ranges 0 to 55,295 (0x0000 to 0xD7FF) and 57,344 to 1,114,111 (0xE000 to 0x10FFFF). The code values of unicode scalar values are all representable as 21-bit integers. The 21-bit integers which are not the code value of any unicode scalar value are the ranges 55,296 to 57,343 (0xD800 to 0xDFFF) and 1,114,112 to 2,097,151 (0x110000 to 0x1FFFFF). (Note that 21-bit code values are exactly what can be encoded by the lead and trail byte scheme of UTF-8 restricted to 4-byte sequences. Note also that code value 0x10FFFF is the largest value that can be decoded from a surrogate pair.) From this perspective, the Tcl 8.1 string values were no longer seen as proper Unicode, but as a legacy system called UCS-2 which suffered from two flaws. It lacked support for the supplemntary planes of the full Unicode character set beyond **U+FFFF**. It also allowed for the presence of symbols with codes between 55,296 and 57,343 (0xD800 to 0xDFFF). Neither system could encode the other in direct, simple, efficient representations. Unicode also came to define a collection of encodings. The simplest to understand is UTF-32, where each unicode scalar value in a Unicode sequence is directly represented by a 4-octet (32-bit) value matching the code. UTF-32 offers simplicity and fixed-width representation for efficient random access indexing. It also offers considerable waste of storage and transmission. Properly defined, UTF-32 does not include representations of symbols that are not unicode scalar values. However, the extension to include them is not difficult to imagine. The name UCS-4 comes down from the ISO 10646 effort, and has evolved to mean the same thing as UTF-32. The UTF-16 encoding uses 2-octet (16-bit) code units in variable length patterns to represent each unicode scalar value. The Unicode scalar values from ranges 0x0000 - 0xD7FF and 0xE000 - 0xFFFF are represented by themselves in a single code unit. The Unicode scalar values from the range 0x10000 - 0x10FFFF are each represented by a pair of 16-bit code units, the first from the range 0xD800 - 0xDBFF and the second from the range 0xDC00 - 0xDFFF. Every 2-octet code unit may appear somewhere in the proper UTF-16 encoding of some Unicode text. There are no code unit values that are forbidden. There are UCS-2 sequences that are not valid UTF-16, precisely those UCS-2 sequences that include surrogates not arranged in properly formed pairs. Note the difficulties if we try to design an encoding with 2-octet code units to encode all sequences over the union of the UCS-2 and Unicode alphabets. This can certainly be done, but the result will bear little resemblence to UTF-16. All valid UTF-16 sequences are used up representing valid Unicode. To also represent the strings of UCS-2 that are not valid Unicode, we would need to use 2-octet sequences that are not valid UTF-16. Any such scheme will have some point of discontinuity with the legacy UCS-2 system. Likewise, since every 2-octet code unit sequence is used in UCS-2 to represent itself, there is no room to create representation for supplemental planes of unicode in a 2-octet encoding without introducing a discontinuity. There will have to be some 2-octet sequence that used to mean one thing and now means another thing at the transition, or there will have to be some approach of preserving both systems and taking care to distinguish at each point which is in use. This sticking point is a major difficulty in managing a migration in Tcl's Unicode support, since public interfaces exist that transfer 2-octet encoded data. Unicode also defines the UTF-8 encoding of unicode scalar values into single-octet (8-bit) (byte-oriented) code units in variable length patterns. The Unicode scalar values from the range 0x0000 - 0x007F are represented by themselves in a single code unit. Each value from the range 0x0080 - 0x07FF is represented by a two-byte sequence of a leading byte from the range 0xC2 - 0xDF and a trailing byte from the range 0x80 - 0xBF. Each value from the ranges 0x0800 - 0xD7FF and 0xE000 - 0xFFFF is represented by a three-byte sequence starting with a leading byte from the range 0xE0 - 0xEF and two trailing bytes as before. (Some such three-byte sequences would encode surrogates, and are therefore not valid UTF-8 sequences). Each value from the range 0x10000 - 0x10FFFF is represented by a four-byte sequence starting with a leading byte from the range 0xF0 - 0xF4 and three trailing bytes as before. (Some such four-byte sequences would decode to a codepoint greater than 0x10FFF, and are therefore not valid UTF-8 sequences). Note the implication that the bytes 0xC0, 0xC1, and 0xF5 - 0xFF can never appear in a proper UTF-8 byte sequence. Besides those forbidden bytes, there are many sequences also forbidden, including any trailing byte where one does not belong, any leading byte where one does not belong, or any multi-byte sequence that when decoded would produce a value outside the domain for that sequence length (for example, the four-byte sequence 0xF0 0x80 0x80 0x80 that would appear to encode **U+0000**, which is in the domain of unicode scalar values properly encoded by a single-byte). The last of these constraints was formally imposed by Unicode 3.1.0. Earlier versions of Unicode explictly approved of UTF-8 decoders that accepted overlong byte sequences, and even included such decoders in their sample implementations. Here the news is better when it comes to thinking about representations of all sequences over the union of UCS-2 and Unicode alphabets. The three-byte sequences that might encode surrogates that are forbidden in UTF-8 are available to encode those values of UCS-2 without interference with UTF-8 encoding of everything else. The WTF-8 variation of UTF-8 is one approach in this area, though the details require careful examination. The byte sequence 0xF0 0x90 0x80 0x80 (representing **U+10000**) and the byte sequence 0xED 0xA0 0x80 0xED 0xB0 0x80 (representing **U+D8000 U+DC00**, which in turn is the surrogate pair representation of **U+10000**) are distinct, allowing the required distinction. WTF-8 is defined to disallow this, however, because of the lack of a continuing ability to distinguish after conversion to UTF-16. That remains the critical sticking point. Tcl 8.1 documentation began the claim that Tcl used UTF-8 to store its string values and as the encoding with which to pass byte-oriented string values through interfaces, but that was never strictly true. More on that below. Unicode 6.0.0 (October 2010) was notable as the first version to include assignments for emoji symbols. Their popularity on mobile devices would end the days when interest in Unicode characters above **U+FFFF** could be said to be rare. Also in October 2010 Tcl was awakened out of its Unicode support slumber, still offering only partial Unicode 3.1.0 support. At that point, Tcl was brought up to date with Unicode 6.0.0, but Tcl remained without surrogate pair support. Since that implied Tcl was also without emoji symbol support, this increasingly became a mark against Tcl as a langauge to use for good Unicode programming. The partial support of Unicode 6 was first released in Tcl 8.5.10 in June 2011. Since then it has been customary to update Tcl's partial support for the new versions of the Unicode Standard as they are released. Tcl 8.5.19 was released February 2016 with partial support of Unicode 8.0. Tcl 8.6.0 was released December 2012 with partial support of Unicode 6.2. Tcl 8.6.6 was released July 2016 with partial support of Unicode 9.0. In 2017, the trunk of Tcl development was turned over to work on the Tcl 9.0 release. This focused attention again on how best to revise the string representation for that new milestone. Following the example of earlier alphabet expansions, it appears clear that for improved Unicode support, we need the Tcl 9 alphabet to include the set of all unicode scalar values. Also following prior examples, it seems desirable to strictly grow the alphabet so that all strings representable in Tcl 8 remain representable in Tcl 9. Achieving both would mean a Tcl 9 alphabet that is a (super?)set of the union of UCS-2 and Unicode scalar values. The history of Tcl string values in releases 8.6.7 and later felt the influence of that focus, and are best discussed after some attention to Tcl's string representations. ## Representations Tcl 7 strings are represented directly as C strings. The representation is one-to-one and complete. Every Tcl 7 string has a representation by exactly one C string, and every C string represents exactly one proper Tcl 7 string. There is no C string that can be rejected as not representing a Tcl 7 string. This is very simple. Also, any interface using the passing of C string values to implement the conceptual passing of Tcl 7 string values can be designed with this knowledge. There is no need to provide for error handling when the possibility of error is defined out of existence. Because of this, many of Tcl's interface routines that date back the longest offer no capability to report errors in their string arguments. The C string representation is also a fixed width encoding of Tcl 7 strings. This allows for efficient indexing. Tcl 8.0 strings are represented directly as the pair of a byte array and a length stored as a C **signed int**. No Tcl 8.0 string with length greater than **INT\_MAX** can be accommodated by this representation. Other than that limitation, the same direct, complete, and one-to-one nature of representation is present as for Tcl 7 strings. A string too long is the only error condition that needs to be considered, and that error can be prevented by constraining size of arguments. Again, many interfaces offer no detection or handling mechanisms to deal with an errors or invalidity in string value arguments. The Tcl 8.0 representations also continue to offer efficient indexing via fixed-width storage. Tcl 8.0 strings are a superset of Tcl 7 strings. When a Tcl 7 string value is represented in the Tcl 8.0 manner, the byte array has the same contents as the C string representation from Tcl 7. The only difference is whether there must be a terminating **NUL** byte. In many places in the Tcl 8.0 representation, such a terminating **NUL** byte was used to allow easier interoperability with legacy routines written to Tcl 7 expectations. Most notably the *bytes* and *length* fields of a **Tcl\_Obj** struct implement the Tcl 8.0 string representation and require the terminating **NULL** to be present at *bytes*[*length*]. The expansion of the Tcl alphabet in Tcl 8.1 to all two-byte codes brought about new string value representations in the Tcl library. The simplest was the creation of the **Tcl\_UniChar** type, a two-byte integer able to contain a single symbol from the Tcl alphabet. This representation of a Tcl symbol is used in the new Tcl 8.1 routines **Tcl\_UniCharToUpper**, **Tcl\_UniCharToLower**, **Tcl\_UniCharToTitle**, **Tcl\_UtfToUniChar**, **Tcl\_UniCharToUtf** and **Tcl\_UniCharAtIndex**. In each of these routines a symbol is passed in as a value of type **int** and it is passed out as a value of type **Tcl\_UniChar**. A Tcl string is then trivially represented as an array of **Tcl\_UniChar**. It has the familiar properties of being a simple, complete, one-to-one, fixed-width representation of all abstract Tcl 8.1 strings, with all the familiar comforts and benefits. This representation is used in the new Tcl 8.1 routines **Tcl\_UniCharLen**, **Tcl\_UniCharNcmp**, **Tcl\_UniCharToUtfDString**, and **Tcl\_UtfToUniCharDString**. The first two are utility routines that can be applied to this representation. The second two are conversion routines between the **Tcl\_UniChar** array representation and another representation still to be described. None of the public routines in Tcl 8.1 passed an argument into or out of Tcl as a Tcl value in the **Tcl\_UniChar** array representation. The Tcl 8.1 interface continues to pass Tcl values in and out by use of either the C strings of Tcl 7, or the counted strings of Tcl 8.0 (often encapsulated in a **Tcl\_Obj** struct). These interfaces could only represent the extended alphabet of Tcl 8.1 by way of a reinterpretation of how the string values are encoded in the byte sequences. The routine **Tcl\_UniCharToUtf** generates these encoded byte sequences. Tcl 8.1 documentation claims that the byte sequence encoding is UTF-8. This was never strictly true. Examination of the body of **Tcl\_UniCharToUtf** reveals that it encodes something closer to the FSS-UTF encoding from Unicode 1.1, with encoding into sequences of more than 3 bytes disabled. The limit on byte sequence length is available from the public Tcl header as **TCL\_UTF\_MAX**, with the suggestion that configuration to limits other than 3 might be available. Tcl 8.1 did not offer any such working customizations. At best, any variants triggered by a value of **TCL\_UTF\_MAX** other than 3 in the Tcl 8.1 source code could be viewed as speculations about what might be offered in later Tcl releases. FSS-UTF calls for 3-byte sequences to encode all codepoints in the range U+0800 to U+FFFF, and calls for 4-byte sequences to encode all codepoints in the range U+10000 to U+1FFFFF, and longer byte sequences to encode codepoints up to U+7FFFFFFF, providing room for up to 31-bits of codepoints that ISO 10646 still proposed at the time. The UTF-8 encoding defined in Unicode 2.0 proposed a different use for 4-byte sequences. They would be used to encode codepoint pairs using the surrogate extension mechanism. The refusal of the Tcl 8.1 encoder to produce 4-byte sequences marks another way in which Tcl's encoding was never UTF-8. The UTF-8 spec in Unicode 2.0 was silent about the encoding of unpaired surrogate codepoints, but the sample implementation included as an example did encode such codepoints into 3-byte sequences. Later revisions to the UTF-8 spec eliminated the encoding of unpaired surrogates. The encoding implemented by **Tcl\_UniCharToUtf** encodes all surrogate codepoints, paired or not, into 3-byte sequences. Tcl 7 strings cannot contain the byte 0x00, while UTF-8 (and FSS-UTF) specifies that **U+0000** is to be encoded by 0x00. The early UTF-8 specs allowed UTF-8 decoders to decode the two-byte sequence 0xC0 0x80 into **U+0000** (decoding of all overlong encodings was officially tolerated at the time). so **Tcl\_UniCharToUtf** is implemented to use that modified encoding of **U+0000**. Any Tcl 8.1 string encoded in the way described above can be passed out of Tcl in the encoded form. Callers of Tcl have to adjust their expectations to treat the value properly. Earlier versions of Tcl would pass out arbitrary byte sequences representing strings with one byte per symbol. The new representation has a variable number of bytes per symbol. Each Tcl caller had to review its functionality and purpose to decide what adaptations were needed. Tcl 8.1 provided a large number of utility routines to process byte sequences in the variable width encoding. Tcl 8.1 also provided routines to support processing of arbitrary byte array values for those Tcl callers that really needed to process arbitrary byte arrays exchanged as Tcl values. Callers of Tcl also pass Tcl string values into the library as either C strings or counted strings. These callers likewise had to adjust to the new encoding expectations. When they pass in byte sequences in agreement with the encoding, they can expect to gain access to the whole extended string set of Tcl 8.1. However, the encoding does not produce all byte sequences. Many byte sequences do not correspond to the proper encoding of anything. The Tcl 8.1 rules for decoding are implemented in the routine **Tcl\_UtfToUniChar**. All byte sequences generated by the encoder are reliably decoded back to the original symbol. Round trips through the encoding are lossless. In addition, following recommendations of the time, the decoder accepted all overlong sequences otherwise structured with appropriate sequences of lead bytes and trail bytes. The decoder did not decode any byte sequences of length greater than three. Any byte leading a sequence not recognized by those rules would be decoded as itself, effectively making use of ISO-8859-1 as a fallback encoding. All of these conventions of the decoder provided opportunities for very different byte sequences to be seen as representations of the same string after decoding, with all of the risks of violations of fundmental string behavior that go with that. The UTF-8 encoding as strictly specified benefits from the property that sorts on the encoded form are the same as sorts on the original sequence, but Tcl's modified approach loses that benefit. This possibility for multiple accepted representations of the same string is the root of most failures of string fundamentals in Tcl 8.1. Why accept these improper byte sequences at all? One key factor is the large existing set of routines that accept string arguments with no mechanism for error handling, because in earlier Tcl releases no errors were possible. All byte sequences were proper strings until Tcl 8.1. That said, it is less clear why a fallback to ISO-8859-1 was considered wiser than something like generation of U+FFFD, the REPLACEMENT character. One factor might be the ability to pass 4-byte UTF-8 through without conversion or loss, at the expense of it being seen as 4 symbols instead of 2 (UTF-16) code units or one symbol. All that said, roundtrips from the byte-encoded format to **Tcl\_UniChar** array and back are not lossless. This might be another of those arguable mis-steps seen more clearly in hindsight. It is the responsibility of each caller of **Tcl\_UtfToUniChar** that the byte sequence passed in is either **NUL** terminated, or has sufficient length that decoding governed by lead byte values will not read beyond valid memory limits. The utility routine > **int** **Tcl\_UtfCharComplete**(**char** *_src_, **int** _len_) is provided as a tool to check for this condition. It is known by all callers that this routine will return 1 (true) whenever (_len_ > **TCL\_UTF\_MAX**), so many callers omit calls in that circumstance. Callers of **Tcl\_UniCharToUtf** are expected to provide an output buffer of at least **TCL\_UTF\_MAX** bytes, and to expect a return count up to (but never more than) **TCL\_UTF\_MAX**. The encoder and decoder routines allow translation back and forth between the two Tcl 8.1 representations for strings, the variable-width byte encoding falsely described as UTF-8, and the fixed-width encoding as elements of a **Tcl\_UniChar** array. The translation is one symbol at a time, either reading or writing one **Tcl\_UniChar** element with each call. This nails down the conception of the array representation as UCS-2 strings, with no decoding of multi-element sequences. The documentation is clear about this when it declares that *n* calls to **Tcl\_UtfNext** must produce the same result as one call to **Tcl\_UtfAtIndex**(*n*). Indexing and iteration must agree. For extensions and apps, the migration from Tcl 8.0 to Tcl 8.1 was pretty abrupt. The new primary string representation was a variable-width encoding of an enlarged alphabet, requiring the use of new utility routines to achieve basic tasks like iterating or indexing, in many cases with a substantial performance penalty. After conversion to the new expectations, it was rare that the revised code could still be used with Tcl 8.0. A non-trivial conversion with little support for retreat or delay, and a performance hit as a reward. There was plenty of unhappiness expressed about the change. Tcl 8.2 was released just 4 months later, in August 1999. The string representations were unchanged, but much more caching of the **Tcl\_UniChar** array representation was added internally to better amortize the performance penalties of some operations on sequences stored in a variable-width encoding. Different strings were also represented differently internally based on their content to offer performance and efficiency benefits. The result was good enough to make the Tcl migration to international strings a success, if a bumpy one. New interfaces were also supplied by Tcl 8.2 to allow Tcl callers to exchange UCS-2 strings stored as **Tcl\_UniChar** arrays with the Tcl library: **Tcl\_NewUnicodeObj**, **Tcl\_SetUnicodeObj**, **Tcl\_GetUnicode**, **Tcl\_GetUniChar**, and **Tcl\_AppendUnicodeToObj**. Here it is unmistakable that **Tcl\_GetUnicode** was a mis-step, re-opening the "**NUL** as terminator" problem on the new alphabet. The undefined nature of **Tcl\_GetUniChar** for an index out of range was unwise. And overall the use of "Unicode" in the routine names that act on what came to be understood as UCS-2 strings brings about a festival of confusion. Tcl 8.3 (released February 2000) added no new interfaces making use of **Tcl\_UniChar**, and string representations remained unchanged. Tcl 8.4 (released September 2002) has already been noted not for the changes it made, but for the changes in Unicode it declined to acknowledge. Two new utility routines **Tcl\_UniCharNcasecmp** and **Tcl\_UniCharCaseMatch** accepted **Tcl\_UniChar** array arguments. The routine **Tcl\_GetUnicodeFromObj** was also added as the properly defined replacement for **Tcl\_GetUnicode**. Otherwise, the universe of Tcl strings and their representations remained unchanged. Tcl release 8.4.4 (July, 2003) made the first stab at supporting a value for **TCL\_UTF\_MAX** greater than 3. Comments suggested another possible value of 6, with a conditional typedef for a 4-byte **Tcl\_UniChar**. However, Tcl continued to use the FSS-UTF approach to encoding longer byte sequences in such custom builds. In Tcl release 8.4.14 (October 2006), comments were revised to declare for the first time a plan to make use of UTF-16. Nevertheless the encoder remained one patterned after FSS-UTF. The same state of affairs persisted through all the remaining releases of Tcl 8.4. It also remained effectively the same in all releases from Tcl 8.5.0 (December 2007) through Tcl 8.5.19 (February 2016). (A comment was added in *tclStringObj.c* during a commit otherwise largely devoted to formatting and whitespace issues. The new comment also made suggestions about claiming use of UTF-16 without making any changes in the code that would make it in any way conformant to UTF-16.) It also remained the same in releases Tcl 8.6.0 (December 2012) through Tcl 8.6.4 (March 2015). In all these releases, Tcl's encoder into byte sequences continued to be implemented on the old FSS-UTF model. Setting **TCL\_UTF\_MAX** to 6 might provide representations for astral characters, but nothing would encode them into or decode them from their UTF-16 encoding. Tcl 8.6.5 (February 2016) for the first time broke away from FSS-UTF in its encoder in custom builds with values of **TCL\_UTF\_MAX** greater than 3. Five- and six-byte sequences were no longer generated, though the decoder continued to decode them. The upper limit of U+10FFFF was imposed in the encoder with larger values replaced by U+FFFD. Several of the classification functions were extended in custom builds to apply to astral characters. The other novelty in this release was the carving out of the **TCL\_UTF\_MAX == 4** custom builds to offer astral character support, but not through use of a 32-bit **Tcl\_UniChar**. For the first time, interpretation of surrogate pairs appears in some parts of some custom builds. Also for the first time, the callers of **Tcl_UtfToUniChar** were required in some circumstances to manage a dance of two calls to decode a single byte sequence, because some byte sequences were for the first time represented by a sequence of two **Tcl\_UniChar** values. This was the first step toward use of UTF-16 in the Tcl library. Like most first attempts, it didn't get everything right out of the gate, but improvements would continue to come. Because these changes were only visible in custom builds, there was little controversy about their appearance in a patch release. Tcl 8.6.6 (July 2016) remained the same. Tcl 8.6.7 (August 2017) was the first release since Tcl 8.1.0 to change the decoding of string value byte sequences in the default build. This release brought an end to Tcl's acceptance of overlong byte sequences (other than the modified encoding for U+0000). This change fixes many EIAS violations created by decoding multiple encoded sequences to mean the same symbols. In custom builds, Tcl 8.6.7 also ended the decoding of the five- and six-byte sequences of FSS-UTF. Routines where TCL\_UTF\_MAX is relevant: Tcl\_UtfBackslash, Tcl\_UtfPrev, Tcl\_WinTCharToUtf All Utf.3 (sort of) [gets] and [read] and "unicode" string rep buffer sizing Encoding routines (Tcl\_ExternalToUtf,...) |