u8_strcmp(3c) 맨 페이지 - 윈디하나의 솔라나라

개요

섹션
맨 페이지 이름
검색(S)

u8_strcmp(3c)

u8_strcmp(3C)            Standard C Library Functions            u8_strcmp(3C)

NAME
       u8_strcmp - UTF-8 string comparison function

SYNOPSIS
       #include <sys/u8_textprep.h>

       int u8_strcmp(const char *s1, const char *s2, size_t n,
            int flag, size_t version, int *errnum);

PARAMETERS
       s1, s2

           Pointers to null-terminated UTF-8 strings


       n

           The maximum number of bytes to be compared. If 0, the comparison is
           performed  until  either or both of the strings are examined to the
           string terminating null byte.


       flag

           The possible comparison options constructed  by  a  bit-wise-inclu‐
           sive-OR of the following values:


           U8_STRCMP_CS          Perform   case-sensitive  string  comparison.
                                 This is the default.


           U8_STRCMP_CI_UPPER    Perform  case-insensitive  string  comparison
                                 based  on Unicode uppercase converted results
                                 of s1 and s2.


           U8_STRCMP_CI_LOWER    Perform  case-insensitive  string  comparison
                                 based  on Unicode lowercase converted results
                                 of s1 and s2.


           U8_STRCMP_NFD         Perform string comparison  after  s1  and  s2
                                 have been normalized by using Unicode Normal‐
                                 ization Form D.


           U8_STRCMP_NFC         Perform  string  comparison  after  s1 and s2
                                 have been normalized by using Unicode Normal‐
                                 ization Form C.


           U8_STRCMP_NFKD        Perform string comparison  after  s1  and  s2
                                 have been normalized by using Unicode Normal‐
                                 ization Form KD.


           U8_STRCMP_NFKC        Perform  string  comparison  after  s1 and s2
                                 have been normalized by using Unicode Normal‐
                                 ization Form KC.

           Only one case-sensitive or case-insensitive option is allowed. Only
           one Unicode Normalization option is allowed.


       version

           The version of Unicode data that should be used during  comparison.
           The following values are supported:

           U8_UNICODE_320          Use Unicode 3.2.0 data during comparison.


           U8_UNICODE_500          Use Unicode 5.0.0 data during comparison.


           U8_UNICODE_1400_ORCL    Use  Unicode 14.0.0 data during comparison.
                                   (See NOTE below.)


           U8_UNICODE_LATEST       Use the latest Unicode version data  avail‐
                                   able, which is currently Unicode 14.0.0.



       errnum

           A  non-zero  value indicates that an error has occurred during com‐
           parison. The following values are supported:

           EBADF     The specified option values are conflicting and cannot be
                     supported.


           EILSEQ    There was an illegal character at s1, s2, or both.


           EINVAL    There was an incomplete character at s1, s2, or both.


           ERANGE    The specified Unicode version value is not supported.



DESCRIPTION
       The u8_strcmp() function internally processes UTF-8 strings pointed  to
       by s1 and s2 based on the corresponding version of the Unicode Standard
       and  other  input arguments and compares the result strings in byte-by-
       byte, machine ordering.


       When multiple comparison options are specified,  Unicode  Normalization
       is  performed  after  case-sensitive  or case-insensitive processing is
       performed.

NOTE
       U8_UNICODE_1400_ORCL uses a slightly modified version  of  the  Unicode
       14.0.0  tables. Where Unicode 14.0.0 says that the uppercase equivalent
       of U+0131 LATIN SMALL LETTER DOTLESS I is U+0049 LATIN  CAPITAL  LETTER
       I,  this implementation does not; it leaves U+0131 without an uppercase
       equivalent. This change helps to reduce conflicts between  English  and
       Turkish uses of dotted and dotless I.

RETURN VALUES
       The  u8_strcmp() function returns an integer greater than, equal to, or
       less than 0 if the string pointed to by s1 is greater than,  equal  to,
       or less than the string pointed to by s2, respectively.


       When u8_strcmp() detects an illegal or incomplete character, such char‐
       acter  causes  the function to set errnum to indicate the error. After‐
       ward, the comparison is still performed on the resultant strings and  a
       value based on byte-by-byte comparison is always returned.

EXAMPLES
       Example 1 Perform simple default string comparison


         #include <sys/u8_textprep.h>

         int
         docmp_default(const char *u1, const char *u2) {
             int result;
             int errnum;

             result = u8_strcmp(u1, u2, 0, 0, U8_UNICODE_LATEST, &errnum);
             if (errnum == EILSEQ)
                 return (-1);
             if (errnum == EINVAL)
                 return (-2);
             if (errnum == EBADF)
                 return (-3);
             if (errnum == ERANGE)
                 return (-4);
         }


       Example 2 Perform case-insensitive comparison



       Perform  uppercase based case-insensitive comparison with Unicode 3.2.0
       data.


         #include <sys/u8_textprep.h>

         int
         docmp_caseinsensitive_u320(const char *u1, const char *u2) {
             int result;
             int errnum;

             result = u8_strcmp(u1, u2, 0, U8_STRCMP_CI_UPPER,
                 U8_UNICODE_320, &errnum);
             if (errnum == EILSEQ)
                 return (-1);
             if (errnum == EINVAL)
                 return (-2);
             if (errnum == EBADF)
                 return (-3);
             if (errnum == ERANGE)
                 return (-4);

             return (result);
         }


       Example 3 Perform Unicode Normalization Form D



       Perform Unicode Normalization Form D and uppercase based  case-insensi‐
       tive comparison with Unicode 3.2.0 data.


         #include <sys/u8_textprep.h>

         int
         docmp_nfd_caseinsensitive_u320(const char *u1, const char *u2) {
             int result;
             int errnum;

             result = u8_strcmp(u1, u2, 0,
                 (U8_STRCMP_NFD|U8_STRCMP_CI_UPPER), U8_UNICODE_320,
                 &errnum);
             if (errnum == EILSEQ)
                 return (-1);
             if (errnum == EINVAL)
                 return (-2);
             if (errnum == EBADF)
                 return (-3);
             if (errnum == ERANGE)
                 return (-4);

             return (result);
         }


ATTRIBUTES
       See attributes(7) for descriptions of the following attributes:

       tab()  box; cw(2.75i) |cw(2.75i) lw(2.75i) |lw(2.75i) ATTRIBUTE TYPEAT‐
       TRIBUTE VALUE _ Interface StabilityCommitted _ MT-LevelMT-Safe


SEE ALSO
       u8_textprep_str(3C),  u8_validate(3C),  attributes(7),   u8_strcmp(9F),
       u8_textprep_str(9F), u8_validate(9F)


       Converting  Codesets  in Internationalizing and Localizing Applications
       in Oracle Solaris


       The Unicode Standard (https://www.unicode.org/standard/standard.html)

HISTORY
       The u8_strcmp() function was introduced in Oracle Solaris 11.0.0.

Oracle Solaris 11.4               25 Nov 2024                    u8_strcmp(3C)
맨 페이지 내용의 저작권은 맨 페이지 작성자에게 있습니다.
RSS ATOM XHTML 5 CSS3