A language processing method, device, apparatus and computer readable storage medium

By determining the basic area encoding combination relationship of agglutinative text, the target character is automatically selected and displayed, which solves the problem of inaccurate font display in agglutinative text and improves the readability of the text.

CN113705162BActive Publication Date: 2026-05-01TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2021-03-04
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing OCR and input methods cannot automatically select character types for agglutinative languages, resulting in insufficient accuracy in displaying agglutinative text.

Method used

By obtaining at least two basic area codes of the target text, determining the combination relationship between two adjacent basic area codes, and determining the target character code from the associated character codes based on the combination relationship, automatic character selection and correct display are achieved.

Benefits of technology

It improves the accuracy of font display in agglutinative text and enhances text readability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113705162B_ABST
    Figure CN113705162B_ABST
Patent Text Reader

Abstract

The application discloses a language processing method, device and equipment and a computer readable storage medium. The method comprises the following steps: obtaining at least two basic area encodings of a target text, wherein the basic area encodings are associated with at least two font encodings; determining a combination relationship of two adjacent basic area encodings in the at least two basic area encodings, wherein the combination relationship represents whether the two adjacent basic area encodings correspond to the same target font encoding; determining a target font encoding corresponding to each basic area encoding from the at least two font encodings associated with each basic area encoding based on the combination relationship; obtaining a target character corresponding to the target font encoding; and displaying the target text based on the target character. The technical scheme provided by the embodiment of the application can at least realize correct display of the font of the target text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a language processing method, apparatus, device, and computer-readable storage medium. Background Technology

[0002] Existing Optical Character Recognition (OCR) technology and input methods can process various languages, including agglutinative languages. Because agglutinative languages ​​contain multiple basic letters, each corresponding to at least two different glyphs, the glyphs displayed in a word depend on factors such as the position of the basic letters within the word or the combination of basic letters. Therefore, in processing agglutinative languages, existing OCR and input methods cannot automatically select the appropriate glyphs when displaying the text, resulting in an accuracy rate that often falls short of expectations. Summary of the Invention

[0003] This application provides a language processing method, apparatus, device, and computer-readable storage medium, which can at least solve the technical problems of being unable to automatically select glyphs for agglutinative languages ​​and being unable to correctly display glyphs for agglutinative languages.

[0004] On the one hand, this application provides a language processing method, the method comprising:

[0005] Obtain at least two basic region codes of the target text, wherein the basic region codes are associated with at least two character encodings;

[0006] Determine the combination relationship between two adjacent basic area codes in the at least two basic area codes, wherein the combination relationship characterizes whether two adjacent basic area codes correspond to the same target character code;

[0007] Based on the combination relationship, the target character encoding corresponding to each basic area encoding is determined from at least two character encodings associated with each basic area encoding;

[0008] Obtain the target character corresponding to the target character encoding;

[0009] Based on the target character, display the target text.

[0010] On the other hand, a language processing apparatus is provided, the apparatus comprising:

[0011] The first acquisition module acquires at least two basic region codes of the target text, wherein the basic region codes are associated with at least two character encodings;

[0012] The first determining module is used to determine the combination relationship between two adjacent basic area codes in the at least two basic area codes, wherein the combination relationship characterizes whether two adjacent basic area codes correspond to the same target character code.

[0013] The second determining module is used to determine the target character encoding corresponding to each basic area encoding from at least two character encodings associated with each basic area encoding based on the combination relationship;

[0014] The second acquisition module is used to acquire the target character corresponding to the target character encoding.

[0015] The display module is used to display the target text based on the target character.

[0016] On the other hand, a language processing device is provided, the device including a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the language processing method as described above.

[0017] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction, at least one program, code set, or instruction set is stored therein, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the language processing method described above.

[0018] The language processing method, apparatus, device, and computer-readable storage medium provided in this application have the following technical effects:

[0019] This application obtains at least two basic area codes of the target text, determines the combination relationship between two adjacent basic area codes, and, based on the combination relationship, determines the target font code corresponding to each basic area code from at least two font codes associated with each basic area code, thereby achieving automatic selection of at least two basic area codes; and displays the target text according to the target character corresponding to the target font code, thereby achieving correct font display of the target text and improving the readability of the target text. Attached Figure Description

[0020] To more clearly illustrate the technical solutions and advantages in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1This is a flowchart illustrating a language processing method provided in an embodiment of this application;

[0022] Figure 2 This is a flowchart illustrating a method for determining the combination relationship between two adjacent basic area codes provided in an embodiment of this application;

[0023] Figure 3 This is a flowchart illustrating another method for determining the combination relationship between the codes of two adjacent basic regions provided in an embodiment of this application;

[0024] Figure 4 This is a schematic diagram of a basic area coding combination provided in an embodiment of this application;

[0025] Figure 5 This is a flowchart illustrating another method for determining the combination relationship between the codes of two adjacent basic regions provided in an embodiment of this application;

[0026] Figure 6 This is a flowchart illustrating another method for determining the combination relationship between two adjacent basic area codes provided in an embodiment of this application;

[0027] Figure 7 This is a schematic diagram of the character encoding corresponding to different syntax information of a basic area code "0645" provided in the embodiments of this application;

[0028] Figure 8 This is a comparison diagram of the correct and incorrect display of words in a target text provided in an embodiment of this application;

[0029] Figure 9 This is a flowchart illustrating an application example of a language processing method provided in an embodiment of this application;

[0030] Figure 10 This is a schematic diagram of the structure of a language processing device provided in an embodiment of this application;

[0031] Figure 11 This is a schematic diagram of the structure of a server provided in an embodiment of this application. Detailed Implementation

[0032] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0033] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0034] The following describes a language processing method proposed in this application. Figure 1 This is a flowchart illustrating a language processing method provided in an embodiment of this application. This specification provides the operational steps of the method as described in the embodiments or flowchart, but based on conventional or non-inventive labor, more or fewer operational steps may be included. The order of steps listed in the embodiments is merely one possible execution order among many and does not represent the only possible execution order. In actual system or server product execution, the method can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment) as shown in the embodiments or accompanying drawings. Specifically, as... Figure 1 As shown, the method may include:

[0035] S101: Obtain at least two basic area codes of the target text, wherein the basic area codes are associated with at least two character encodings.

[0036] In the embodiments described in this specification, the target text may be the encoded text corresponding to the agglutinative text.

[0037] In practical applications, agglutinative text can include at least two different agglutinative characters, each with a corresponding basic letter. Computers can recognize agglutinative text in images using OCR technology, and can also obtain externally input agglutinative text through an agglutinative input method. When storing the agglutinative text, the encodings of the basic letters corresponding to the glyphs of at least two agglutinative characters are stored to obtain the encoded text corresponding to the agglutinative text.

[0038] In the embodiments of this specification, the at least two basic area codes of the target text correspond to the codes of at least two basic letters, and the at least two font codes associated with the basic area codes correspond to the at least two fonts associated with the basic letters.

[0039] In some embodiments, the encoding types of the at least two basic area codes and the at least two character codes may include, but are not limited to, Unicode.

[0040] Taking the encoded text corresponding to a target text in an Arabic alphabet-based language (an agglutinative language, hereinafter referred to as the language) as an example:

[0041] Ethnic minority languages ​​consist of 32 basic letters, each of which is associated with at least two ethnic minority script characters, resulting in a total of 128 ethnic minority script characters. Computers can recognize ethnic minority language text in images using OCR technology, and can also obtain externally input ethnic minority language text through ethnic minority language input methods. When storing the ethnic minority language text, the codes of the basic letters corresponding to at least two ethnic minority script characters are stored to obtain the encoded text corresponding to the ethnic minority language text.

[0042] In the embodiments of this specification, obtaining at least two basic area codes of the target text can achieve the correct display of the target text's font by automatically selecting the font code of the at least two basic area codes (automatic selection).

[0043] S103: Determine the combination relationship between two adjacent basic area codes in the at least two basic area codes, wherein the combination relationship characterizes whether two adjacent basic area codes correspond to the same target character code.

[0044] In the embodiments of this specification, in order to automatically select at least two basic area codes for the target text, it is necessary to consider the combination relationship between two adjacent basic area codes. For two adjacent basic area codes that can be combined, they need to be treated as a code combination to determine the target character type corresponding to the code combination.

[0045] In the embodiments of this specification, in order to determine the combination relationship between two adjacent basic area codes in at least two basic area codes, the method further includes:

[0046] The processing priority of the at least two basic area codes is determined based on the writing order of the target text.

[0047] Traverse the encodings of at least two basic regions;

[0048] Determine the first code that is adjacent to the currently traversed basic area code and has a higher processing priority than the currently traversed basic area code;

[0049] Determine a second code that is adjacent to the currently traversed basic area code and has a lower processing priority than the currently traversed basic area code.

[0050] In the embodiments of this specification, based on the processing priority of at least two basic area codes, the combination relationship between the first code and the currently traversed basic area code is first determined. Correspondingly, as... Figure 2 As shown, determining the combination relationship between two adjacent basic area codes in the at least two basic area codes includes:

[0051] S201: Obtain at least one preset basic area code combination, wherein the two basic area codes in the basic area code combination correspond to the same target character code.

[0052] In the embodiments of this specification, the two basic region codes that can establish a combination relationship can be exhaustively enumerated according to the language rules of the agglutinative language to obtain at least one basic region code combination, and at least one basic region code combination is preset in the computer.

[0053] S203: Determine whether a first target encoding combination exists in the at least one basic area encoding combination; wherein, the first target encoding combination is a combination of the first encoding and the currently traversed basic area encoding.

[0054] S205: If it is determined that the first target code combination exists, determine whether the currently traversed basic area code and the first code have established a combination relationship. If it is determined that no combination relationship has been established, establish the combination relationship between the currently traversed basic area code and the first code.

[0055] In this embodiment of the specification, if the first code can establish a combination relationship with the currently traversed basic area code, in order to avoid the repeated establishment of the combination relationship, it is also necessary to determine whether the first code has already established a combination relationship with the currently traversed basic area code. If the first code and the currently traversed basic area code have not established a combination relationship, then a combination relationship between the first code and the currently traversed basic area code is established.

[0056] S207: At the end of the traversal, the combination relationship established during the traversal is taken as the combination relationship between two adjacent basic area codes in the at least two basic area codes.

[0057] In the embodiments of this specification, the codes of two adjacent basic regions that establish a combination relationship are used as the combination of basic region codes.

[0058] In the embodiments of this specification, if a combination relationship has already been established between the first code and the currently traversed basic area code, there is no need to process the currently traversed basic area code. Correspondingly, as... Figure 3 As shown, determining the combination relationship between two adjacent basic area codes in the at least two basic area codes includes:

[0059] S301: If it is determined that the first target code combination exists and the currently traversed basic area code has established a combination relationship with the first code, then traverse the next basic area code.

[0060] Taking ethnic languages ​​as an example, such as Figure 4 The diagram shown is a schematic of a basic area encoding combination provided in an embodiment of this specification. It consists of two adjacent basic area codes, "0644" and "0627", and the corresponding character encoding is "FEDF+FE8E".

[0061] Assuming the current basic area code is "0627", the first code is "0644", and the second code is "0634", we can determine that "0644" can be combined with "0627". We then further determine whether "0644" has already been combined with "0627". If no combination has been formed, we establish the combination relationship between "0644" and "0627". If a combination has been formed, we continue to traverse the second code "0634" after "0627".

[0062] In the embodiments of this specification, by determining whether a combination relationship has been established between the first code and the currently traversed basic area code, the problem of repeatedly establishing a combination relationship between the first code and the currently traversed basic area code can be avoided, thereby avoiding the problem of repeatedly performing automatic selection on the first combination code and improving the rigor of the automatic selection scheme of this application.

[0063] Furthermore, without establishing a combination relationship between "0644" and "0627", displaying the characters according to the fonts corresponding to these two basic area codes will result in incorrect characters. Given the established combination relationship between "0644" and "0627", the target character encoding corresponding to the resulting basic area encoding combination is determined to be "FEDF+FE8E". Displaying the character according to the character encoding "FEDF+FE8E" will yield the correct character.

[0064] In the embodiments of this specification, by using the codes of two adjacent basic regions that establish a combination relationship as the basic region code combination, the font corresponding to the basic region code combination is automatically selected, which can improve the accuracy of the font of the displayed target text and greatly increase the readability of the target text.

[0065] In the embodiments of this specification, based on the processing priority of at least two basic area codes, if the first code cannot be combined with the currently traversed basic area code, the combination relationship between the currently traversed basic area code and the second code is then determined. Correspondingly, as... Figure 5 As shown, determining the combination relationship between two adjacent basic area codes in the at least two basic area codes includes:

[0066] S501: If it is determined that the first target encoding combination does not exist, determine whether a second target encoding combination exists among the at least one basic area encoding combination; wherein, the second target encoding combination is a combination of the currently traversed encoding and the second encoding;

[0067] S503: If it is determined that the second target code combination exists, establish the combination relationship between the currently traversed basic area code and the second code.

[0068] Taking ethnic languages ​​as an example, such as Figure 6 As shown, this is a basic area encoding combination provided in an embodiment of this specification, which consists of two adjacent basic area codes "0644" and "0627", and the corresponding character encoding is "FEDF+FE8E".

[0069] Assuming the current traversed basic area code is "0644", the first code is "0634", and the second code is "0627", it can be determined that "0634" and "0644" cannot be combined, but "0644" can be combined with "0627", thus establishing the combination relationship between "0644" and "0627".

[0070] In the embodiments of this specification, based on the processing priority, it is first determined whether the first code and the currently traversed basic area code can be combined. If the first code cannot be combined with the currently traversed basic area code, then the currently traversed basic area code is combined. This can avoid the currently traversed basic area code from being incorrectly combined with the second code when it can be combined with the first code, thereby improving the rigor of the automatic selection scheme.

[0071] In the embodiments of this specification, when the currently traversed basic area code cannot establish a combination relationship with the second code, the corresponding approach is as follows: Figure 6 As shown, determining the combination relationship between two adjacent basic area codes in the at least two basic area codes includes:

[0072] S601: If it is determined that neither the first target encoding combination nor the second target encoding combination exists, the target syntax information of the currently traversed basic area encoding in the target text is determined;

[0073] S603: Based on the target syntax information, determine the target character encoding corresponding to the currently traversed basic area encoding from at least two character encodings associated with the currently traversed basic area encoding.

[0074] In the embodiments of this specification, for two adjacent basic area codes that can establish a combination relationship, the two adjacent basic area codes are treated as a whole (basic area code combination) and automatically selected. However, for cases where the currently traversed basic area code cannot be combined with the corresponding first or second code, automatic selection is performed based on the target syntax information of the currently traversed basic area code in the target text.

[0075] In the embodiments of this specification, the target syntax information of the currently traversed basic area encoding in the target text can be one of the different syntax information mentioned above.

[0076] In the embodiments of this specification, the factors affecting the selection of target character encoding corresponding to the basic area encoding are fully considered, which can improve the accuracy and rigor of the automatic character selection scheme.

[0077] In an optional embodiment, to determine the target syntax information of the currently traversed basic region encoding in the target text, determining the target syntax information of the currently traversed basic region encoding in the target text includes:

[0078] Determine the types of the first and second codes, including text types and symbol types;

[0079] Based on the types of the first and second codes, the target syntax information of the target text is determined by the basic area code currently being traversed.

[0080] In the embodiments of this specification, at least one basic region code of the target text may include a symbol-type basic region code and / or a text-type basic region code, wherein the symbol-type basic region code is used to separate the text-type basic region codes. Based on this, according to the types of the first code and the second code, the target syntax information of the currently traversed basic region code in the target text can be determined.

[0081] In one specific embodiment, determining the target grammatical information of the currently traversed basic region encoding in the words of the target text based on the types of the first encoding and the second encoding includes:

[0082] When both the first and second encodings are symbolic, the target syntax information of the currently traversed basic area encoding is determined to be: the currently traversed basic area encoding is a word of the target text;

[0083] When the type of the first encoding is symbol type and the type of the second encoding is text type, the target syntax information of the currently traversed basic area encoding is determined as follows: the currently traversed basic area encoding is located at the beginning of a word in the target text.

[0084] When both the first and second encodings are text types, the target syntax information of the currently traversed basic area encoding is determined as follows: the currently traversed basic area encoding is located in the word of the target text.

[0085] When the type of the first encoding is text and the type of the second encoding is symbol, the target syntax information of the currently traversed basic area encoding is determined as follows: the currently traversed basic area encoding is located at the end of the word of the target text.

[0086] In the embodiments described in this specification, the target character encoding is different depending on the grammatical information of the basic area encoding currently being traversed in the target text.

[0087] In an optional embodiment, in order to automatically select the basic region encoding based on the target syntax information of the basic region encoding in the target text, the method further includes:

[0088] A second mapping information is pre-set, which represents the mapping relationship between different syntactic information of the currently traversed basic area code and the corresponding target character code; wherein, the different syntactic information of the currently traversed basic area code includes at least two of the following: the currently traversed basic area code is a word of the target text; the currently traversed basic area code is located at the beginning of a word in the target text; the currently traversed basic area code is located in the middle of a word in the target text; the currently traversed basic area code is located at the end of a word in the target text.

[0089] Accordingly, determining the target character encoding corresponding to the currently traversed basic area encoding from at least two character encodings associated with the currently traversed basic area encoding based on the target syntax information includes:

[0090] Based on the second mapping relationship and the target syntax information, the target character encoding corresponding to the currently traversed basic area encoding is determined from at least two character encodings associated with the currently traversed basic area encoding.

[0091] In the embodiments of this specification, based on the language rules of the agglutinative language corresponding to the target text, the character encoding corresponding to different grammatical information of the currently traversed basic area encoding is determined, and a mapping relationship between different grammatical information and the corresponding character encoding is pre-established to obtain the second mapping information.

[0092] Taking ethnic languages ​​as an example, such as Figure 7The diagram shown illustrates the character encoding corresponding to different grammatical information of a basic area code "0645" provided in this embodiment. Specifically, the character encoding corresponding to a word in the target text is an independent encoding; the character encoding corresponding to the beginning of a word in the target text is a follow-up encoding; the character encoding corresponding to the middle of a word in the target text is a front-back encoding; and the character encoding corresponding to the end of a word in the target text is a front-back encoding. Furthermore, Figure 7 It also displays the characters corresponding to each glyph encoding.

[0093] In this case, the basic letters in the corresponding ethnic language of the basic area code are consonants, which is the same as the independent code. If the basic letters in the corresponding ethnic language of the basic area code are vowels, which is the same as the independent code, the at least two related character codes also include simple independent codes, which are not enumerated here.

[0094] In the embodiments of this specification, by pre-setting a second mapping relationship, the target character encoding can be quickly determined from at least two character encodings associated with the currently traversed basic area encoding in the target text based on the target syntax information of the currently traversed basic area encoding, thereby improving the language processing speed.

[0095] In the embodiments of this specification, at least one basic region code of the target text may also include at least one special basic region code. This special basic region code cannot be combined with adjacent first and second codes, nor can its corresponding target character code be determined solely based on its target grammar information in the target text. For this type of special basic region code, it is necessary to determine its corresponding target character code by considering which specific first code(s) it corresponds to.

[0096] In an optional embodiment, if the currently traversed basic region code is a special basic region code, the method further includes:

[0097] At least one set of codes is pre-defined;

[0098] A third mapping information is preset, which represents the mapping relationship between the first encoding and the corresponding target character encoding when the first encoding is an encoding in a preset encoding set;

[0099] Accordingly, determining the target character encoding corresponding to the currently traversed basic area encoding from at least two character encodings associated with the currently traversed basic area encoding based on the target syntax information further includes:

[0100] If the currently traversed basic area code is located in a word in the target text, determine whether the first code belongs to at least one preset code set;

[0101] Based on the judgment result, determine the encoding set to which the first encoding belongs from at least one encoding set;

[0102] Based on the third mapping information and the encoding set to which the first encoding belongs, the target character encoding corresponding to the currently traversed basic area encoding is determined.

[0103] In this embodiment of the specification, when the currently traversed basic region code is located within a word in the target text, its corresponding target character code still needs to be determined in conjunction with the specific code of the first code. It can be understood that the type of the first code is a text type, and the type of the basic region code in at least one code set is a text type.

[0104] In the embodiments of this specification, based on the language rules of the agglutinative language corresponding to the target text, if the basic area code currently traversed is located in a word in the target text and the corresponding target character code is a different character code, it can be used as the (basic area) code set of the first code. A mapping relationship between different character codes and the corresponding basic area code set is established in advance to obtain the third mapping relationship.

[0105] To facilitate the representation and processing of the encoding set, in an optional embodiment, all glyph codes can be re-encoded starting from 1. Taking a minority language as an example, which has 128 glyph codes, these glyph codes can be re-encoded according to 1-128.

[0106] Taking the basic area code "FBF4" as an example, it is re-encoded to obtain a new code "6". When it is located in a word in the target text, if the first code belongs to the set {9,10,11,12,13,65,66,67,68,69} represented by the new code, the new code corresponding to "FBF4" is 50, and the target character code corresponding to the new code 50 is "06C8"; among them, the character codes corresponding to the new codes 9-13 are "062F, 0631, 0632, 0698, 06CB" respectively, and the character codes corresponding to the new codes 65-69 are "FEAA, FEAE, FEB0, FE8B, FBDF" respectively.

[0107] S105: Based on the combination relationship, determine the target character encoding corresponding to each basic area encoding from at least two character encodings associated with each basic area encoding.

[0108] In the embodiments of this specification, for two adjacent basic area codes that can establish a combination relationship, the two adjacent basic area codes are treated as a whole (basic area code combination) and the character code corresponding to the basic area code combination is automatically selected as the target character code.

[0109] In an optional embodiment, to determine the character encoding corresponding to the basic area encoding combination, the method further includes:

[0110] A first mapping information is pre-set, which represents the mapping relationship between each basic area encoding combination and the corresponding target character encoding;

[0111] Accordingly, determining the target character encoding corresponding to each basic area encoding from at least two character encodings associated with each basic area encoding based on the combination relationship includes:

[0112] When the currently traversed basic area code has a combination relationship with the second code, the target character code corresponding to the second target code combination is determined from the target character codes corresponding to each basic area code combination according to the first mapping information.

[0113] In the embodiments of this specification, based on the language rules of the agglutinative language corresponding to the target text, the character encoding corresponding to at least one basic area encoding combination is determined, and a mapping relationship between at least one basic area encoding and the corresponding character encoding is pre-established to obtain the first mapping information.

[0114] Specifically, the first mapping information can be found by referring to Figure 8 As shown.

[0115] S107: Obtain the target character corresponding to the target character encoding;

[0116] In the embodiments of this specification, different character encodings correspond to different characters.

[0117] In an optional embodiment, a mapping relationship between different font codes and their corresponding characters is pre-established to obtain fourth mapping information. Based on the fourth mapping information, the target character corresponding to the target font code is obtained.

[0118] S109: Display the target text based on the target character.

[0119] In the embodiments of this specification, after determining the target font encoding corresponding to at least two basic area encodings of the target text, the target text is displayed according to the target character corresponding to the target font encoding.

[0120] like Figure 8 The image shown is a comparison chart of correct and incorrect display of words in a target text provided in an embodiment of this specification:

[0121] Taking ethnic minority languages ​​as an example, the writing order of ethnic minority languages ​​is from right to left.

[0122] Figure 8 The text involves a word composed of three basic area codes, located at the beginning, middle, and end of the word. If no automatic transformation of the three basic area codes is performed, the word will be displayed based on the characters corresponding to the basic area codes, resulting in... Figure 8 The text does not display any selection errors; if the language processing method in the embodiments of this specification is followed, the target character encoding corresponding to the three basic area codes is automatically selected, and the words are displayed according to the target characters corresponding to the target character encoding, then... Figure 8 The correct selection is displayed.

[0123] In the embodiments of this specification, by obtaining at least two basic area codes of the target text and automatically converting the at least two basic area codes, the correct display of the target text's font can be achieved, thereby improving the readability of the target text.

[0124] To illustrate in detail the application of the language processing methods provided in the embodiments of the above specification, such as Figure 9 As shown, taking a minority language as an example, this paper introduces an application example of the language processing method provided in this application. The specific process includes:

[0125] S901: Obtain the complete basic zone code of a word in a minority language: "0644+0627+……+0649"; where "+" is the separator between each basic zone code to facilitate the differentiation of each basic zone code.

[0126] S902: Obtain at least one basic area coding combination pre-set according to the language rules of the ethnic language;

[0127] S903: Traverse each basic zone code in “0644+0627+……+0649”, and determine “0644” and “0627” as two adjacent basic zone codes that can establish a combination relationship in “0644+0627+……+0649” according to at least one preset basic zone code combination.

[0128] S904: Establish the combination relationship between "0644" and "0627", and represent the target code combination of "0644" and "0627" as "(0644+0627)";

[0129] S905: For the basic area code "0649" which cannot establish a combination relationship with the adjacent basic area codes, the target grammatical information of "0649" is determined to be the word ending of a word in the ethnic language;

[0130] S906: Based on the preset first mapping information, determine the target character encoding corresponding to "(0644+0627)" as "(FEDF+FE8E)";

[0131] S907: Based on the preset second mapping information, if the target grammar information of “0649” is determined to be located at the end of a word in a minority language, the corresponding target character encoding is its preceding encoding “FEF0”.

[0132] S908: Display the word in the ethnic language based on the target character corresponding to “(FEDF+FE8E)+……+FEF0”.

[0133] The above S901 to S908 can be understood as the process of forward encoding of all basic region codes of words in ethnic languages.

[0134] In an optional embodiment, to ensure that the forward encoding process is error-free, the above process may further include a reverse decoding process, the specific process of which is as follows:

[0135] After executing S907, execute S909.

[0136] S909: Based on the preset first mapping information and second mapping information, reverse decode “(FEDF+FE8E)+……+FEF0”. If the multiple basic area codes obtained by reverse decoding are “0644+0627+……+0649”, it is determined that there were no errors in the forward encoding process; execute S908.

[0137] In the embodiments of this specification, the target text is encoded in both forward and reverse directions according to the language processing method described above. The results of forward encoding and reverse encoding are mapped to each other, which can reduce language processing errors and ensure the correct display of the target text's font.

[0138] like Figure 10 The image shown illustrates a language processing apparatus provided in an embodiment of this application. (Refer to...) Figure 10 The device includes:

[0139] The first acquisition module 1001 acquires at least two basic area codes of the target text, wherein the basic area codes are associated with at least two character encodings;

[0140] The first determining module 1002 is used to determine the combination relationship between two adjacent basic area codes in the at least two basic area codes, wherein the combination relationship indicates whether two adjacent basic area codes correspond to the same target character code;

[0141] The second determining module 1003 is used to determine the target character encoding corresponding to each basic area encoding from at least two character encodings associated with each basic area encoding based on the combination relationship.

[0142] The second acquisition module 1004 is used to acquire the target character corresponding to the target character encoding.

[0143] Display module 1005 is used to display the target text based on the target character.

[0144] In some embodiments, the apparatus further includes:

[0145] The third determining module is used to determine the processing priority of the at least two basic area codes according to the writing order of the target text;

[0146] The traversal module is used to traverse the at least two basic area codes;

[0147] The fourth determining module is used to determine the first code that is adjacent to the currently traversed basic area code and has a higher processing priority than the currently traversed basic area code;

[0148] The fifth determining module is used to determine a second code that is adjacent to the currently traversed basic area code and has a lower processing priority than the currently traversed basic area code;

[0149] Accordingly, the first determining module 1002 further includes:

[0150] The acquisition unit is used to acquire at least one preset basic area code combination, wherein the two basic area codes in the basic area code combination correspond to the same target character code.

[0151] The first determination unit is used to determine whether a first target encoding combination exists in the at least one basic area encoding combination; wherein, the first target encoding combination is a combination of the first encoding and the currently traversed basic area encoding;

[0152] The first establishment unit is used to determine whether the currently traversed basic area code and the first code have been established as a combination relationship when it is determined that the first target code combination exists; and to establish the combination relationship between the currently traversed basic area code and the first code when it is determined that no combination relationship has been established.

[0153] The first determining unit is used to, at the end of the traversal, take the combination relationship established during the traversal as the combination relationship between two adjacent basic area codes in the at least two basic area codes.

[0154] Accordingly, the first determining module 1002 further includes:

[0155] The second determination unit is used to determine whether a second target encoding combination exists among the at least one basic area encoding combination when it is determined that the first target encoding combination does not exist; wherein, the second target encoding combination is a combination of the currently traversed encoding and the second encoding;

[0156] The second establishment unit, upon determining that the second target code combination exists, establishes a combination relationship between the currently traversed basic area code and the second code.

[0157] Accordingly, the first determining module 1002 further includes:

[0158] The traversal unit is used to traverse the next basic area code when it is determined that the first target code combination exists and the currently traversed basic area code and the first code have established a combination relationship.

[0159] In some embodiments, the apparatus further includes:

[0160] The first mapping module is used to pre-set the first mapping information, which represents the mapping relationship between each basic area encoding combination and the corresponding target character encoding.

[0161] Accordingly, the second determining module 1003 includes:

[0162] The determining unit is used to determine the target character encoding corresponding to the second target encoding combination from the target character encodings corresponding to each basic area encoding combination, based on the first mapping information, when the currently traversed basic area encoding has a combination relationship with the second encoding.

[0163] In some embodiments, the first determining module 1002 further includes:

[0164] The second determining unit is used to determine the target syntax information of the currently traversed basic area code in the target text when it is determined that neither the first target code combination nor the second target code combination exists.

[0165] The third determining unit is used to determine the target character encoding corresponding to the currently traversed basic area encoding from at least two character encodings associated with the currently traversed basic area encoding based on the target syntax information.

[0166] In some embodiments, the third determining unit includes:

[0167] The first determining subunit is used to determine the types of the first code and the second code, the types including text type and symbol type;

[0168] The second determining subunit is used to determine the target syntax information of the currently traversed basic area encoding in the target text based on the types of the first encoding and the second encoding.

[0169] In some embodiments, the second determining subunit is specifically used for:

[0170] When both the first and second encodings are symbolic, the target syntax information of the currently traversed basic area encoding is determined to be: the currently traversed basic area encoding is a word of the target text;

[0171] When the type of the first encoding is symbol type and the type of the second encoding is text type, the target syntax information of the currently traversed basic area encoding is determined as follows: the currently traversed basic area encoding is located at the beginning of a word in the target text.

[0172] When both the first and second encodings are text types, the target syntax information of the currently traversed basic area encoding is determined as follows: the currently traversed basic area encoding is located in the word of the target text.

[0173] When the type of the first encoding is text and the type of the second encoding is symbol, the target syntax information of the currently traversed basic area encoding is determined as follows: the currently traversed basic area encoding is located at the end of the word of the target text.

[0174] In some embodiments, the apparatus further includes:

[0175] The second mapping module is used to pre-set second mapping information, which represents the mapping relationship between different syntactic information of the currently traversed basic area code and the corresponding target character code; wherein, the different syntactic information of the currently traversed basic area code includes at least two of the following: the currently traversed basic area code is a word of the target text; the currently traversed basic area code is located at the beginning of a word in the target text; the currently traversed basic area code is located in the middle of a word in the target text; the currently traversed basic area code is located at the end of a word in the target text.

[0176] Accordingly, the third determining unit further includes:

[0177] The third determining subunit is used to determine the target character encoding corresponding to the currently traversed basic area encoding from at least two character encodings associated with the currently traversed basic area encoding, based on the second mapping relationship and the target syntax information.

[0178] In some embodiments, the apparatus further includes:

[0179] The configuration module is used to pre-configure at least one set of encodings;

[0180] The third mapping module is used to pre-set third mapping information, which represents the mapping relationship between the first code and the corresponding target character code when the first code is a code in a preset set of codes.

[0181] Correspondingly, the third determining unit further includes:

[0182] The judgment subunit is used to determine whether the first code belongs to at least one preset code set when the currently traversed basic area code is located in the word of the target text.

[0183] The fourth determining subunit is used to determine the encoding set to which the first encoding belongs from at least one encoding set based on the judgment result;

[0184] The fifth determining subunit is used to determine the target character encoding corresponding to the currently traversed basic area encoding based on the third mapping information and the encoding set to which the first encoding belongs.

[0185] The apparatus and method embodiments described herein are based on the same inventive concept.

[0186] This application also provides a computer-readable storage medium storing at least one instruction or at least one program segment, which is loaded and executed by a processor to implement the aforementioned language processing method.

[0187] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various alternative implementations described above.

[0188] This application provides a language processing server, which includes a processor and a memory. The memory stores at least one instruction, at least one program, code set, or instruction set. The at least one instruction, at least one program, code set, or instruction set is loaded and executed by the processor to implement the language processing method provided in the above method embodiments.

[0189] Memory can be used to store software programs and modules. The processor executes various functional applications and language processing by running the software programs and modules stored in the memory. Memory can primarily include a program storage area and a data storage area. The program storage area can store the operating system, application programs required for functions, etc.; the data storage area can store data created based on the use of the device, etc. Furthermore, memory can include high-speed random access memory, and can also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, memory can also include a memory controller to provide the processor with access to the memory.

[0190] The methods and embodiments provided in this application can be executed on mobile terminals, computer terminals, servers, or similar computing devices. Taking running on a server as an example, Figure 11 This is a hardware structure block diagram of a server for a language processing method provided in an embodiment of this application. For example... Figure 11 As shown, the server 1100 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 1110 (CPUs 1110 may include, but are not limited to, microprocessors (MCUs) or programmable logic devices (FPGAs), a memory 1130 for storing data, and one or more storage media 1120 (e.g., one or more mass storage devices) for storing application programs 1123 or data 1122. The memory 1130 and storage media 1120 may be temporary or persistent storage. The program stored in the storage media 1120 may include one or more modules, each module including a series of instruction operations on the server. Furthermore, the CPU 1110 may be configured to communicate with the storage media 1120 and execute the series of instruction operations stored in the storage media 1120 on the server 1100. Server 1100 may also include one or more power supplies 1160, one or more wired or wireless network interfaces 1150, one or more input / output interfaces 1140, and / or one or more operating systems 1121, such as Windows Server™, Mac OSX™, Unix™, Linux™, FreeBSD™, etc.

[0191] The input / output interface 1140 can be used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of server 1100. In one example, the input / output interface 1140 includes a network interface controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the input / output interface 1140 may be a radio frequency (RF) module for wireless communication with the Internet.

[0192] Those skilled in the art will understand that Figure 11 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, server 1100 may also include... Figure 11 The more or fewer components shown, or having the same Figure 11 The different configurations shown.

[0193] The embodiments of this application also provide a storage medium, which can be disposed in a server to store at least one instruction, at least one program, code set or instruction set related to implementing a language processing method in the method embodiments. The at least one instruction, the at least one program, the code set or instruction set is loaded and executed by the processor to implement the language processing method provided in the above method embodiments.

[0194] Optionally, in this embodiment, the storage medium may be located at at least one of the multiple network servers in a computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0195] As can be seen from the embodiments of the language processing method, apparatus, server or storage medium provided in this application, by determining the combination relationship between two adjacent basic area codes in at least two basic area codes, the at least two basic area codes of the target text can be automatically selected based on the combination relationship, so as to achieve the correct display of the font of the target text and greatly improve the readability of the target text.

[0196] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, specific embodiments have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0197] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0198] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0199] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A language processing method, characterized in that, The method includes: Obtain at least two basic region codes of the target text, wherein the basic region codes are associated with at least two character encodings; Determine the combination relationship between two adjacent basic area codes in the at least two basic area codes, wherein the combination relationship characterizes whether two adjacent basic area codes correspond to the same target character code; Based on the combination relationship, the target character encoding corresponding to each basic area encoding is determined from at least two character encodings associated with each basic area encoding; Obtain the target character corresponding to the target character encoding; Based on the target character, display the target text; The step of determining the target character encoding corresponding to each basic area encoding from at least two character encodings associated with each basic area encoding based on the combination relationship includes: The target character encoding corresponding to the basic area encoding combination is determined as the target character encoding corresponding to the two adjacent basic area encodings constituting the basic area encoding combination. The basic area encoding combination is composed of two adjacent basic area encodings that establish a combination relationship. The two basic area encodings in the basic area encoding combination correspond to the same target character encoding. The target character encoding corresponding to the basic area encoding combination is determined according to a pre-set first mapping information. The first mapping information represents the mapping relationship between each basic area encoding combination and the corresponding target character encoding.

2. The method according to claim 1, characterized in that, The method further includes: The processing priority of the at least two basic area codes is determined based on the writing order of the target text. Traverse the encodings of at least two basic regions; Determine the first code that is adjacent to the currently traversed basic area code and has a higher processing priority than the currently traversed basic area code; Determine a second code that is adjacent to the currently traversed basic area code and has a lower processing priority than the currently traversed basic area code; Accordingly, determining the combination relationship between two adjacent basic area codes in the at least two basic area codes includes: Obtain at least one preset basic area code combination, wherein two basic area codes in the basic area code combination correspond to the same target character code; Determine whether a first target encoding combination exists among the at least one basic area encoding combination; wherein, the first target encoding combination is a combination of the first encoding and the currently traversed basic area encoding; If it is determined that the first target code combination exists, it is determined whether the currently traversed basic area code and the first code have established a combination relationship. If it is determined that no combination relationship has been established, the combination relationship between the currently traversed basic area code and the first code is established. At the end of the traversal, the combination relationship established during the traversal is taken as the combination relationship between two adjacent basic area codes in the at least two basic area codes.

3. The method according to claim 2, characterized in that, Determining the combination relationship between two adjacent basic region codes in the at least two basic region codes further includes: If it is determined that the first target encoding combination does not exist, it is determined whether a second target encoding combination exists among the at least one basic area encoding combination; wherein, the second target encoding combination is a combination of the currently traversed encoding and the second encoding; If the existence of the second target code combination is determined, a combination relationship is established between the currently traversed basic area code and the second code.

4. The method according to claim 2, characterized in that, Determining the combination relationship between two adjacent basic region codes in the at least two basic region codes further includes: If it is determined that the first target code combination exists and that the currently traversed basic area code has established a combination relationship with the first code, then traverse the next basic area code.

5. The method according to claim 3, characterized in that, The method further includes: A first mapping information is pre-set, which represents the mapping relationship between each basic area encoding combination and the corresponding target character encoding; Accordingly, determining the target character encoding corresponding to each basic area encoding from at least two character encodings associated with each basic area encoding based on the combination relationship includes: When the currently traversed basic area code has a combination relationship with the second code, the target character code corresponding to the second target code combination is determined from the target character codes corresponding to each basic area code combination according to the first mapping information.

6. The method according to claim 3, characterized in that, Determining the combination relationship between two adjacent basic region codes in the at least two basic region codes further includes: If it is determined that neither the first target encoding combination nor the second target encoding combination exists, the target syntax information of the currently traversed basic area encoding in the target text is determined; Based on the target syntax information, the target character encoding corresponding to the currently traversed basic area encoding is determined from at least two character encodings associated with the currently traversed basic area encoding.

7. The method according to claim 6, characterized in that, Determining the target syntax information of the currently traversed basic region encoding in the target text includes: Determine the types of the first and second codes, including text types and symbol types; Based on the types of the first and second codes, the target syntax information of the target text is determined by the basic area code currently being traversed.

8. The method according to claim 7, characterized in that, The step of determining the target grammar information of the currently traversed basic region code in the words of the target text based on the types of the first and second codes includes: When both the first and second encodings are symbolic, the target syntax information of the currently traversed basic area encoding is determined to be: the currently traversed basic area encoding is a word of the target text; When the type of the first encoding is symbol type and the type of the second encoding is text type, the target syntax information of the currently traversed basic area encoding is determined as follows: the currently traversed basic area encoding is located at the beginning of a word in the target text. When both the first and second encodings are text types, the target syntax information of the currently traversed basic area encoding is determined as follows: the currently traversed basic area encoding is located in the word of the target text. When the type of the first encoding is text and the type of the second encoding is symbol, the target syntax information of the currently traversed basic area encoding is determined as follows: the currently traversed basic area encoding is located at the end of the word of the target text.

9. The method according to claim 8, characterized in that, The method further includes: A second mapping information is pre-set, which represents the mapping relationship between different syntactic information of the currently traversed basic area code and the corresponding target character code; wherein, the different syntactic information of the currently traversed basic area code includes at least two of the following: the currently traversed basic area code is a word of the target text; the currently traversed basic area code is located at the beginning of a word in the target text; the currently traversed basic area code is located in the middle of a word in the target text; the currently traversed basic area code is located at the end of a word in the target text. Accordingly, determining the target character encoding corresponding to the currently traversed basic area encoding from at least two character encodings associated with the currently traversed basic area encoding based on the target syntax information includes: Based on the second mapping information and the target syntax information, the target character encoding corresponding to the currently traversed basic area encoding is determined from at least two character encodings associated with the currently traversed basic area encoding.

10. The method according to claim 8, characterized in that, The method further includes: At least one set of codes is pre-defined; A third mapping information is preset, which represents the mapping relationship between the first encoding and the corresponding target character encoding when the first encoding is an encoding in a preset encoding set; Accordingly, determining the target character encoding corresponding to the currently traversed basic area encoding from at least two character encodings associated with the currently traversed basic area encoding based on the target syntax information further includes: If the currently traversed basic area code is located in a word in the target text, determine whether the first code belongs to at least one preset code set; Based on the judgment result, determine the encoding set to which the first encoding belongs from at least one encoding set; Based on the third mapping information and the encoding set to which the first encoding belongs, the target character encoding corresponding to the currently traversed basic area encoding is determined.

11. A language processing device, characterized in that, The device includes: The first acquisition module acquires at least two basic region codes of the target text, wherein the basic region codes are associated with at least two character encodings; The first determining module is used to determine the combination relationship between two adjacent basic area codes in the at least two basic area codes, wherein the combination relationship characterizes whether two adjacent basic area codes correspond to the same target character code. The second determining module is used to determine the target character encoding corresponding to each basic area encoding from at least two character encodings associated with each basic area encoding based on the combination relationship. The second acquisition module is used to acquire the target character corresponding to the target character encoding. The display module is used to display the target text based on the target character; The step of determining the target character encoding corresponding to each basic area encoding from at least two character encodings associated with each basic area encoding based on the combination relationship includes: The target character encoding corresponding to the basic area encoding combination is determined as the target character encoding corresponding to the two adjacent basic area encodings constituting the basic area encoding combination. The basic area encoding combination is composed of two adjacent basic area encodings that establish a combination relationship. The two basic area encodings in the basic area encoding combination correspond to the same target character encoding. The target character encoding corresponding to the basic area encoding combination is determined according to a pre-set first mapping information. The first mapping information represents the mapping relationship between each basic area encoding combination and the corresponding target character encoding.

12. A language processing device, characterized in that, The device includes a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the language processing method as described in any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction or at least one program segment, which is loaded and executed by a processor to implement the language processing method as described in any one of claims 1 to 10.

Citation Information

Patent Citations

  • Thai display method, device and system

    CN102508822A

  • Uyghur Kazak and Kirghiz display method and application

    CN103870439A