Reading method estimation device, reading method estimation method, reading method estimation program, and recording medium

The pronunciation estimation device addresses the challenges of dictionary dependency and labor-intensive updates by using a neural machine translation model to convert character strings into pronunciations, achieving robust and efficient pronunciation estimation for trademarks and other character strings.

JP7692715B2Active Publication Date: 2025-06-16GENERAL FINANCIAL INSTITUTION JAPAN CHARTERED INFORMATION AGENCY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2021053896
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-03-26
Publication Date
2025-06-16
Estimated Expiration
2041-03-26

AI Technical Summary

Technical Problem

Existing pronunciation estimation devices for character strings, particularly trademarks, face challenges in robustness against unknown words and require significant labor for dictionary tuning and updating.

Method used

A pronunciation estimation device that splits input character strings into characters and converts them into pronunciations using a neural machine translation model, specifically a transformer, without relying on a dictionary database.

Benefits of technology

Enables accurate estimation of pronunciations for various character strings, including trademarks, without the need for a dictionary database, thereby reducing labor and improving robustness against unknown words.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007692715000003
    Figure 0007692715000003
  • Figure 0007692715000004
    Figure 0007692715000004
  • Figure 0007692715000005
    Figure 0007692715000005
Patent Text Reader

Abstract

To provide a reading way estimation device capable of estimating a way of reading various character strings, a name of a trademark in particular without creating a lexical database about a way of reading character strings.SOLUTION: A reading way estimation device 100 divides an inputted character string into character units, and converts characters into a reading way composed of Katakana or Hiragana by allowing a transformer having undergone learning with arrangement standardization about trademarks to sequentially generate Katakana characters or Hiragana characters one by one from the character string divided into character units.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a reading method estimation device, a reading method estimation method, a reading method estimation program, and a recording medium. In particular, the present invention relates to a device, a method, a program, and a recording medium for estimating the reading method of a character string composed of alphabet, hiragana, katakana, kanji, numbers, and other character types.

Background Art

[0002] In recent years, since the Internet has become established as a social infrastructure, text-based information transmission has increased regardless of whether it is in Japanese or a foreign language. On the other hand, since phoneme information of a character string is generally not transmitted along with the text information, although the character string can be easily transmitted, there are many cases where the reading method of the character string is unknown, and the demand for a character string reading method estimation device or method is increasing.

[0003] Examples of character strings composed of alphabet, hiragana, katakana, kanji, numbers, and other character types, for which it is difficult to determine the reading method, include personal names, place names, company / group names, kanji, foreign proper nouns, and trademarks.

[0004] By the way, the scope of similarity of trademarks is determined by comprehensively considering the impressions, memories, associations, etc. given to consumers, etc. by the appearance, concept, name, etc. of the trademark. Therefore, conventionally, in order to construct a trademark search system, businesses, etc. have created "display trademarks" and "names" from character trademarks and stored them in a database.

[0005] In this specification, the "character trademark" is defined to include not only trademarks consisting only of character elements but also the character element parts in combined trademarks of graphics and characters. Also, the "name" of a trademark is defined as the reading method of the trademark recognized by consumers in contact with the trademark. When a consumer reads a trademark, there may be cases where general nouns without self-other discrimination ability, such as "Limited Company", are excluded from the reading. In this regard, the "name" of a trademark is different from the reading method of ordinary character strings.

[0006] Estimating the pronunciation from a word mark itself does not impose a significant workload on the operator. However, creating pronunciations for a large number of trademarks as large as building a database is a very heavy workload, and its automation has been desired conventionally.

[0007] Patent Document 1 describes a pronunciation estimation device using morphological analysis and a reading dictionary.

Prior Art Documents

Patent Documents

[0008]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0009] However, since the technique described in Patent Document 1 uses a method of registering candidates registered in advance for each morpheme as a dictionary when generating a pronunciation, its robustness against unknown words is low, and furthermore, in actual operation, there has been a problem that the labor of tuning / updating the dictionary occurs.

[0010] Therefore, an object of the present invention is to provide a pronunciation estimation device, a pronunciation estimation method, etc. that can estimate the pronunciations of various character strings, particularly the appellations of trademarks, without creating a dictionary database related to the pronunciations of character strings.

Means for Solving the Problems

[0011] According to one embodiment, a pronunciation estimation device is provided. This pronunciation estimation device includes a string splitter that splits the input character string into characters, and a converter that converts the character string split into characters into a pronunciation consisting of katakana or hiragana characters by sequentially generating one katakana or hiragana character at a time from the character string split into characters.

[0012] The objects and advantages of the present invention are realized and achieved by the elements and combinations particularly pointed out in the claims. It should be understood that both the foregoing general description and the following detailed description are exemplary and explanatory and are not restrictive of the invention, as claimed.

Effect of the Invention

[0013] The reading method estimation apparatus and method disclosed in this specification can provide a reading method estimation apparatus and method that can estimate the reading methods of various character strings, particularly the appellations of trademarks, without creating a special dictionary database regarding the reading methods of character strings.

Brief Description of the Drawings

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Mode for Carrying Out the Invention

[0015] As used herein, the term "character string" means a string consisting of alphabet (including alphabet with diacritical marks such as French accents and German umlauts), hiragana, katakana, Chinese characters (including simplified Chinese characters and traditional Chinese characters in addition to Chinese characters used in Japanese), numbers, special characters constituting display trademarks, and other character types, as well as combinations thereof, but is not limited thereto. Examples of other character types include character types used in minor languages such as Cyrillic characters, and special characters constituting display trademarks. In addition, the objects indicated by the character string include, but are not limited to, personal names, place names, names of companies and organizations, Chinese characters, proper nouns in foreign languages, and trademarks.

Example

[0016] FIG. 1 is a block diagram of a computer system 300 for implementing an aspect according to an embodiment of the present disclosure. The mechanisms and apparatuses of various embodiments disclosed herein may be applied to any suitable computing system. The main components of the computer system 300 include one or more processors 302, a memory 304, a terminal interface 312, a storage interface 314, an I / O (input / output) device interface 316, and a network interface 318. These components may be interconnected via a memory bus 306, an I / O bus 308, a bus interface unit 309, and an I / O bus interface unit 310.

[0017] The computer system 300 may include one or more general-purpose programmable central processing units (CPUs) 302A and 302B collectively referred to as the processor 302. In certain embodiments, the computer system 300 may comprise multiple processors, and in other embodiments, the computer system 300 may be a single CPU system. Each processor 302 executes instructions stored in the memory 304 and may include an on-board cache.

[0018] In one embodiment, the memory 304 may include a random access semiconductor memory, a storage device, or a storage medium (either volatile or non-volatile) for storing data and programs. In one embodiment, the memory 304 represents the entire virtual memory of the computer system 300 and may include the virtual memory of other computer systems connected to the computer system 300 via a network. The memory 304 may conceptually be regarded as a single entity, but in other embodiments, this memory 304 may have a more complex configuration, such as a hierarchy of caches and other memory devices. For example, the memory may exist as multiple levels of caches, and these caches may be divided by function. As a result, one cache may hold instructions, and another cache may hold non-instruction data used by the processor. The memory may be distributed and associated with various different CPUs, such as in a so-called NUMA (Non-Uniform Memory Access) computer architecture.

[0019] Memory 304 may store all or part of the programs, modules, and data structures that implement the functions described herein. For example, memory 304 may store the reading method estimation application 350. In certain embodiments, the reading method estimation application 350 may include instructions or descriptions that execute the functions described below on the processor 302, or may include instructions or descriptions that are interpreted by other instructions or descriptions. In certain embodiments, the reading method estimation application 350 may be implemented in hardware via semiconductor devices, chips, logic gates, circuits, circuit cards, and / or other physical hardware devices instead of or in addition to a processor-based system. In certain embodiments, the reading method estimation application 350 may include data other than instructions or descriptions. In certain embodiments, a camera, sensor, or other data input device (not shown) may be provided to communicate directly with the bus interface unit 309, the processor 302, or other hardware of the computer system 300. In such a configuration, the need for the processor 302 to access the memory 304 and the latent factor identification application may be reduced.

[0020] Computer system 300 may include a bus interface unit 309 that facilitates communication between a processor 302, a memory 304, a display system 324, and an I / O bus interface unit 310. The I / O bus interface unit 310 may be coupled to an I / O bus 308 for transferring data between various I / O units. The I / O bus interface unit 310 may communicate via the I / O bus 308 with a plurality of I / O interface units 312, 314, 316, and 318, also known as I / O processors (IOPs) or I / O adapters (IOAs). The display system 324 may include a display controller, a display memory, or both. The display controller may be capable of providing video, audio, or both data to a display device 326. Additionally, computer system 300 may include devices such as one or more sensors configured to collect data and provide the data to processor 302. The display memory may be dedicated memory for buffering video data. The display system 324 may be connected to a display device 326 such as a single display screen, a television, a tablet, or a portable device. In some embodiments, the display device 326 may include speakers for rendering audio. Alternatively, speakers for rendering audio may be connected to an I / O interface unit. In other embodiments, the functionality provided by the display system 324 may be implemented by an integrated circuit including processor 302. Similarly, the functionality provided by the bus interface unit 309 may be implemented by an integrated circuit including processor 302.

[0021] The I / O interface unit is equipped with the function of communicating with various storage or I / O devices. For example, the terminal interface unit 312 can be attached with user I / O devices 320 such as user output devices like video display devices, speaker TVs, etc., and user input devices like keyboards, mice, keypads, touch pads, trackballs, buttons, light pens, or other pointing devices. The user can use the user interface to operate the user input device to input input data and instructions to the user I / O device 320 and the computer system 300, and may also receive output data from the computer system 300. The user interface may be displayed on a display device, played by a speaker, or printed via a printer, for example, via the user I / O device 320.

[0022] The storage interface 314 can be attached with one or more disk drives and direct access storage devices 322 (usually magnetic disk drive storage devices, but may also be an array of disk drives configured to appear as a single disk drive or other storage devices). In certain embodiments, the storage device 322 may be implemented as any secondary storage device. The content of the memory 304 may be stored in the storage device 322 and read from the storage device 322 as needed. The I / O device interface 316 may provide an interface to other I / O devices such as printers, fax machines, etc. The network interface 318 may provide a communication path so that the computer system 300 and other devices can communicate with each other. This communication path may be, for example, the network 330.

[0023] The computer system 300 shown in FIG. 1 includes a bus structure that provides direct communication paths among a processor 302, a memory 304, a bus interface 309, a display system 324, and an I / O bus interface unit 310. However, in other embodiments, the computer system 300 may include point-to-point links in a hierarchical, star, or web configuration, multiple hierarchical buses, parallel or redundant communication paths. Further, although the I / O bus interface unit 310 and the I / O bus 308 are shown as a single unit, in practice, the computer system 300 may include multiple I / O bus interface units 310 or multiple I / O buses 308. Also, multiple I / O interface units are shown for separating the I / O bus 308 from various communication paths connecting to various I / O devices. However, in other embodiments, some or all of the I / O devices may be directly connected to one system I / O bus.

[0024] In certain embodiments, the computer system 300 may be a device that receives requests from other computer systems (clients) without a direct user interface, such as a multi-user mainframe computer system, a single-user system, or a server computer. In other embodiments, the computer system 300 may be a desktop computer, a portable computer, a notebook computer, a tablet computer, a pocket computer, a telephone, a smartphone, or any other suitable electronic device.

[0025] Next, with reference to FIG. 2, the basic configuration of the reading method estimation device 100 for character trademarks will be described. The reading method estimation device 100 of this embodiment includes a character string splitter 101, a converter 102, and a semantic and intonational chunk splitter 103.

[0026] The string splitter 101 splits the input string into characters. The strings that are expected as input include alphabets (including alphabets with diacritical marks such as French accents and German umlauts), hiragana, katakana, Chinese characters (including Chinese characters used in Japanese, as well as simplified and traditional Chinese characters), numbers, other character types, and combinations thereof.

[0027] The converter 102 includes a learning model learned by the learner 200 based on the collation and normalization data. The converter 102 generates the reading of the trademark string split into characters using the learned model. In this embodiment, since a neural machine translation model called a transformer composed of an encoder and a decoder with an attention mechanism is used, the learning of the conversion model and the generation of the appellation by the model in this embodiment are performed in the same process as the process in which neural machine translation generates a translated sentence.

[0028] Next, with reference to FIG. 3, the preprocessing and learning method of the learning data of the converter 102 will be described.

[0029] Similar to the learning of the neural machine translation model based on the set of the original text and the translated text, what is necessary for the learning in this embodiment is a corpus having a large amount of sets of stringified trademarks and appellations. This corpus can be generated from the collation and normalization data as shown in FIG. 2.

[0030] As shown in FIG. 3, in the collation and normalization data, for character trademarks, "display trademark" and "appellation" are stored respectively. The display trademark is the result of transcribing the string included in the trademark containing characters, combined with special symbols representing its form, etc.

[0031] The pre-processing of the organized and standardized data will be explained with reference to Figure 3. First, in step 1, pairs of display trademarks and pronunciations are extracted from the organized and standardized data. Display trademarks frequently contain character strings that should not be pronounced, such as special symbols that indicate the form of the trademark (§, ∞, ▲, ▼, ¢, \), but as will be described later, there is no problem if these are not deleted.

[0032] In the organized standardized data, multiple pronunciations are stored for one display trademark. Here, only the first pronunciation of the multiple pronunciations is used for learning. This is because if all stored pronunciations were to be used as learning data, multiple different readings would be defined for one display trademark, which would prevent the neural network from learning properly. For example, in the case of the display trademark shown in this figure, "Japan Patent Information Organization" (registered trademark), four corresponding pronunciations are stored, but only the first pronunciation "Nippon Tokkyo Jo Hokkiko" is used for learning. In addition to organized standardized data, data on bibliographic and historical information such as standard patent information data can be used.

[0033] Next, in step 2, the extracted two are tokenized (divided into components) on a character-by-character basis. For example, the display trademark "Japan Patent Information Organization" is tokenized as "Japan Patent Information Organization," and the pronunciation "Nippon Tokyo Jo Ho Ki Ko" is tokenized as "Nippon Tokyo Jo Ho Ki Ko." Character-by-character tokenization is not effective in the field of natural language processing, including machine translation, which aims to grasp meanings and concepts, and is therefore rarely adopted in any task. However, for character strings such as trademarks, which contain a wide variety of character types, character-by-character tokenization is the most effective tokenization method. This is because, in character strings that contain a wide variety of character types, it is difficult to determine the so-called morpheme divisions, or creating a dictionary to define morpheme divisions may be inappropriate.

[0034] Next, in step 3, the character-by-character string is converted into a dictionary ID string. For example, "Japan Patent Information Organization" is converted into "142 153 482 991 3416 389 379 900," and "Nippon Tokkyo Joho Ki Ko" is converted into "38 8 60 2 7 8 21 32 16 32 1 63 1 21 18 1." The dictionary used here is created on the input side of the neural network by counting all characters contained in the display trademark without duplication. The dictionary is created on the output side of the neural network in the same way, but the characters contained in the output dictionary are all katakana characters, including lowercase characters. The number of registered items for both is 4,098 and 84, respectively.

[0035] Here, the number of parameters of a neural network for machine translation, such as a transformer, is proportional to the square of the number of items registered in the dictionary. The number of dictionary items for a neural network used in normal machine translation is about 10,000, so the size of the neural network used in the present invention is much slimmer than those neural networks.

[0036] Finally, learning is performed in step 4. Learning by the learning device 200 is performed in the same manner as in neural machine translation using a normal transformer. However, unlike normal neural machine translation in which the input unit is a morpheme or a subword obtained by further dividing a morpheme, the input unit of the present invention is a character. Therefore, when a corpus is generated from organized standardized data, the input character type is assumed to be alphabet, hiragana, katakana, kanji, special characters, etc., and the output character type is assumed to be katakana.

[0037] As mentioned above, the neural network of the present invention is extremely slim compared to neural networks used for machine translation, so if training is performed under the same conditions as when training a machine translator, training can be completed in about one day.

[0038] In the present invention, a Transformer, which has advantages over other neural networks in terms of both accuracy and calculation speed, is adopted as the conversion model. However, other neural networks used in neural machine translation, such as LSTM (Long Short-Term Memory) and RNN (Recurrent Neural Network), may also be used.

[0039] In addition, as the corpus, in addition to data related to bibliographic and progress information such as organized and standardized data and patent information standard data, data in which strings such as a Japanese dictionary and a foreign language dictionary are set with the Japanese reading may also be used.

[0040] Next, with reference to FIG. 4, the process of generating a title by the converter 102 will be described. The converter 102 estimates a title using the learning model obtained by the learning of the learner 200.

[0041] First, the converter 102 generates the first character of the title from the information of the divided trademark string. Here, the character type constituting the title is katakana in accordance with the organized and standardized data, but other syllabic characters starting with hiragana may also be used. In FIG. 4, "JAPIO" (registered trademark) is illustrated as the input string, but the character generated in the first title generation step (1) is "ji".

[0042] Thereafter, based on the already generated katakana or hiragana character name sequence and the information from the encoder, the converter 102 continues to sequentially generate the next character type (katakana) by the decoder until a token indicating the end is output. In this specification, "sequentially generate" means generating one character at a time following the order of the reading string. The converter 102 generates a katakana or hiragana character that becomes the first character of the reading from the character string divided into character units, and then, based on the one or more already generated katakana or hiragana characters, sequentially generates a katakana or hiragana character that is the next character to appear after the already generated one or more katakana or hiragana characters, thereby converting the character string divided into character units into a reading consisting of katakana or hiragana.

[0043] In the name generation step (2) of FIG. 4, "ya" as the next name character is generated from the input "ji", and in the name generation step (3) of FIG. 4, "pi" as the next name character is generated from the input "ja", and so on.

[0044] Examples of the names estimated by the reading estimation device 100 are shown in Table 1. Since the division unit in the string splitter 101 is a character and the neural network of the converter 102 is a transformer, it can be understood that it is possible to handle all languages. Among the character trademarks in Table 1, "Japan Patent Information Organization" and "JAPIO" are registered trademarks.

[0045]

Table 1

[0046] The present invention is capable of supporting multiple languages because the unit of tokenization is set to characters and a transformer is adopted for the neural network. When the unit of tokenization is set to words or word pieces, it becomes difficult to perform conversion in accordance with the language assumed by the character string creator when the words or word pieces after tokenization are associated with items in a dictionary different from the language assumed by the character string creator. On the other hand, when the unit of tokenization is set to characters, such inconvenience does not occur. It will be easily understood that the learning model learned from a set of trademarks composed of characters tokenized in this way and appellations corresponding to the Japanese readings is not limited in character types. For example, even if a character trademark is composed of Cyrillic characters, Hangul characters, Arabic characters, Thai characters, Devanagari characters, etc. of a minority language, if it has a learning model learned from a set of a plurality of consecutive character strings and the appellations corresponding to these character strings, the reading estimation device 100 can estimate the appellations even for these other character types. Further, in the present embodiment, the reading estimation device 100 for character trademarks is described, but it will be easily understood that the reading estimation device according to the present invention is not limited to character trademarks and can also estimate appellations such as personal names, place names, company / group names, Chinese characters, and proper nouns in foreign languages. That is, the reading of these personal names and the like can be estimated by a learning model learned from a set of personal names, place names, company / group names, Chinese characters, proper nouns in foreign languages, etc. composed of characters to be tokenized and appellations corresponding to the Japanese readings.

[0047] In addition, a neural network adopting an attention mechanism such as a transformer can refer to all the information on the encoder side when generating tokens on the decoder side, and conversion based on the features (such as language) indicated by the order of the entire character string becomes possible. Note that the memory gate of LSTM can also achieve a similar effect, although the performance is inferior.

[0048] When accuracy verification was performed using the organized and standardized data that was not used as learning data, the accuracy was 99.3%. Furthermore, since the pairs of trademarks for display and appellations are used as learning data, as shown in Table 2, effects such as excluding character strings with no discriminative power such as "General Incorporated Foundation", "Corporation", and "Limited Liability Company" from the target of appellation generation can also be achieved. Also, when generating the appellation for the word mark "JAPIO じゃぴお", it is possible to combine the two overlapping appellations of "ジャピオ" into one appellation. Among the word marks in Table 2, "Japan Patent Information Organization, General Incorporated Foundation" is a registered trademark.

[0049]

Table 2

[0050] The reason such appellation generation is possible is that the learning data generated in each step of Figure 3 is learned using a transformer. Regarding the translation process by a transformer, compared with conventional statistical machine translation, technical issues such as translation omissions caused by translation omissions in the learning data have been reported. However, in cases where such "conversion omissions" must be deliberately caused, this technical drawback of the transformer can instead become an advantage.

[0051] For the same reason, control characters also disappear due to conversion, so there is no need to specially delete them in the preprocessing. This is made possible by a mechanism called the attention mechanism of the transformer. Although the performance is inferior, the same effect can also be achieved with the memory gate of LSTM.

[0052] Next, with reference to FIG. 5, the operation of the semantic and intonation chunk splitter 103 will be described. When the semantic and intonation chunk splitter 103 determines that it is possible to delimit the appellation generated by the converter 102 from the aspects such as semantics or intonation, it delimits and outputs the appellation. The determination of whether delimitation is possible can be made based on dictionary data or the like that registers semantic and intonation chunks frequently occurring in the appellation. For example, "Japan Patent Information Organization" is converted by the converter 102 into "Nippon Tokkyo Johou Kiko", but the semantic and intonation chunk splitter 103 further delimits and outputs it as "Nippon Tokkyo Johou Kiko".

[0053] Generally, when a character trademark is redundant, a part of the semantic and intonation unit in the entire appellation of the trademark may be established as the appellation. By the division by the semantic and intonation chunk splitter 103, it becomes possible to generate the appellation of a redundant character trademark.

[0054] Note that the above is also a reading estimation method for estimating the reading of a character string from the character string. This reading estimation method includes a step of dividing the input character string into characters, and a step of converting the character string divided into characters into a reading consisting of katakana characters or hiragana characters by sequentially generating one katakana character or hiragana character at a time from the character string divided into characters by a neural network. Further, the present invention includes a program for executing each step in the reading estimation method and a computer-readable recording medium on which the program is recorded. Thereby, the reading estimation apparatus and method according to the present invention can estimate the readings of various character strings, particularly the appellations of trademarks, without creating a special dictionary database regarding the readings of character strings.

Example

[0055] The reading method estimation device 100 includes a character string splitter 101 and a converter 102. Note that the same components as those in the first embodiment are denoted by the same reference numerals, and the description thereof is omitted. The character string splitter 101 receives the character string split by the character type splitter 400. The splitting of the appellation may be realized by preprocessing. When performing the splitting in the preprocessing, first, the character trademark is split at the position where the character type switches. Specifically, as shown in FIG. 6, first, when raising the split appellation of the character trademark "JAPIO TOP" (registered trademark), the character trademark is split by the character type splitter 400 at the position where the character type switches from alphabet to Chinese characters, and the character trademark "JAPIO TOP" is split into two character trademarks: (1) "JAPIO" and (2) "TOP".

[0056] In the second embodiment, for each of the split character trademarks, the appellation is raised. In the case of FIG. 6, two appellations, (1) "JAPIO" and (2) "ITADAKI", are generated respectively.

[0057] When viewed from the input character trademark side, the above-mentioned semantic and intonation chunks are often defined with the position where the character type switches as a delimiter in a language in which a plurality of character types are mixed like Japanese. By paying attention to this characteristic and performing preprocessing, it becomes possible to perform redundant appellation raising of character trademarks without creating dictionary data or the like in which semantic and intonation chunks are registered, as in the first embodiment.

Explanation of reference numerals

[0058] 100 Reading method estimation device 101 Character string splitter 102 Converter 103 Semantic and intonation chunk splitter 200 Learner 300 Computer system 302 Processor 302A, 302B General-purpose programmable central processing unit (CPU) 304 Memory 306 Memory bus 308 I / O bus 309 Bus IF 310 I / O Bus IF 312 Terminal Interface 314 Storage Interface 316 I / O Device Interface 318 Network Interface 320 User IO Device 322 Storage Device 324 Display System 326 Display Device 330 Network 350 Reading Estimation Application 400 Character Type Splitter

Claims

1. A reading estimation device for estimating the reading of a given string, comprising: A string splitter for splitting the input string into characters; A converter for converting the string split into characters into a reading consisting of katakana or hiragana by sequentially generating one katakana or hiragana character at a time from the string split into characters by a neural network.

2. The string consists of alphabets, hiragana, katakana, kanji, numbers, other character types, and combinations thereof. The reading estimation device according to claim 1.

3. The string is a character trademark, and the reading is the name of the character trademark. The reading estimation device according to claim 1 or 2.

4. The neural network is trained with the display trademark in the trademark-related data split into characters and the name in the trademark-related data split into characters. The reading estimation device according to any one of claims 1 to 3.

5. The neural network is a transformer. The reading estimation device according to any one of claims 1 to 4.

6. comprising a semantic and intonational chunk splitter for splitting the reading into semantic and intonational chunks. The reading estimation device according to any one of claims 1 to 5.

7. comprising a character type splitter for splitting the input string at locations where the character type changes, and the string splitter and the converter perform processing on a plurality of strings split by the character type splitter at locations where the character type changes. The reading method estimation device according to any one of claims 1 to 6.

8. A reading method estimation method for estimating the reading method of a character string, which is executed by a computer including a memory for storing a program having the following steps and a processor for reading and executing the program from the memory, a step of splitting the input character string into characters; a step of converting the character string split into characters into a reading consisting of katakana or hiragana by sequentially generating one katakana character or one hiragana character at a time from the character string split into characters by a neural network.

9. The character string consists of alphabets, hiragana, katakana, kanji, numerals, other character types, and combinations thereof. The reading method estimation method according to claim 8.

10. The character string is a character trademark, and the reading is the name of the character trademark. The reading method estimation method according to claim 8 or 9.

11. The neural network is trained with the display trademarks in the trademark-related data split into characters and the names in the trademark-related data split into characters. The reading method estimation method according to any one of claims 8 to 10.

12. The neural network is a transformer. The reading method estimation method according to any one of claims 8 to 11.

13. a step of splitting the reading into semantic and intonational chunks. The reading method estimation method according to any one of claims 8 to 12.

14. A step of splitting the input string at positions where the character type changes. In the step of splitting the input string into characters and the step of converting the characters into a reading consisting of katakana or hiragana, processing is performed on a plurality of strings split at positions where the character type changes. The reading estimation method according to any one of claims 8 to 13.

15. A program for causing each step in the reading estimation method according to any one of claims 8 to 14 to be executed.

16. A computer-readable recording medium recording a program for causing each step in the reading estimation method according to any one of claims 8 to 14 to be executed.

Citation Information

Patent Citations

  • Reading information determination method, device, and program

    JP2004206659A

  • Transliteration device, transliteration program, computer-readable recording medium in which transliteration program is recorded and method of transliteration

    JP2012185679A