Information Processing Method, Information Processing Apparatus, and Program

By concatenating word sequences with classification labels to create input data, a single language model can be learned efficiently, addressing the challenge of representing multiple classifications and reducing processing costs.

JP7690876B2Active Publication Date: 2025-06-11RICOH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2021202764
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-12-14
Publication Date
2025-06-11
Estimated Expiration
2041-12-14

AI Technical Summary

Technical Problem

Conventional technologies face challenges in learning a single language model that can effectively represent word sequences belonging to multiple classifications, and the time and processing cost increase significantly as the number of classifications grows.

Method used

An information processing method where an apparatus inputs classification labels for word sequences, creates input data by concatenating the word sequences and classification labels, and learns a single language model capable of calculating word appearance probabilities using this data.

Benefits of technology

This approach allows for the efficient learning of a single language model that can handle word sequences across multiple classifications, reducing the time and processing costs associated with learning separate models for each classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007690876000004
    Figure 0007690876000004
  • Figure 0007690876000005
    Figure 0007690876000005
  • Figure 0007690876000006
    Figure 0007690876000006
Patent Text Reader

Abstract

To provide an information processing method, an information processing device, and a program for learning a single language model that corresponds to a word string belonging to a plurality of categories.SOLUTION: In an information processing system including an information processing device and a terminal device connected to a communication network such as an office LAN, an information processing device 5 includes a word string input unit 10 that receives an input of a word string, a category label input unit 11 that receives an input of at least one category label to which the word string belongs, a connection unit 12 that creates input data obtained by connecting the word string and the category label, a learning unit 13 that, by using a plurality of the input data sets, learns a single language model capable of calculating the appearance probability of a word in a word string belonging to a classification label, and a language model storage unit 4 that stores a parameter of the learned language model.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing method, an information processing apparatus, and a program.

Background Art

[0002] Conventionally, there are techniques for calculating vectors (word vectors) that represent the meanings of words constituting a word sequence using a language model learned using a word sequence, and techniques for performing text generation based on the appearance probability of words. Patent Document 1 discloses a technique for learning a language model used to appropriately predict the next word using information on a long context.

Summary of the Invention

Problems to be Solved by the Invention

[0003] However, in the conventional technology, it has not been possible to learn a single language model corresponding to word sequences belonging to a plurality of classifications. Further, when constructing and learning a language model for each classification to which a word sequence belongs, there has been a problem that the time and processing cost required for learning increase as the number of classifications increases.

[0004] An embodiment of the present invention aims to learn a single language model corresponding to word sequences belonging to a plurality of classifications in view of the above problems.

Means for Solving the Problems

[0005] In order to solve the above-described problems, the present invention is an information processing method by an information processing apparatus, wherein the information processing apparatus inputs at least one classification label to which a word sequence belongs, creates input data obtained by concatenating the word sequence and the classification label, and learns a single language model capable of calculating the appearance probability of words in the word sequence belonging to the classification label using a plurality of the input data, and stores parameters of the learned language model.

Effects of the Invention

[0006] According to an embodiment of the present invention, a single language model corresponding to word sequences belonging to a plurality of classifications can be learned.

Brief Description of Drawings

[0007]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Embodiments for Carrying Out the Invention

[0008] Hereinafter, embodiments of an information processing method, an information processing apparatus, and a program according to the present invention will be described in detail with reference to the accompanying drawings.

[0009] [First Embodiment] <System Outline> FIG. 1 is a diagram showing an example of a schematic diagram of an information processing system according to an embodiment of the present invention. The information processing system 1 includes, for example, a terminal device 3 and an information processing device 5 connected to a communication network 2 such as an in-house LAN (Local Area Network). The terminal device 3 transmits, via the communication network 2, a plurality of word sequences necessary for learning a language model and a classification label indicating the classification (field) to which the word sequences belong. Examples of the classification to which the word sequences belong include IPC (International Patent Classification), which is a classification of patent documents. In the prior art, it was necessary to learn hundreds to thousands of language models for word sequences classified by IPC. The information processing device 5 learns a single language model for the word sequences and classification labels received from the terminal device 3, and the language model storage unit 4 stores the parameters of the learned language model. Further, the information processing device 5 can calculate a word vector, a classification vector, and the appearance probability of a word for each of the word sequences and classification labels using the learned language model.

[0010] Note that the system configuration of the information processing system 1 shown in FIG. 1 is an example. For example, although the number of terminal devices 3 included in the information processing system 1 is set to one, it may be any number. Further, the communication network 2 may include a connection section by wireless communication such as mobile communication or wireless LAN.

[0011] Further, although the information processing device 5 is assumed to receive data necessary for learning a language model from the terminal device 3, it is not limited to this. The information processing device 5 may receive data input from an input device or the like included in the information processing device 5. In that case, it may not communicate with the terminal device 3.

[0012] Further, although the language model storage unit 4 is configured as a component included in the information processing device 5, it is not limited to this. The language model storage unit 4 may be provided outside the information processing device 5. Further, the information processing device 5 may be realized by a plurality of information processing devices. In other words, the information processing device 5 may include a plurality of information processing devices.

[0013] <Hardware Configuration Example> FIG. 2 is a diagram showing an example of the hardware configuration of the terminal device 3 and the information processing device 5 according to an embodiment of the present invention. As shown in FIG. 2, the terminal device 3 and the information processing device 5 are constructed by a computer, and include a CPU 501, a ROM 502, a RAM 503, an HD (Hard Disk) 504, an HDD (Hard Disk Drive) controller 505, a display 506, an external device connection I / F (Interface) 508, a network I / F 509, a bus line 510, a keyboard 511, a pointing device 512, a DVD-RW (Digital Versatile Disk Rewritable) drive 514, and a media I / F 516.

[0014] Among these, the CPU 501 controls the operations of the entire terminal device 3 and the information processing device 5. The ROM 502 stores programs used for driving the CPU 501 such as the IPL. The RAM 503 is used as a work area for the CPU 501. The HD 504 stores various data such as programs. The HDD controller 505 controls the reading or writing of various data to and from the HD 504 according to the control of the CPU 501. The display 506 displays various information such as a cursor, a menu, a window, characters, or an image. The external device connection I / F 508 is an interface for connecting various external devices. The external devices in this case are, for example, a USB (Universal Serial Bus) memory, a printer, and the like. The network I / F 509 is an interface for performing data communication using the communication network 2. The bus line 510 is an address bus, a data bus, etc. for electrically connecting the components such as the CPU 501 shown in FIG. 2.

[0015] In addition, the keyboard 511 is a type of input means having a plurality of keys used for inputting characters, numerical values, various instructions, and the like. The pointing device 512 is a type of input means for selecting and executing various instructions, selecting a processing target, moving a cursor, and the like. The DVD-RW drive 514 controls reading or writing of various data with respect to the DVD-RW 513 as an example of a removable recording medium. Note that the DVD-RW drive 514 is not limited to the DVD-RW, and may be a DVD-R or the like. The media I / F 516 controls reading or writing (storage) of data with respect to the recording medium 515 such as a flash memory.

[0016] <Regarding functions> FIG. 3 is a diagram showing an example of a configuration diagram of functional blocks in the information processing apparatus 5 according to the first embodiment of the present invention. The information processing apparatus 5 includes a word sequence input unit 10, a classification label input unit 11, a concatenation unit 12, a learning unit 13, a vector calculation unit 14, and a probability calculation unit 15. Each of these units is a function or means realized by the CPU 501 executing instructions included in one or more programs installed in the information processing apparatus 5. The language model storage unit 4 can be realized by a storage device such as the HD 504 included in the information processing apparatus 5, for example.

[0017] The word sequence input unit 10 receives an input of a word sequence from the terminal device 3 or the like, and outputs the input word sequence to the concatenation unit 12.

[0018] The classification label input unit 11 receives an input of a classification label from the terminal device 3 or the like, and inputs the input classification label to the concatenation unit 12.

[0019] The concatenation unit 12 creates input data to be input to the language model when learning the language model by concatenating the received word sequence and the classification label, and outputs the created input data to the learning unit 13. Here, when the word sequence belongs to a plurality of classifications, the concatenation unit 12 creates input data by concatenating a plurality of classification labels corresponding to the word sequence.

[0020] The learning unit 13 uses the input data received from the connection unit 12 to perform learning of the language model, and transmits the parameters of the learned language model to the language model storage unit 4.

[0021] The vector calculation unit 14 performs vector calculation according to an instruction from the learning unit 13 during the learning of the language model. Details of the calculation method will be described later.

[0022] The probability calculation unit 15 calculates the appearance probability of words according to an instruction from the learning unit 13 during the learning of the language model. Details of the calculation method will be described later.

[0023] The language model storage unit 4 stores the parameters of the language model received from the learning unit 13, or transmits the stored parameters of the language model to the learning unit 13.

[0024] <Language Model Learning Process> Subsequently, the processing in the learning of the language model will be described using a flowchart. FIG. 4 is a diagram showing an example of a flowchart related to the processing in the learning of the language model according to the first embodiment of the present invention. Hereinafter, the processing of each step in FIG. 4 will be described.

[0025] Step S20: The word sequence input unit 10 and the classification label input unit 11 of the information processing apparatus 5 receive the input of the word sequence and the classification label from the terminal device 3 and the like, respectively, and output the input word sequence and classification label to the connection unit 12.

[0026] Step S21: The connection unit 12 of the information processing apparatus 5 creates input data to be input to the language model in the learning of the language model by connecting the received word sequence and the classification label. FIG. 5 is a diagram showing an example of a word sequence, a classification label, and input data according to the first embodiment of the present invention. In FIG. 5, the word sequence 50 has four words, "clothes", "for", "ink", and "jet", and the classification label 51 has two types of classification labels, "B41J2" and "D06P5". The input data 52 is created by arranging all the classification labels 51 from left to right and then continuing to arrange the word sequence 50. Here, a maximum value is determined with respect to the number of elements of the word sequence 50 and the classification label 51. When the number of elements of the input word sequence 50 and the classification label 51 is less than the maximum value, data indicating that there is no data or the like may be inserted. Here, the number of elements is the number of words included in the word sequence 50 and the number of classification labels included in the classification label 51. For example, when the maximum value of the number of elements of the classification label 51 is 3, in the input data 52, after "D06P5", data indicating that there is no data (for example, "blank") or a predetermined value is added as an element of the classification label 51, and then the data of "clothes" in the word sequence 50 is connected.

[0027] FIG. 6 is a diagram showing an example of the configuration of a language model according to the first embodiment of the present invention. As shown in FIG. 6, the language model is, as an example, composed of three layers: a first layer, a second layer, and a third layer. The learning unit 13 calculates a vector (embedding vector) initialized with a random value for each element included in the input data 52 and inputs it to the first layer of the language model. Here, the dimensionality of the embedding vectors for each word and classification label is made the same. Also, for the embedding vectors corresponding to the word sequence 50 and the classification label 51, with respect to the language model shown in FIG. 6, starting from the first (leftmost) column, first the embedding vector of the classification label 51 is input, and then the embedding vector of the word sequence 50 is input. That is, in the language model of FIG. 6, two columns from the first (leftmost) column correspond to the embedding vector of the classification label 51, and the next four columns correspond to the embedding vector of the word sequence 50. Alternatively, as described with reference to FIG. 5, a maximum value is determined with respect to the number of elements of the word sequence 50 and the classification label 51, and when the number of elements of the input word sequence 50 and classification label 51 is less than the maximum value, data indicating no data or a predetermined value may be inserted to determine the number of columns of the language model. Also, in the learning of the language model, the calculation of vectors and occurrence probabilities for classification labels is performed in the same manner without distinction from the calculation for words.

[0028] Furthermore, the learning unit 13 replaces a part of a word with a placeholder represented by "[MASK]" and calculates an embedding vector in the same manner as for a word and inputs it to the first layer. In FIG. 6, the word "clothes" is replaced with "[MASK]".

[0029] The second and third layers of the language model are composed of a word vector corresponding to the word input in the first layer and a classification vector corresponding to the classification label, which are denoted as "Vec" in FIG. 6. The word vector and the classification vector may also be collectively simply called vectors.

[0030] At the uppermost stage of the language model, the occurrence probability for the word or classification label corresponding to the position of that column is output. Returning to FIG. 4 for explanation.

[0031] Step S22: The learning unit 13 of the information processing apparatus 5 calculates the vector and the appearance probability of the word in order to perform the learning of the language model. First, the learning unit 13 causes the vector calculation unit 14 to calculate the vectors of the second layer and the third layer shown in FIG. 6. The vector calculation unit 14 calculates the vector using the following mathematical formula 1.

[0032]

Equation

[0033] Next, the learning unit 13 causes the probability calculation unit 15 to calculate the appearance probability of the word. The probability calculation unit 15 calculates the appearance probability Pk of the k-th word from the head (left side) in the third layer of the language model in FIG. 6 using the following mathematical formula 2. Here, it is assumed that k is the position of the stakeholder (k = 3).

[0034]

Equation

[0035] Also, the Softmax function is expressed as an equation that satisfies the following mathematical formula 3 with the elements of the vectors y and x being y i and x i respectively.

[0036]

Equation

[0037] Step S24: The learning unit 13 of the information processing apparatus 5 executes the process of step S25 for all the input data 52 for learning the language model. That is, if the learning unit 13 has not executed the process of step 23 for all the input data 52, the process is transitioned to step S20, and if not, the process is transitioned to step S25.

[0038] Step S25: The learning unit 13 of the information processing apparatus 5 transmits the parameters of the language model learned in the processes of steps S20 to S24 to the language model storage unit 4. The language model storage unit 4 stores the received parameters of the language model.

[0039] Through the above processing, the information processing apparatus 5 can create input data by concatenating the input word sequence and the classification label indicating the classification to which the word sequence belongs, and can perform learning of a single language model using the created input data. That is, by creating and using input data that includes not only the word sequence but also the classification label indicating the classification to which the word sequence belongs, it is possible to learn a single language model corresponding to word sequences belonging to a plurality of classifications without learning a language model for each classification to which the word sequence belongs. Further, in the information processing apparatus 5, it is also possible to set the maximum number of elements of each of the number of elements of the word sequence 50 and the classification label 51 in the input data to the language model and the number of elements of the embedding vector in the first layer of the corresponding language model to a predetermined maximum value. That is, first, a predetermined maximum value is set for the number of elements of the word sequence and the classification label in the input data, respectively. Then, when the number of elements is smaller than the maximum value, input data is created by adding elements having a predetermined value so that the number of elements becomes equal to the maximum value. The process of adding elements having a predetermined value may be performed by the concatenation unit 12, or may be performed by the word sequence input unit 10 or the classification label input unit 11. By this process, it becomes possible to learn a single language model corresponding to word sequences and classification labels having various numbers of elements within a predetermined maximum value.

[0040] [Second Embodiment] In the second embodiment, the information processing apparatus 5 calculates vectors (word vectors and classification vectors) using the language model learned in the first embodiment. FIG. 7 is a diagram showing an example of a configuration diagram of functional blocks in the information processing apparatus 5 according to the second embodiment of the present invention. In FIG. 7, the processing of the word sequence input unit 10 and the classification label input unit 11 is the same as the processing described with reference to FIG. 3. Regarding the concatenation unit 12, the method of creating the input data 52 is the same as the method described with reference to FIG. 3, but the difference is that the created input data 52 is output to the vector calculation unit 14 instead of the learning unit 13.

[0041] The vector calculation unit 14 calculates vectors (word vectors and classification vectors) by the same procedure as the method in step S22 of FIG. 4 and Equation 1 during the learning of the language model. Here, the parameters of the language model required for the calculation are received from the language model storage unit 4.

[0042] The language model storage unit 4 transmits the stored parameters of the language model to the vector calculation unit 14.

[0043] FIG. 8 is a diagram showing an example of a flowchart related to the processing according to the second embodiment of the present invention. Steps S30 and S31 execute the processing by the same procedure as steps S20 and S21 of FIG. 4.

[0044] Step S32: The vector calculation unit 14 of the information processing device 5 requests the parameters of the language model from the language model storage unit 4 and receives the parameters of the language model from the language model storage unit 4.

[0045] Step S33: The vector calculation unit 14 of the information processing device 5 calculates vectors (word vectors and classification vectors) by the same procedure as the method in step S22 of FIG. 4 and Equation 1 during the learning of the language model. Here, the vector calculation unit 14 may calculate only word vectors or only classification vectors. Further, the vector calculation unit 14 may transmit the calculated vectors to the terminal device 3, or display them on the display device of the information processing device 5, or store them in the storage device of the information processing device 5.

[0046] Through the above processing, the information processing device 5 can calculate vectors (word vectors and classification vectors) for word sequences and classification labels using the parameters of a single language model learned for word sequences belonging to a plurality of classifications. The word vectors and classification vectors can be used for analyzing the semantic similarity between words and the similarity between classifications, respectively. For example, the semantic similarity of words is analyzed based on the distance between word vectors. Alternatively, the similarity between two classifications (fields) is analyzed based on the distance between classification vectors.

[0047] [Third Embodiment] In the third embodiment, the information processing apparatus 5 calculates the appearance probability of words using the language model learned in the first embodiment. Here, regarding the word sequence 50 and the classification label 51 in the input data to the language model, the classification label 51 changed from the received classification label 51 is input. FIG. 9 is a diagram showing an example of the configuration diagram of the functional blocks in the information processing apparatus 5 according to the third embodiment of the present invention. In FIG. 9, the processing of the word sequence input unit 10 is the same as the processing described in FIG. 3. Regarding the classification label input unit 11, the classification label 51 obtained by changing the classification label 51 received from the terminal device 3 or the like is input to the connection unit 12. That is, the classification label input unit 11 inputs a classification label 51 different from the combination of the received word sequence 50 and classification label 51 to the connection unit 12. Regarding the connection unit 12, the method of creating the input data 52 is the same as the method described in FIG. 3, but the difference is that the created input data 52 is output to the probability calculation unit 15 instead of the learning unit 13.

[0048] The probability calculation unit 15 calculates the appearance probability Pk of words by the same procedure as the method at the time of learning the language model shown in step S22 and equations 2 and 3 of FIG. 4. Here, the parameters of the language model required for the calculation are received from the language model storage unit 4.

[0049] The language model storage unit 4 transmits the stored parameters of the language model to the probability calculation unit 15.

[0050] FIG. 10 is a diagram showing an example of a flowchart regarding the processing according to the third embodiment of the present invention.

[0051] Step S40: The word sequence input unit 10 and the classification label input unit 11 of the information processing apparatus 5 receive the input of the word sequence 50 and the classification label 51 from the terminal device 3 or the like, respectively. The word sequence input unit 10 outputs the input word sequence 50 to the concatenation unit 12. The classification label input unit 11 outputs the changed classification label 51 to the concatenation unit 12. The value of the classification label 51 to be changed may be specified, for example, from an input device of the terminal device 3 or the information processing apparatus 5.

[0052] Step S41: Execute the process in the same procedure as step S21 in FIG. 4.

[0053] Step S42: The probability calculation unit 15 of the information processing apparatus 5 requests the language model parameters from the language model storage unit 4 and receives the language model parameters from the language model storage unit 4.

[0054] Step S43: The probability calculation unit 15 of the information processing apparatus 5 calculates the appearance probability of the word by the same procedure as the method in the learning of the language model shown in step S22 and formulas 2 and 3 in FIG. 4. Further, the probability calculation unit 15 may transmit the calculated appearance probability of the word to the terminal device 3, or display it on the display device of the information processing apparatus 5, or store it in the storage device of the information processing apparatus 5.

[0055] Through the above processing, the information processing apparatus 5 can input the concatenated word sequence and the changed classification label as input data to a single language model learned for word sequences belonging to a plurality of classifications, and calculate the appearance probability of the word. Here, if the changed classification label is used in the learning of the language model and the learning of the language model is properly performed, the calculated appearance probability of the word corresponds to the appearance probability of the word in the changed classification label. That is, in the information processing apparatus 5, without using the language model learned for each classification, by inputting the specified classification label to a single language model, the appearance probability of the word in the specified classification label can be calculated.

[0056] As described above, several embodiments for carrying out the present invention have been explained. However, the present invention is not limited to such embodiments at all, and various modifications and substitutions can be made without departing from the gist of the present invention.

[0057] For example, an example of the configuration diagrams of the functional blocks shown in FIGS. 3, 7, and 9 is divided according to the main functions in order to facilitate the understanding of the processing by the information processing system 1 and the information processing apparatus 5. The present invention is not limited by the way of dividing the processing units or the names. The processing in the information processing system 1 and the information processing apparatus 5 can be further divided into more processing units according to the processing content. Also, one processing unit can be divided so as to include more processing.

[0058] Further, each function of the embodiments described above can be realized by one or a plurality of processing circuits. Here, the "processing circuit" in this specification means a processor programmed to execute each function by software like a processor implemented by an electronic circuit, an ASIC (Application Specific Integrated Circuit) designed to execute each function described above, a DSP (digital signal processor), an FPGA (field programmable gate array), or a device such as a conventional circuit module.

[0059] Also, the described device group only shows one of a plurality of computing environments for implementing the embodiments disclosed in this specification. In one embodiment, the information processing system 1 and the information processing apparatus 5 include a plurality of computing devices such as a server cluster. The plurality of computing devices are configured to communicate with each other via an arbitrary type of communication link including a network or a shared memory, and execute the processing disclosed in this specification.

[0060] Also, the word sequence input unit 10 and the classification label input unit 11 may simply be called input units, and the vector calculation unit 14 and the probability calculation unit 15 may simply be called calculation units.

Explanation of Signs

[0061] 1 Information processing system 2 Communication network 3 Terminal device 4 Language model storage unit 5 Information processing device 10 Word sequence input unit 11 Classification label input unit 12 Concatenation unit 13 Learning unit 14 Vector calculation unit 15 Probability calculation unit 50 Word sequence 51 Classification label

Prior Art Documents

Patent Documents

[0062]

Patent Document 1

Claims

1. An information processing method by an information processing apparatus, wherein the information processing apparatus inputs at least one classification label to which a word sequence belongs, creates input data by concatenating the word sequence and the classification label, learns a single language model capable of calculating the appearance probability of words in the word sequence belonging to the classification label using a plurality of the input data, stores parameters of the learned language model, the information processing method.

2. The information processing method according to claim 1, wherein in the language model, the classification label is treated equivalently to the words of the word sequence.

3. For the number of elements of the word sequence and the classification label in the input data, respectively set a predetermined maximum value, and when the number of elements is smaller than the maximum value, add elements having a predetermined value so that the number of elements becomes equal to the maximum value to create the input data, the information processing method according to claim 1.

4. Calculating a classification vector corresponding to the classification label using the learned language model, the information processing method according to any one of claims 1 to 3.

5. Calculating a word vector corresponding to the word sequence belonging to the classification label using the learned language model, the information processing method according to any one of claims 1 to 3.

6. By changing the classification label in the input data and inputting it to the learned language model, calculating the appearance probability of words in the changed classification label, the information processing method according to any one of claims 1 to 3.

7. Inputs at least one classification label to which a word sequence belongs, creates input data by concatenating the word sequence and the classification label, learns a single language model capable of calculating the appearance probability of words in the word sequence belonging to the classification label using a plurality of the input data, A program for causing an information processing apparatus to execute a process of storing parameters of the learned language model.

8. A classification label input unit that inputs at least one classification label to which a word sequence belongs, A concatenation unit that creates input data by concatenating the word sequence and the classification label, A learning unit that learns a single language model capable of calculating the appearance probability of words in the word sequence belonging to the classification label using a plurality of the input data, A language model storage unit that stores parameters of the learned language model, An information processing apparatus having

Citation Information

Patent Citations

  • Manufacture of forge-welded steel pipe

    JP1989095814A

  • Difficulty estimation learning device, difficulty estimation model learning device, and device, method and program for estimating difficulty

    JP2016152033A

  • Text generation method based on semantic expression, text generation device based on semantic expression, electronic equipment, non-temporary computer-readable storage medium and computer program

    JP2021117985A