Information processing device, information providing method, and program

The information processing device enhances creative work by generating and comparing vectors to find technical document data that can improve a specified index when combined with a specified technology, addressing the limitations of conventional systems.

JP7718056B2Active Publication Date: 2025-08-05RICOH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021010161
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-01-26
Publication Date
2025-08-05
Estimated Expiration
2041-01-26

AI Technical Summary

Technical Problem

Conventional information retrieval systems fail to support creative work by searching for technical document data that can increase a specified index by combining it with a specified technology from among multiple existing technical document data.

Method used

An information processing device that generates distributed representation vectors for technical and index document data, calculates similarity between composite vectors, and outputs information on document data that can enhance a specified index by combining with a specified technology.

Benefits of technology

Enables the search for technical document data that can enhance a specified index by combining with a specified technology, facilitating creative work and innovation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007718056000001
    Figure 0007718056000001
  • Figure 0007718056000002
    Figure 0007718056000002
  • Figure 0007718056000003
    Figure 0007718056000003
Patent Text Reader

Abstract

To provide an information processing apparatus configured to retrieve a technical document than can extend a designated index in combination with a designated technology from existing multiple technical documents.SOLUTION: An information processing apparatus includes: a receiving unit which receives settings for technical document data and index document data; a word embedding vector generation unit which generates word embedding vectors using multiple pieces of document data to be retrieved; a first vector generation unit which generates a first vector with distributed representation which represents the technical document data; a second vector generation unit which generates a second vector with distributed representation which represents the index document data; and an output unit which outputs information on one or more pieces of document data to be retrieved that can extend an index represented by the index document data in combination with a technology represented by the technical document data, out of the multiple pieces of document data to be retrieved, by using the word embedding vector, the first vector, and the second vector.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information providing method, and a program. [Background technology]

[0002] 2. Description of the Related Art There is an information retrieval system that searches through a plurality of document data for document data that meets a user's request and outputs the search results.

[0003] For example, an information search system is known that searches for books to recommend to a customer by calculating the similarity between a customer profile based on information about books purchased by the customer and a keyword vector for each book based on a book database (see, for example, Patent Document 1). Summary of the Invention [Problem to be solved by the invention]

[0004] In recent years, efforts to reform work styles have become more active, and various tools have been developed to make work more efficient. Most of these business improvement tools are designed to automate simple or repetitive tasks or to facilitate communication. While the aim is to use the time saved by using these tools to carry out work that creates new value, the current situation is that there are almost no tools that support creative work.

[0005] Innovation is described as a "new combination" brought about by new combinations of existing technologies. Therefore, there is a demand for support tools for creative work that present new value by combining, for example, a company's own technology with existing technology that is known to the public.

[0006] However, conventional information retrieval systems such as those shown in Cited Document 1 are unable to search for technical document data that can increase a specified index by combining it with a specified technology from among multiple existing technical document data.

[0007] One embodiment of the present invention has been made in consideration of the above-mentioned problems, and provides an information processing device that can search for technical documents from multiple existing technical documents that can increase a specified index by combining them with a specified technology. [Means for solving the problem]

[0008] In order to solve the above problem, an information processing apparatus according to an embodiment is provided. Setting operation , and index document data of setting Setting operation and a receiving unit that receives the document data to be searched using the document data. , a vector of distributed representations representing the plurality of target document data a word embedding vector generation unit that generates a word embedding vector; a first vector generation unit that generates a first vector of distributed representations that represent the technical document data; and a second vector generation unit that generates a second vector of distributed representations that represent the index document data; a calculation unit that calculates a similarity between the second vector and a plurality of composite vectors obtained by combining the first vector and each of the word embedding vectors; and a search result that includes information on the target document data corresponding to one or more composite vectors among the plurality of composite vectors whose similarity is higher than that of other composite vectors. and an output unit that outputs the signal. [Effects of the Invention]

[0009] According to one embodiment of the present invention, an information processing device can be provided that can search for technical document data from multiple existing technical document data that can increase a specified index by combining it with a specified technology. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a diagram illustrating an example of a system configuration of an information processing system according to an embodiment. [Figure 2] FIG. 1 is a diagram illustrating an overview of processing performed by an information processing device according to an embodiment. [Figure 3] FIG. 10 is a diagram illustrating vectorization of document data according to an embodiment. [Figure 4] FIG. 2 is a diagram illustrating an example of a hardware configuration of a computer according to an embodiment. [Figure 5]FIG. 1 is a diagram illustrating an example of a functional configuration of an information processing device according to an embodiment. [Figure 6] FIG. 1 is a sequence diagram (1) showing an example of processing by the information processing device according to the first embodiment. [Figure 7] FIG. 10 is a sequence diagram (2) illustrating an example of processing by the information processing device according to the first embodiment. [Figure 8] FIG. 11 is a sequence diagram showing an example of a process for acquiring target document data according to the second embodiment. [Figure 9] FIG. 11 is a sequence diagram illustrating an example of a process for generating word embedding vectors according to the third embodiment. [Figure 10] FIG. 13 is a sequence diagram illustrating an example of a process for generating word embedding vectors according to the fourth embodiment. [Figure 11] FIG. 13 is a sequence diagram showing an example of a process of downloading target document data according to the fifth embodiment. [Figure 12] FIG. 1 is a diagram (1) showing an example of a result display screen according to an embodiment. [Figure 13] FIG. 10 is a diagram (2) showing an example of a result display screen according to an embodiment. [Figure 14] FIG. 10 is a diagram illustrating an example of a setting screen according to an embodiment. [Figure 15] FIG. 10 is a diagram illustrating an example of a registration screen for target document data according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. <System configuration> Fig. 1 is a diagram showing an example of a system configuration of an information processing system according to an embodiment. The information processing system 1 includes, for example, an information processing device 100, an operation terminal 101, a local data server 102, and one or more external data servers 103. In the example of Fig. 1(A), the information processing device 100, the operation terminal 101, and the local data server 102 are communicably connected to each other via a local network 10 such as a LAN (Local Area Network). Furthermore, the one or more external data servers 103 are communicably connected to the information processing device 100 via, for example, an external network 11 such as the Internet and the local network 10.

[0012] The information processing device 100 is an information processing device having a computer configuration or a system including a plurality of computers, and provides the information providing service according to this embodiment to the user by executing a predetermined program.

[0013] The operation terminal 101 is, for example, an information terminal such as a PC (Personal Computer), a tablet terminal, or a smartphone, and a user can use, for example, the information providing service provided by the information processing device 100 using the operation terminal 101. However, this is not limited to this, and the user may, for example, operate the information processing device 100 to use the information providing service provided by the information processing device 100.

[0014] The local data server 102 is, for example, an information processing device having a computer configuration, or a system including multiple computers, and stores some or all of multiple document data to be searched (hereinafter referred to as target document data). As an example, the local data server 102 stores, among the multiple target document data, various document data related to technologies and the like that are owned by a company or the like.

[0015] The one or more external data servers 103 are, for example, information processing devices having a computer configuration or systems including multiple computers, and store some or all of the target document data to be searched. The external data server 103 stores, for example, a database of patent documents, a database of papers in various fields, or a database of technical publications published by other companies, etc.

[0016] The information processing device 100 can acquire target document data such as patent documents, papers, technical journals, etc. relating to various existing technologies from a local data server 102 and one or more external data servers 103.

[0017] 1(B), the information processing device 100 may be, for example, a server device connected to an external network 11, a cloud service, or the like. In this case, a user can use an information providing service provided by the information processing device 100 by accessing the information processing device 100 using an operation terminal 101.

[0018] (Processing Overview) 2 is a diagram for explaining an outline of processing by the information processing device. The information processing device 100 sets technical document data corresponding to a predetermined technology and index document data corresponding to a predetermined index, and provides information on an existing technology that can increase the index corresponding to the index document data by combining it with the technology corresponding to the technical document.

[0019] As a specific example, suppose that a department developing rearview monitors for automobiles in the 1990s and 2000s decided in its strategy review to increase (improve) the indicator of "ease of operation" for the "rearview monitor" technology. In this case, the user sets information on the operation screen (or setting screen) 200 provided by the information processing device 100, for example, as shown in FIG. 2(A), such as "rearview monitor" as the "technology to be increased" and "ease of operation" as the "indicator to be increased."

[0020] For example, it is assumed that the user inputs the character string "back monitor" into an input field 201 for a name corresponding to a technology on the operation screen 200, and sets a file name of technical document data related to the back monitor in a file setting field 202. It is also assumed that the user inputs the character string "ease of operation" into an input field 203 for a name corresponding to an indicator on the operation screen 200, and sets a file name of indicator document data explaining "ease of operation" in the file setting field 202. In this case, when the user selects the "Execute" button 205 on the operation screen 200, the information processing device 100 provides, for example, as shown in FIG. 2(B), information on a combination of an existing technology, "back monitor + multi-viewpoint image synthesis" 213, which can improve the indicator "ease of operation" 212 by combining it with the technology "back monitor" 211.

[0021] Preferably, the information processing device 100 provides one or more technical document data related to "multiple-viewpoint image synthesis" or information related to the technical document data as information related to "back monitor + multi-viewpoint image synthesis." This makes it easier for the user to come up with new technologies, such as an "around view monitor," from the technology "back monitor" that the user wants to develop and the technical document related to "multiple-viewpoint image synthesis" provided by the information processing device 100.

[0022] In this way, the information processing device 100 can provide information such as existing technologies that can improve the indicators set by the user by combining them with the company's own technology set by the user, as well as document data related to the technologies. Furthermore, the user can create new technologies that can improve the indicators set by combining the information provided by the information processing device 100 with the company's own technology. In other words, when the company's (company's) technology to be improved and the direction in which it should be improved are decided, the information processing system 1 according to this embodiment can present the user with an optimal solution to the problem of what existing technologies should be combined as elements.

[0023] (Regarding document data vectorization) For example, the information processing device 100 vectorizes the set technical document data and index document data, and calculates the similarity between the vectorized index and a composite vector of each of the vectorized technology and existing technology vectors. Furthermore, the information processing device 100 identifies an existing technology that can further improve the set index by combining it with the set technology by identifying a composite vector that has a higher similarity to the vectorized index.

[0024] In this embodiment, the method for vectorizing document data is not particularly limited, but here, as an example, an overview of vectorizing document data using a method such as Word2Vec or Doc2Vec will be described.

[0025] 3 is a diagram illustrating vectorization according to an embodiment. In this embodiment, generating vectors of distributed representations representing document data is referred to as vectorization of document data. As a method for vectorizing document data, for example, the Word2Vec method as shown in FIG. 3(A) can be applied.

[0026] For example, if the vocabulary size of a word is S, that word can be represented by an S-dimensional one-hot vector. However, the one-hot vectors representing these words are not distributed representation vectors, and the vectors are not related to each other at all. In addition, the cosine similarity between all vectors is 0, so all vectors are dissimilar.

[0027] Document data is not simply a collection of words; its meaning changes depending on the order in which the words appear. Natural language processing techniques such as Word2Vec improve accuracy by incorporating this word order as important information. For example, the words before and after a word contained in document data are extracted, and the last word is called the Target, and the set of words before that is called the Context. Furthermore, by learning an embedding matrix (weights) W for predicting the Target when a certain Context is given, each column vector of the word embedding matrix W can be used as a vector of the distributed representation of the word.

[0028] For example, the word embedding matrix W is defined as W=(w1w2 w i ··· w S )(w i is a column vector) and a certain context (the i-th word of the document data) as shown in Figure 3(A). In addition, when the one-hot vector of the i-th word is multiplied by the embedding matrix W, the i-th column vector wi = (w i,1 ,w i,2 ,···,w i,N )T (the hidden layer in Figure 3(A)). Here, the one-hot vector is a column vector in which only the i-th element is 1 and the other elements are 0.

[0029] Also, the intermediate layer column vector w i Then, we add the matrix W, which is the transpose of the embedding matrix W as shown in Figure 3(B). T By multiplying by this, the column vector of the output layer (output vector) p = (p1, p2, , p s )T is obtained. However, the output vector p is passed through the softmax function to randomize it. This output vector p and the correct value vector p' = (p'1, p'2, , p' s By updating the embedding matrix W so as to reduce the difference (error) between the Context and the Target, it is possible to obtain the embedding matrix W for predicting the Target from the Context.

[0030] As mentioned above, when the one-hot vector representing the context is multiplied by the embedding matrix W, the i-th column vector wi is extracted, and only this column vector wi is used in subsequent calculations. Therefore, this column vector wi can be considered as a distributed representation vector that indicates the characteristics of the i-th word.

[0031] Using the same concept, we can obtain a vector di of distributed representations that represent the entire document data by applying the Doc2Vec method. In this case, instead of the embedding matrix W described above, we use the embedding matrix D = (d1d2 d i ··· d U ) (di is a column vector) may be used. Note that the method for vectorizing document data is not limited to Word2Vec, and other methods such as Doc2Vec may also be used.

[0032] <Hardware configuration> The information processing device 100, the operation terminal 101, the local data server 102, the external data server 103, etc. in FIG. 1 have, for example, the hardware configuration of a computer 400 as shown in FIG.

[0033] Fig. 4 is a diagram showing an example of the hardware configuration of a computer according to an embodiment. As shown in Fig. 4, the computer 400 includes, for example, a central processing unit (CPU) 401, a read-only memory (ROM) 402, a random access memory (RAM) 403, a hard disk (HD) 404, a hard disk drive (HDD) controller 405, a display 406, an external device connection interface (I / F) 407, one or more network I / Fs 408, a keyboard 409, a pointing device 410, a digital versatile disk rewritable (DVD-RW) drive 412, a media I / F 414, and a bus line 415.

[0034] Of these, the CPU 401 controls the overall operation of the computer 400. The ROM 402 stores programs used to start up the computer 400, such as an IPL (Initial Program Loader). The RAM 403 is used, for example, as a work area for the CPU 401. The HD 404 stores programs such as an OS (Operating System), applications, and device drivers, as well as various data. The HDD controller 405 controls the reading and writing of various data from and to the HD 404, for example, under the control of the CPU 401.

[0035] The display 406 displays various types of information such as a cursor, menus, windows, characters, or images. The display 406 may be provided outside the computer 400. The external device connection I / F 407 is an interface such as a USB (Universal Serial Bus) that connects various external devices to the computer 400. One or more network I / Fs 408 are interfaces for communicating with other devices using, for example, the local network 10 or the external network 11.

[0036] The keyboard 409 is a type of input means having multiple keys for inputting characters, numbers, various instructions, etc. The pointing device 410 is a type of input means for selecting and executing various instructions, selecting a processing target, moving a cursor, etc. The keyboard 409 and pointing device 410 may be provided outside the computer 400.

[0037] The DVD-RW drive 412 controls reading and writing of various data from and to a DVD-RW 411, which is an example of a removable recording medium. The DVD-RW 411 is not limited to a DVD-RW, and may be a DVD-R or the like. The media I / F 414 controls reading and writing (storing) of data from and to a medium 413, such as a flash memory. The bus line 415 includes an address bus, a data bus, various control signals, and the like, for electrically connecting the above-mentioned components.

[0038] 4 is an example of the hardware configuration of the computer 400. The computer 400 may have any configuration as long as it has, for example, a CPU 401, a ROM 402, a RAM 403, one or more network I / Fs 408, and a bus line 415.

[0039] <Functional configuration> 5 is a diagram illustrating an example of the functional configuration of an information processing device according to an embodiment. For example, the information processing device 100 implements an acquisition unit 501, a word embedding vector generation unit 502, a first vector generation unit 503, a second vector generation unit 504, a calculation unit 505, an output unit 506, a reception unit 507, a data management unit 508, and a storage unit 509 by executing a predetermined program on the CPU 401 in FIG. 4. Note that at least a portion of the above functional configurations may be implemented by hardware.

[0040] The acquisition unit 501 executes an acquisition process for acquiring multiple pieces of target document data to be searched. The acquisition unit 501 acquires target document data, which is document data of existing technologies such as patent documents, papers, and technical journals, from one or more external data servers 103 or the local data server 102, and stores the acquired data (acquired data 511) in the storage unit 509, etc. Note that the acquisition unit 501 may acquire feature amount data of the target document data instead of (or in addition to) the target document data, and store the data as acquired data 511 in the storage unit 509, etc.

[0041] The word embedding vector generation unit 502 executes a word embedding vector generation process to generate vectors of distributed representations (hereinafter referred to as word embedding vectors) representing multiple target document data sets to be searched. The word embedding vector generation unit 502 vectorizes multiple target document data sets using, for example, a method such as Word2Vec or Doc2Vec described in FIG. 3.

[0042] The first vector generation unit 503 executes a first vector generation process to generate a vector (hereinafter referred to as a first vector) of a distributed representation representing technical document data received on the operation screen 200, as shown in Fig. 2(A) for example. The first vector generation unit 503 vectorizes the technical document data using a method such as Word2Vec or Doc2Vec described in Fig. 3.

[0043] 2A, the second vector generation unit 504 executes a second vector generation process to generate vectors (hereinafter referred to as second vectors) of distributed representations representing index document data received on the operation screen 200. The second vector generation unit 504 vectorizes the index document data using a method such as Word2Vec or Doc2Vec described in FIG.

[0044] The calculation unit 505 executes calculation processes such as adding and subtracting between vectors, calculating the similarity between vectors, etc. The calculation unit 505 also performs calculations such as calculating the similarity between a second vector and a plurality of composite vectors formed by combining a first vector with each of the word embedding vectors of a plurality of target document data.

[0045] The output unit 506 executes an output process using the word embedding vector, the first vector, and the second vector to output information on one or more target document data that can increase the index represented by the index document data by combining the word embedding vector, the first vector, and the second vector with the technology represented by the technical document data. For example, the output unit 506 uses the calculation unit 505 to calculate the similarity between the second vector and multiple composite vectors formed by combining the first vector with each of the word embedding vectors of the multiple target document data. The output unit 506 also outputs search results that include information on target document data corresponding to one or more composite vectors among the multiple composite vectors that have a higher similarity to the second vector than the other composite vectors. The output unit 506 may, for example, transmit (output) the output results to another device such as the operation terminal 101 for display, or may output the search results to the storage unit 509 for storage.

[0046] The reception unit 507, for example, displays an operation screen 200 as shown in FIG. 2(A) on the operation terminal 101 or the like, and executes a reception process for receiving setting operations by the user. For example, on the operation screen 200 as shown in FIG. 2(A), the user sets technical document data of the company that the user wants to use to create new value in the file setting field 202 corresponding to "technology." As a result, the reception unit 507 receives the file set in the file setting field 202 as technical document data. Also, on the operation screen 200 as shown in FIG. 2(A), the user sets index document data that describes a direction (index) that the user wants to develop in the file setting field 204 corresponding to "index." As a result, the reception unit 507 receives the file set in the file setting field 204 as index document data.

[0047] The directionality to be improved may be, for example, an image of a KPI (Key Performance Indicator), etc. Furthermore, the indicator document data may include, for example, character strings that explain indicators such as "ease of operation," for example, as in a Japanese dictionary.

[0048] The data management unit 508 stores and manages various data and information such as acquired data 511, word embedding vectors 512, vector data 513, and document data 514 in, for example, the storage unit 509. Here, the acquired data 511 includes multiple target document data acquired by the acquisition unit 501, or data such as features corresponding to the multiple target document data. The word embedding vectors 512 include word embedding vectors generated by the word embedding vector generation unit 502. The vector data 513 includes a first vector generated by the first vector generation unit 503 and a second vector generated by the second vector generation unit 504. The document data 514 includes, for example, technical document data and index document data accepted by the acceptance unit 507.

[0049] The data management unit 508 may store the above-mentioned data in a storage device, a storage server, a cloud service, or the like external to the information processing device 100. Here, as an example, the following description will be given assuming that the data management unit 508 stores and manages the above-mentioned data in the storage unit 509.

[0050] The storage unit 509 is realized by, for example, the program executed by the CPU 401 in FIG. 4, the HD 404, the HDD controller 405, etc., and stores various information, data, programs, etc., including each piece of data managed by the data management unit 508.

[0051] The functional configuration of the information processing device 100 shown in Fig. 5 is an example. For example, at least some of the functional configurations of the information processing device 100 shown in Fig. 5 may be provided outside the information processing device 100. Furthermore, the functional configurations of the information processing device 100 shown in Fig. 5 may be distributed and provided in multiple computers 400.

[0052] <Processing flow> Next, the process flow of the information providing method according to this embodiment will be described by exemplifying a number of embodiments.

[0053] [First embodiment] 6 and 7 are sequence diagrams showing an example of processing by the information processing device according to the first embodiment. This processing shows an example of processing executed by the information processing device 100 when a user sets technical document data, index document data, etc. on the operation screen 200 as shown in FIG. 2(A) and selects the "Execute" button 205.

[0054] In steps S601 and S602, when the receiving unit 507 receives a setting operation by the user on the operation screen 200 as shown in Fig. 2(A), the receiving unit 507 notifies the acquisition unit 501 of the received setting information. Note that this setting information includes, for example, technical document data of the company that is desired to be used to create new value, and index document data (or link information for acquiring the technical document data and index document data) that describes the direction (index) that the company wants to develop.

[0055] In step S603, the acquisition unit 603 acquires multiple pieces of target document data to be searched, for example, from one or more external data servers 103 or the local data server 102. For example, the acquisition unit 501 acquires document data related to existing technologies, such as patent documents, papers, or technical journals. If the setting information received from the reception unit 507 includes link information for acquiring the technical document data and the index document data, the acquisition unit 603 also acquires the technical document data and the index document data using the link information.

[0056] In step S604, the acquisition unit 603 stores in the storage unit 509 the target document data acquired from one or more external data servers 103 or the local data server 102, and the acquired technical document data and index document data.

[0057] In step S605, the acquisition unit 501 requests the word embedding vector generation unit 502 to generate a word embedding vector.

[0058] In steps S606 and S607, the word embedding vector generation unit 502 retrieves the target document data retrieved by the retrieval unit 501 from the storage unit 509, and generates word embedding vectors representing the retrieved target document data. In steps S608 and S609, the word embedding vector generation unit 502 stores the generated word embedding vectors in the storage unit 509, and sends a completion notification to the retrieval unit 501, which is the requesting source, indicating that the generation of the word embedding vectors has been completed.

[0059] In step S610, the acquisition unit 501 requests the first vector generation unit 503 to generate a first vector of a distributed representation representing the technical document data.

[0060] In steps S611 and S612, the first vector generation unit 503 retrieves the word embedding vectors stored in the storage unit 509 and uses the retrieved word embedding vectors to generate a first vector of a distributed representation representing the technical document data. In steps S613 and S614, the first vector generation unit 503 stores the generated first vector in the storage unit 509 and sends a completion notification to the requesting acquisition unit 501 indicating that the generation of the first vector has been completed.

[0061] In step S615, the acquiring unit 501 requests the second vector generating unit 504 to generate a second vector of a distributed representation representing the technical document data.

[0062] In steps S616 and S617, the second vector generation unit 504 retrieves the word embedding vectors stored in the storage unit 509 and uses the retrieved word embedding vectors to generate second vectors of distributed representations representing the index document data. In steps S618 and S619, the second vector generation unit 504 stores the generated second vectors in the storage unit 509 and sends a completion notification to the requesting acquisition unit 501 indicating that the generation of the second vectors is complete.

[0063] In the above example, the acquisition unit 501 controls the word embedding vector generation unit 502, the first vector generation unit 503, the second vector generation unit 504, etc., but this is just an example. For example, the information processing device 100 may have a separate unit control unit that controls the acquisition unit 501, the word embedding vector generation unit 502, the first vector generation unit 503, the second vector generation unit 504, etc. Alternatively, in the information processing device 100, the reception unit 507 may control the word embedding vector generation unit 502, the first vector generation unit 503, the second vector generation unit 504, etc.

[0064] 6, the acquisition unit 501 requests the output unit 506 to output the search results. In response to this, the output unit 506 transmits, for example, in step S621, a calculation request to the calculation unit to request a calculation of similarity.

[0065] In steps S622 to S624, when the calculation unit 505 receives a calculation request from the output unit 506, it retrieves the word embedding vector, the first vector, and the second vector from the storage unit 509. Furthermore, using the retrieved word embedding vector, the first vector, and the second vector, the calculation unit 505 executes the processes of steps S625 and S626 for each of the word embedding vectors representing the multiple target document data.

[0066] In steps S625 and S626, the calculation unit 505 generates a composite vector of one of the word embedding vectors representing the multiple target document data and the first vector, and calculates the similarity between the generated composite vector and the second vector. Here, the similarity is a value that indicates how close the composite vector is to the second vector representing the index document data specified by the user, and is calculated using, for example, the angle, distance, or cosine similarity between the composite vector and the second vector.

[0067] In step S627, the calculation unit 505 performs the processes of steps S625 and S626 on each of the word embedding vectors representing the plurality of target document data, and then transmits the calculation results to the output unit 506.

[0068] In step S628, the output unit 506 outputs a search result including information on one or more target document data items that, when combined with technical document data, can increase the index represented by the index document data compared to other target document data items. For example, as shown in FIG. 2B, the output unit 506 may use vectors to illustrate a combination of existing technology, "back monitor + multi-viewpoint image synthesis" 213, which can increase the index "ease of operation" 212 when combined with a technology, "back monitor" 211. Alternatively, the output unit 506 may output a search result that lists, in order of similarity, information on target document data items corresponding to one or more composite vectors among multiple composite vectors that have a higher similarity than other composite vectors.

[0069] Through the above processing, the information processing device 100 according to the first embodiment can search for technical document data from among multiple existing technical document data that can increase a specified index when combined with a specified technology, and output the search results.

[0070] [Second embodiment] Fig. 8 is a sequence diagram showing an example of target document data acquisition processing according to the second embodiment. This processing shows an example of target document data acquisition processing executed by the acquisition unit 501 in step S603 in Fig. 6, for example. In the second embodiment, an example of processing will be described in which the acquisition destination and acquisition method of target document data of the information processing device 100 can be set.

[0071] At the start of the process shown in FIG. 8, if the local data server 102 in FIG. 1A is set as the acquisition destination of the target document data, the information processing system 1 executes an acquisition process 810 shown in steps S811 to S813.

[0072] In step S811, the acquisition unit 501 transmits an acquisition request for requesting acquisition of multiple target document data or feature amounts of multiple target document data to the local data server 102. Note that if the acquisition method for target document data of the information processing device 100 is set to "feature amount," the acquisition request transmitted by the acquisition unit 501 includes information requesting transmission of feature amounts.

[0073] If the acquisition request sent by the acquisition unit 501 includes information requesting the transmission of features, in step S812, the local data server 102 converts the target document data into features that can be used to generate word embedding vectors. By converting the target document data into features, the information processing device 100 can generate word embedding vectors for the target document data held by the local data server 102, including, for example, private target document data.

[0074] On the other hand, if the acquisition request sent by the acquisition unit 501 does not include information requesting transmission of feature amounts, the local data server 102 skips the process of step S812 and executes the process of step S813.

[0075] In step S813, the local data server 102 transmits the plurality of target document data or the feature quantities of the plurality of target document data to the acquisition unit 501 that has made the request.

[0076] Also, at the start of the processing of Figure 8, if external data server 103a, one of the external data servers 103 in Figure 1 (A), is set as the acquisition destination for the target document data, the information processing system 1 executes acquisition processing 820 shown in steps S821 to S823.

[0077] In step S821, the acquisition unit 501 sends an acquisition request to the external data server 103a to request acquisition of multiple target document data or feature quantities of multiple target document data. If the acquisition request sent by the acquisition unit 501 includes information requesting the transmission of feature quantities, in step S822, the external data server 103a converts the multiple target document data into feature quantities that can be used to generate word embedding vectors.

[0078] On the other hand, if the acquisition request sent by the acquisition unit 501 does not include information requesting transmission of feature amounts, the external data server 103a skips the process of step S822 and executes the process of step S823.

[0079] In step S823, the external data server 103a transmits a plurality of target document data or feature quantities of a plurality of target document data to the acquisition unit 501 that has made the request.

[0080] Similarly, at the start of the processing of Figure 8, if external data server 103b, one of the external data servers 103 in Figure 1(A), is set as the destination for obtaining the target document data, the information processing system 1 executes the obtaining processing 830 shown in steps S831 to S833.

[0081] In step S831, the acquisition unit 501 sends an acquisition request to the external data server 103b to request acquisition of multiple target document data or feature quantities of multiple target document data. If the acquisition request sent by the acquisition unit 501 includes information requesting the transmission of feature quantities, in step S832, the external data server 103b converts the multiple target document data into feature quantities that can be used to generate word embedding vectors.

[0082] On the other hand, if the acquisition request sent by the acquisition unit 501 does not include information requesting transmission of feature amounts, the external data server 103b skips the process of step S832 and executes the process of step S833.

[0083] In step S833, the external data server 103b transmits a plurality of target document data or feature quantities of a plurality of target document data to the acquisition unit 501 that has made the request.

[0084] By the above process, the user can obtain multiple target document data from the configured server device. Furthermore, even if the user cannot obtain the original target document data from the data source due to security or rights reasons, the user can convert the data into features and obtain the data.

[0085] [Third embodiment] 9 is a sequence diagram illustrating an example of a process for generating word embedding vectors according to the third embodiment. This process illustrates another example of the process for generating word embedding vectors executed by the information processing device 100 in steps S603 to S609 of FIG.

[0086] In step S900, the acquiring unit 501 determines whether or not there is a word embedding vector stored in the storage unit 509. If there is no word embedding vector, the information processing device 100 executes the same processes as steps S603 to S609 in FIG.

[0087] On the other hand, if there is a word embedding vector, the information processing device 100 executes the processes of steps S901 to S905 in FIG.

[0088] In step S901, the acquiring unit 501 sends a request to the word embedding vector generating unit 502 to update the word embedding vector.

[0089] In steps S902 and S903, when the word embedding vector generation unit 502 receives a request to update the word embedding vector, it retrieves the word embedding vector stored in the storage unit 509 and determines whether an update is necessary. For example, if the generated word embedding vector cannot represent the technical document data specified by the user, the word embedding vector generation unit 502 determines that an update of the word embedding vector is necessary. Note that whether the generated word embedding vector can represent the technical document data specified by the user can be determined, for example, by whether the vector length when the generated word embedding vector is used to vectorize the technical document specified by the user is greater than or equal to a threshold.

[0090] If an update is necessary, in step S904, the information processing device 100 executes a process of generating a word embedding vector, for example, as shown in steps S603 to S609 in Fig. 6. On the other hand, if an update is not necessary, the information processing device 100 skips the process of step S904 and executes the process of step S905.

[0091] In step 905, the word embedding vector generation unit 502 sends a completion notification to the requesting acquisition unit 501 indicating that the update of the word embedding vector has been completed.

[0092] By performing the above process, the information processing device 100 can suppress the execution of the process of updating the word embedding vector when updating the word embedding vector is not necessary.

[0093] [Fourth embodiment] FIG. 10 is a sequence diagram illustrating an example of a process for generating word embedding vectors according to the fourth embodiment. The information processing device 100 may periodically perform the process for updating word embedding vectors shown in FIG. 10. For example, the information processing device 100 may perform the process for updating word embedding vectors shown in FIG. 10 during times of low usage, such as at night, or may perform the process for updating word embedding vectors shown in FIG. 10 every time a search process is performed a predetermined number of times. Furthermore, the information processing device 100 may perform the process for updating word embedding vectors shown in FIG. 10 by combining multiple conditions.

[0094] In step S1001, the word embedding vector generation unit 502 requests the acquisition unit 501 to acquire target document data in order to update multiple target document data.

[0095] In step S1002, the acquisition unit 501 acquires multiple pieces of target document data, for example, from one or more external data servers 103 or the local data server 102. In steps S1003 and S1004, the acquisition unit 501 stores the acquired multiple pieces of target document data in the storage unit 509, and notifies the word embedding vector generation unit 502 of an acquisition completion notification indicating that acquisition of the target document data has been completed.

[0096] In steps S1005 and S1006, the word embedding vector generation unit 502 acquires the target document data acquired by the acquisition unit 501 from the storage unit 509, and generates word embedding vectors representing the acquired target document data. In step S1007, the word embedding vector generation unit 502 stores the generated word embedding vectors in the storage unit 509 or the like.

[0097] By periodically performing the above process, the information processing device 100 can reduce the frequency with which the word embedding vector generation process is performed, for example, in the word embedding vector generation process shown in Fig. 9. Alternatively, by periodically performing the above process, the information processing device 100 may omit the word embedding vector generation process shown in Fig. 9.

[0098] [Fifth embodiment] 11 is a sequence diagram showing an example of a process of downloading target document data according to the fifth embodiment. Preferably, the information processing device 100 has a function of downloading target document data in response to a user's selection operation on a display screen of search results that the output unit 506 causes the operation terminal 101 to display.

[0099] In step S1101, for example, after executing the process (search process) of the information processing device as shown in FIGS. 6 and 7, the information processing device 100 executes the processes from step S1102 onwards.

[0100] In step S1102, the output unit 506 displays, as an example, a result display screen 1200 as shown in Fig. 12(A). In the example of Fig. 12(A), the result display screen 1200 displays a list of multiple relevant documents (target document data) 1201 extracted by the search process in step S1101 in descending order of relevance (for example, similarity). The user can request downloading of document data of the selected document name by selecting a display element 1202 displaying a document name from the multiple relevant documents 1201.

[0101] In steps S1103 and S1104, upon receiving a document data download request from the user, the receiving unit 507 transmits to the output unit 506 a download request including document information of the requested document data (hereinafter referred to as requested document data).

[0102] If the requested document data requested in the download request is a free document that can be used free of charge, the output unit 506 executes process 1110 for the free document. If the requested document data requested in the download request is a paid document that can be used for a fee, the output unit 506 executes process 1120 for the paid document. If the document data requested in the download request includes both free and paid documents, the output unit 506 executes process 1110 for the free document and process 1120 for the paid document.

[0103] If the requested document data requested in the download request includes a free document, the output unit 506 acquires the requested document data from the storage unit 509 or the like in step S1111.

[0104] If the document data requested in the download request includes a paid document, the output unit 506 executes the processes of steps S1121 to S1123.

[0105] In step S1121, the output unit 506 checks whether the user has a usage right, such as a fixed-price usage right, to use the requested paid document. If the user does not have the usage right, the output unit 506 executes the processes of steps S1122 and S1123. On the other hand, if the user has the usage right, the output unit 506 executes the process of step S1123.

[0106] In step S1122, the output unit 506 performs the billing process using, for example, the reception unit 507. For example, the reception unit 507 displays a purchase screen for document data and receives payment operations by the user.

[0107] In step S1123, the output unit 506 acquires the request document data from the storage unit 509 or the like.

[0108] In step S1131, the output unit 506 transmits the acquired request document data to, for example, the operation terminal 101 or the like.

[0109] Through the above process, the user can easily obtain document data searched by the information processing system 1.

[0110] <Example of display screen> Next, an example of a display screen that the information processing device 100 displays on the operation terminal 101 or the like will be described.

[0111] (Example of the results display screen) 12 and 13 are diagrams showing examples of a result display screen according to an embodiment. Fig. 12(A) shows an example of a result display screen that the information processing device 100 displays on the operation terminal 101 or the like in the first to fifth embodiments, for example. In the example of Fig. 12(A), as described above, the result display screen 1200 displays a list of multiple relevant documents (target document data) 1201 extracted by the search process in descending order of relevance (for example, similarity). A user can select a display element 1202 displaying the document name of a relevant document from the multiple relevant documents 1201, thereby requesting the download of document data of the selected document name, for example.

[0112] 12(B) shows another example of a result display screen that the information processing device 100 displays on the operation terminal 101 or the like. In the example of FIG. 12(B), the result display screen 1210 visually displays a combination of technologies such as "backup monitor + stereo parallax ranging" in a graph with the index "ease of operation" on the horizontal axis and the index "safety" on the vertical axis, so as to indicate the similarity between the two indexes. In this way, the result display screen output by the output unit 506 can be output in various formats.

[0113] 13(A) shows another example of a result display screen that the information processing device 100 displays on the operation terminal 101 or the like. In the example of FIG. 13(A), in addition to the result display screen 1200 described in FIG. 12(A), the result display screen 1300 can include, for example, advertising information 1301 from a company or organization that wants to sell its technology using the information processing system 1. For example, a company or the like that wants to add an advertisement registers the advertising information 1301 when registering target document data in one or more external data servers 103 or the like. In addition, a service provider or the like that operates the information processing system 1 may charge a predetermined fee for registering the advertising information 1301.

[0114] Fig. 13(B) shows another example of a result display screen that the information processing device 100 displays on the operation terminal 101 or the like. In the example of Fig. 13(B), a result display screen 1310 can display, in addition to the result display screen 1200 described in Fig. 12(A), a selection button 1311 for contacting a company, institution, or the like that possesses the technology. When the information processing device 100 accepts a selection operation of the selection button 1311 by the user, it provides the user with a communication means, such as email or chat, for contacting a company, or the like that possesses the corresponding technology.

[0115] Preferably, the output unit 506 prevents the user from selecting the selection button 1311 corresponding to document data for which no contact means is registered, for example, by displaying it in half brightness, graying it out, or hiding it.

[0116] (Example of settings screen) 14 is a diagram showing an example of a setting screen according to an embodiment. The receiving unit 507 may display a setting screen 1400 as shown in FIG. 14, for example, and receive various setting information together with technical document data and index document data.

[0117] 14, a setting screen 1400 is provided with a setting field 1402 in which the number of existing technologies to be combined can be set, in addition to a setting field 1401 in which technical document data is set. For example, in the example of FIG. 2(B), the information processing device 100 provides information on a combination of an existing technology, "back monitor + multi-viewpoint image synthesis" 213, which can be combined with a technology, "back monitor" 211, to increase the indicator, "ease of operation" 212. On the other hand, by setting, for example, "2" in the setting field 1402 in which the number of combinations is set, the information processing device 100 can change the number of existing technologies to be combined, such as "back monitor + multi-viewpoint image synthesis + stereo parallax ranging."

[0118] 14, the setting screen 1400 is provided with a setting field 1403 for setting index document data, a setting field 1404 for setting the number of indexes, which is the number of index document data to be set, and a setting field 1405 for setting the weight of each index document data. For example, when "2" is set in the setting field 1403 for the number of indexes, the information processing device 100 outputs search results for two indexes (e.g., ease of operation and safety) as shown in the result display screen 1210 shown in FIG. 12(B). Furthermore, the setting field 1405 for setting the weight allows the user to set the weight of each index document data.

[0119] (Example of document data registration screen) 15A and 15B are diagrams showing examples of a registration screen for document data according to an embodiment. FIG. 15A shows an example of a registration screen 1500 for registering target document data, for example, in one or more external data servers 103 or a local data server. In the example of FIG. 15A, the registration screen 1500 has a setting field 1501 for setting a file name for the target document data and a setting field 1502 for setting whether the target document data is to be made public or private. Here, if the target document data is made public, the information processing device 100 uses the target document data as one of the existing technologies when another user uses the information processing system 1.

[0120] On the other hand, if the target document data is made private, the information processing device 100 prohibits another user from using the target document data when using the information processing system 1. Alternatively, if the target document data is made private, the information processing device 100 may permit another user to use only the feature amounts of the target document data when using the information processing system 1.

[0121] 15(B) shows an example of a registration screen 1510 when registering target document data in, for example, one or more external data servers 103 or a local data server. In the example of FIG. 15(B), the registration screen 1510 has a setting field 1511 for setting a file name of the target document data, as well as a setting field 1512 for setting a registration destination of the target document data. Using the setting field 1512 for setting the registration destination, the user can select, for example, whether to register the target document data in the external data server 103 available to external users, or in the local data server 102 available only to in-house users.

[0122] The registration screen for the target document data shown in Fig. 15 is an example. The registration screen for the target document data may be provided with a setting field for setting advertising information 1301 to be displayed on a result display screen 1300 as shown in Fig. 13(A), for example. The registration screen for the target document data may also be provided with a setting field for setting a contact method such as email or chat, which is displayed when a selection button 1311 is selected on a result display screen 1310 as shown in Fig. 13(B).

[0123] As described with reference to FIGS. 12 to 15, the information processing system 1 and the information processing device 100 according to this embodiment can be modified and applied in various ways.

[0124] As described above, according to each embodiment of the present invention, it is possible to provide an information processing device 100 and an information processing system 1 that can search for technical document data that can increase a specified index by combining it with a specified technology from among multiple existing technical document data.

[0125] <Supplementary information> Each function of each embodiment described above can be realized by one or more processing circuits. Here, the term "processing circuit" in this specification includes a processor programmed to perform each function by software, such as a processor implemented by an electronic circuit, as well as devices such as an ASIC (Application Specific Integrated Circuit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), and conventional circuit modules designed to perform each of the functions described above.

[0126] Furthermore, the devices described in the examples represent only one of multiple computing environments for implementing the embodiments disclosed herein. In one embodiment, information processing device 100 includes multiple computing devices, such as a server cluster. The multiple computing devices are configured to communicate with each other via any type of communication link, including a network or shared memory, and perform the processes disclosed herein. Furthermore, each element of information processing device 100 may be integrated into a single information processing device or may be separated into multiple information processing devices.

[0127] 5 can be configured to share, for example, various combinations of the processes of the information processing device 100 shown in FIGS. 6 to 11. For example, at least a portion of the processes executed by the information processing device 100 may be executed by the operation terminal 101. Furthermore, at least a portion of the processes executed by the information processing device 100 may be executed by, for example, an external server device, a cloud service, or the like. [Explanation of symbols]

[0128] 1. Information Processing Systems 502 Word Embedding Vector Generation Unit 503 First Vector Generator 504 Second Vector Generator 505 Arithmetic unit 506 Output section 507 Reception [Prior art documents] [Patent documents]

[0129] [Patent Document 1] Japanese Patent Application Laid-Open No. 2001-265808

Claims

1. a reception unit that receives a setting operation for setting technical document data and a setting operation for setting index document data; a word embedding vector generation unit that generates word embedding vectors, which are vectors of distributed representations representing a plurality of target document data sets to be searched, using the plurality of target document data sets; a first vector generation unit that generates a first vector of a distributed representation representing the technical document data; a second vector generation unit that generates a second vector of a distributed representation representing the index document data; a calculation unit that calculates a similarity between the second vector and a plurality of combined vectors obtained by combining the first vector and each of the word embedding vectors; an output unit that outputs search results including information on the target document data corresponding to one or more composite vectors among the plurality of composite vectors, the similarity of which is higher than that of other composite vectors; An information processing device having the above.

2. The information processing apparatus according to claim 1 , wherein the output unit displays a list of information on the target document data corresponding to one or more composite vectors whose similarity is higher than that of other composite vectors, in order of similarity.

3. The information processing device according to claim 1 , wherein the word embedding vector generation unit periodically updates the word embedding vector.

4. The information processing device according to claim 1 , wherein the reception unit displays a setting screen on which the number of the word embedding vectors to be combined into the composite vector can be set.

5. The information processing apparatus according to claim 1 , wherein the reception unit displays a setting screen on which the number of the index document data can be set.

6. The information processing apparatus according to claim 1 , further comprising an acquisition unit configured to acquire the plurality of target document data from one or more server devices in accordance with a setting.

7. The information processing device a reception process for receiving a setting operation for setting technical document data and a setting operation for setting index document data; a word embedding vector generation process for generating word embedding vectors, which are vectors of distributed representations representing a plurality of target document data sets to be searched, using the plurality of target document data sets; a first vector generation process for generating a first vector of a distributed representation representing the technical document data; a second vector generation process for generating a second vector of a distributed representation representing the index document data; a computation process for calculating the similarity between the second vector and a plurality of composite vectors obtained by combining the first vector and each of the word embedding vectors; and an output process for outputting search results including information on the target document data corresponding to one or more composite vectors among the plurality of composite vectors, the similarity of which is higher than that of the other composite vectors; A method of providing information.

8. On the computer, a reception process for receiving a setting operation for setting technical document data and a setting operation for setting index document data; a word embedding vector generation process for generating word embedding vectors, which are vectors of distributed representations representing a plurality of target document data sets to be searched, using the plurality of target document data sets; a first vector generation process for generating a first vector of a distributed representation representing the technical document data; a second vector generation process for generating a second vector of a distributed representation representing the index document data; a computation process for calculating the similarity between the second vector and a plurality of composite vectors obtained by combining the first vector and each of the word embedding vectors; and an output process for outputting search results including information on the target document data corresponding to one or more composite vectors among the plurality of composite vectors, the similarity of which is higher than that of the other composite vectors; A program that executes.

Citation Information

Patent Citations

  • System and method for information retrieval

    JP2001265808A

  • Apparatus and program for supporting systematization of idea

    JP2011204009A

  • System and method on emerging technology product portfolio generation based on firm''s technology capability

    KR101292663B1

  • Apparatus and method of discovering promising convergence technologies based-on network, storage media storing the same

    KR1020190097915A

  • Information processing device, method for generating information for selecting tie-up destination, and program

    WO2008075744A1