Information processing device, information providing method, and program

The information processing device addresses the limitation of conventional systems by generating document and technology vectors to identify technologies that can expand unexplored fields, enhancing creative work through innovative combinations.

JP7757611B2Active Publication Date: 2025-10-22RICOH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2021010160
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-01-26
Publication Date
2025-10-22
Estimated Expiration
2041-01-26

AI Technical Summary

Technical Problem

Conventional information retrieval systems fail to support creative work by searching for technical documents or technical fields that can expand unexplored technological fields by combining a specified technology with existing technical document data.

Method used

An information processing device that receives a setting operation for technical document data, generates document vectors and known technology vectors, and outputs search results including information on technologies that can expand unexplored fields by combining them with specified technologies from multiple existing documents.

Benefits of technology

Enables the search for information on technologies that can expand unexplored technological areas by combining them with a specified technology from multiple existing technical documents, facilitating creative work.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007757611000001
    Figure 0007757611000001
  • Figure 0007757611000002
    Figure 0007757611000002
  • Figure 0007757611000003
    Figure 0007757611000003
Patent Text Reader

Abstract

To provide an information processing apparatus configured to retrieve information on one or more technologies that can extend unexplored technical fields in combination with a designated technology from existing multiple technical documents.SOLUTION: An information processing apparatus includes: a receiving unit which receives settings on technical document data; a word embedding vector generation unit which generates a word embedding vector using multiple document data to be retrieved; a document vector generation unit which generates a document vector with distributed representation which represents the technical document data using the word embedding vector; and an output unit which outputs a retrieval result including information on one or more technologies that can extend a new technical field in combination with technologies represented by known technical vectors representing existing technologies by using the known technical vectors and the document vector.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information providing method, and a program. [Background technology]

[0002] 2. Description of the Related Art There is an information retrieval system that searches through a plurality of document data for document data that meets a user's request and outputs the search results.

[0003] For example, an information search system is known that searches for books to recommend to a customer by calculating the similarity between a customer profile based on information about books purchased by the customer and a keyword vector for each book based on a book database (see, for example, Patent Document 1). Summary of the Invention [Problem to be solved by the invention]

[0004] In recent years, efforts to reform work styles have become more active, and various tools have been developed to make work more efficient. Most of these business improvement tools are designed to automate simple or repetitive tasks or to facilitate communication. While the aim is to use the time saved by using these tools to carry out work that creates new value, the current situation is that there are almost no tools that support creative work.

[0005] Innovation is described as a "new combination" brought about by new combinations of existing technologies. Therefore, there is a demand for support tools for creative work that can present new value by combining existing technologies with existing technologies.

[0006] However, conventional information retrieval systems such as those shown in Cited Document 1 are unable to search for technical documents or technical fields that can expand unexplored technical fields by combining a specified technology with existing technical document data.

[0007] One embodiment of the present invention has been made in consideration of the above-mentioned problems, and provides an information processing device that can search for information on one or more technologies that can expand unexplored technological fields by combining them with a specified technology from multiple existing technical documents. [Means for solving the problem]

[0008] In order to solve the above problem, an information processing apparatus according to one embodiment includes a receiving unit that receives a setting operation for setting technical document data, a document vector generating unit that generates document vectors of distributed representations that represent the technical document data, and a plurality of known technology vectors that represent existing technologies. Each of and the document vector, and an output unit that outputs search results including information about the technology represented by the known technology vector corresponding to one or more of the multiple composite vectors whose similarity is lower than that of the other composite vectors. [Effects of the Invention]

[0009] According to one embodiment of the present invention, an information processing device can be provided that can search for information on one or more technologies that can expand unexplored technological areas by combining them with a specified technology from multiple existing technical documents. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a diagram illustrating an example of a system configuration of an information processing system according to an embodiment. [Figure 2] FIG. 1 is a diagram illustrating an overview of processing performed by an information processing device according to an embodiment. [Figure 3] FIG. 10 is a diagram illustrating vectorization of document data according to an embodiment. [Figure 4] FIG. 2 is a diagram illustrating an example of a hardware configuration of a computer according to an embodiment. [Figure 5]FIG. 1 is a diagram illustrating an example of a functional configuration of an information processing device according to an embodiment. [Figure 6] FIG. 10 is a sequence diagram illustrating an example of processing performed by an information processing device according to an embodiment. [Figure 7] 10 is a flowchart illustrating an example of a search process according to the first embodiment. [Figure 8] FIG. 10 is a sequence diagram illustrating an example of a search process according to the second embodiment. [Figure 9] 10 is a flowchart illustrating an example of an output process according to the third embodiment. [Figure 10] 13 is a flowchart illustrating an example of a search process according to the fourth embodiment. [Figure 11] FIG. 13 is a sequence diagram showing an outline of a search process according to the fifth embodiment. [Figure 12] 13 is a flowchart showing an example of a search process according to the sixth embodiment. [Figure 13] 13 is a flowchart showing an example of search processing according to the seventh embodiment. [Figure 14] FIG. 20 is a sequence diagram showing an example of a download process of document data according to the eighth embodiment. [Figure 15] FIG. 10 is a diagram illustrating an example of a result display screen according to an embodiment. [Figure 16] FIG. 10 is a diagram illustrating an example of a setting screen according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. <System configuration> Fig. 1 is a diagram showing an example of a system configuration of an information processing system according to an embodiment. The information processing system 1 includes, for example, an information processing device 100, an operation terminal 101, a local data server 102, and one or more external data servers 103. In the example of Fig. 1(A), the information processing device 100, the operation terminal 101, and the local data server 102 are communicably connected to each other via a local network 10 such as a LAN (Local Area Network). Furthermore, the one or more external data servers 103 are communicably connected to the information processing device 100 via, for example, an external network 11 such as the Internet and the local network 10.

[0012] The information processing device 100 is an information processing device having a computer configuration or a system including a plurality of computers, and provides the information providing service according to this embodiment to the user by executing a predetermined program.

[0013] The operation terminal 101 is, for example, an information terminal such as a PC (Personal Computer), a tablet terminal, or a smartphone, and a user can use, for example, the information providing service provided by the information processing device 100 using the operation terminal 101. However, this is not limited to this, and the user may, for example, operate the information processing device 100 to use the information providing service provided by the information processing device 100.

[0014] The local data server 102 is, for example, an information processing device having a computer configuration, or a system including multiple computers, and stores some or all of multiple document data to be searched (hereinafter referred to as target document data). As an example, the local data server 102 stores, among the multiple target document data, various document data related to technologies and the like that are owned by a company or the like.

[0015] The one or more external data servers 103 are, for example, information processing devices having a computer configuration or systems including multiple computers, and store some or all of the target document data to be searched. The external data server 103 stores, for example, a database of patent documents, a database of papers in various fields, or a database of technical publications published by other companies, etc.

[0016] The information processing device 100 can acquire target document data such as patent documents, papers, technical journals, etc. relating to various existing technologies from a local data server 102 and one or more external data servers 103.

[0017] 1(B), the information processing device 100 may be, for example, a server device connected to an external network 11, a cloud service, or the like. In this case, a user can use an information providing service provided by the information processing device 100 by accessing the information processing device 100 using an operation terminal 101.

[0018] (Processing Overview) 2 is a diagram for explaining an outline of the processing of the information processing device. The information processing device 100 sets technical document data corresponding to a predetermined technology, and provides information on one or more technologies that can expand an unexplored technical field by combining the technical document data with the technology corresponding to the technical document data.

[0019] As a specific example, suppose a department that developed rearview monitors for automobiles in the 1990s and 2000s wants to strengthen its technical capabilities in "rearview monitor" technology, but is unsure of the direction of the technical field in which it should be developed. In such a case, the user sets information about "rearview monitor" as information about an unexplored technical field on an operation screen (or setting screen) 200, such as that shown in FIG. 2(A), provided by the information processing device 100. For example, on the operation screen 200, the user enters the character string "rearview monitor" in the name input field 201, sets the file name of technical document data related to rearview monitors in the file setting field 202, and selects the "execute" button 203.

[0020] In response to this, the information processing device 100 provides information on technologies that can expand the unexplored technology field by combining it with the set technology "back monitor" 211, for example, as shown in Fig. 2(B). In the example of Fig. 2(B), the information processing device 100 indicates an unexplored technology field "preservability" 212 as information on technologies that can expand the unexplored technology field. As another example, the information processing device 100 may provide information on one or more target document data related to the unexplored technology field "preservability" 212 as information on technologies that can expand the unexplored technology field.

[0021] Preferably, the information processing device 100 may provide information such as "back monitor + video recording" 213, a combination of an existing technology that can expand the searched unexplored technology area "preservability" 212 by combining it with the set technology "back monitor" 211. This makes it easier for the user to come up with new technologies such as a "drive recorder" from the set technology "back monitor" and information on "video recording" provided by the information processing device 100 (for example, related technical literature on video recording).

[0022] In this way, the information processing device 100 can provide information on unexplored technical fields related to the set technology by combining it with the technology set by the user. That is, the information processing system 1 according to this embodiment can provide the user with information on novel technical fields for, for example, an existing in-house technology whose indicators (directions) to be developed are unknown.

[0023] (Regarding document data vectorization) For example, the information processing device 100 vectorizes the set technical document data, calculates a composite vector of each of the vectorized technology and vectors representing existing technology (hereinafter referred to as known technology vectors), and calculates the similarity between the composite vector and the known technology vector. In addition, the information processing device 100 identifies an unexplored technical field in the set technology by extracting one or more composite vectors with a lower calculated similarity (or a higher degree of dissociation).

[0024] In this embodiment, the method for vectorizing document data is not particularly limited, but here, as an example, an overview of vectorizing document data using the Word2Vec method will be described.

[0025] 3 is a diagram illustrating vectorization according to an embodiment. In this embodiment, generating vectors of distributed representations representing document data is referred to as vectorization of document data. As a method for vectorizing document data, for example, the Word2Vec method as shown in FIG. 3(A) can be applied.

[0026] For example, if the vocabulary size of a certain word is S, that word can be represented by an S-dimensional one-hot vector. However, the one-hot vectors representing these words are not distributed representation vectors, and the vectors are not related to each other at all. In addition, the cosine similarity between all vectors is 0, so all vectors are dissimilar.

[0027] Document data is not simply a collection of words; its meaning changes depending on the order in which the words appear. Natural language processing techniques such as Word2Vec improve accuracy by incorporating this word order as important information. For example, the words before and after a word contained in document data are extracted, and the last word is called the Target, and the set of words before that is called the Context. Furthermore, by learning an embedding matrix (weights) W for predicting the Target when a certain Context is given, each column vector of the word embedding matrix W can be used as a vector of the distributed representation of the word.

[0028] For example, the word embedding matrix W is defined as W=(w1w2 w i ··· w S )(w i is a column vector) and a certain context (the i-th word of the document data) is input as shown in Figure 3(A). In addition, when the one-hot vector of the i-th word is multiplied by the embedding matrix W, the i-th column vector wi = (w i,1 ,w i,2 ,···,w i,N )T (the hidden layer in Figure 3(A)). Here, the one-hot vector is a column vector in which only the i-th element is 1 and the other elements are 0.

[0029] Also, the intermediate layer column vector w i Then, we add the matrix W, which is the transpose of the embedding matrix W as shown in Figure 3(B). T By multiplying by this, the column vector of the output layer (output vector) p = (p1, p2, , p s )T is obtained. However, the output vector p is passed through the softmax function to randomize it. This output vector p and the correct value vector p' = (p'1, p'2, , p' s By updating the embedding matrix W so as to reduce the difference (error) between the context and the target, it is possible to obtain the embedding matrix W for predicting the target from the context.

[0030] As mentioned above, when the one-hot vector representing the context is multiplied by the embedding matrix W, the i-th column vector wi is extracted, and only this column vector wi is used in subsequent calculations. Therefore, this column vector wi can be considered as a distributed representation vector that indicates the characteristics of the i-th word.

[0031] Using the same concept, we can obtain a vector di of distributed representations that represent the entire document data by applying the Doc2Vec method. In this case, instead of the embedding matrix W described above, we use the embedding matrix D = (d1d2 d i ··· d U ) (di is a column vector) may be used. Note that the method for vectorizing document data is not limited to Word2Vec, and other methods such as Doc2Vec may also be used.

[0032] <Hardware configuration> The information processing device 100, the operation terminal 101, the local data server 102, the external data server 103, etc. in FIG. 1 have, for example, the hardware configuration of a computer 400 as shown in FIG.

[0033] Fig. 4 is a diagram showing an example of the hardware configuration of a computer according to an embodiment. As shown in Fig. 4, the computer 400 includes, for example, a central processing unit (CPU) 401, a read-only memory (ROM) 402, a random access memory (RAM) 403, a hard disk (HD) 404, a hard disk drive (HDD) controller 405, a display 406, an external device connection interface (I / F) 407, one or more network I / Fs 408, a keyboard 409, a pointing device 410, a digital versatile disk rewritable (DVD-RW) drive 412, a media I / F 414, and a bus line 415.

[0034] Of these, the CPU 401 controls the overall operation of the computer 400. The ROM 402 stores programs used to start up the computer 400, such as an IPL (Initial Program Loader). The RAM 403 is used, for example, as a work area for the CPU 401. The HD 404 stores programs such as an OS (Operating System), applications, and device drivers, as well as various data. The HDD controller 405 controls the reading and writing of various data from and to the HD 404, for example, under the control of the CPU 401.

[0035] The display 406 displays various types of information such as a cursor, menus, windows, characters, or images. The display 406 may be provided outside the computer 400. The external device connection I / F 407 is an interface such as a USB (Universal Serial Bus) that connects various external devices to the computer 400. One or more network I / Fs 408 are interfaces for communicating with other devices using, for example, the local network 10 or the external network 11.

[0036] The keyboard 409 is a type of input means having multiple keys for inputting characters, numbers, various instructions, etc. The pointing device 410 is a type of input means for selecting and executing various instructions, selecting a processing target, moving a cursor, etc. The keyboard 409 and pointing device 410 may be provided outside the computer 400.

[0037] The DVD-RW drive 412 controls reading and writing of various data from and to a DVD-RW 411, which is an example of a removable recording medium. The DVD-RW 411 is not limited to a DVD-RW, and may be a DVD-R or the like. The media I / F 414 controls reading and writing (storing) of data from and to a medium 413, such as a flash memory. The bus line 415 includes an address bus, a data bus, various control signals, and the like, for electrically connecting the above-mentioned components.

[0038] 4 is an example of the hardware configuration of the computer 400. The computer 400 may have any configuration as long as it has, for example, a CPU 401, a ROM 402, a RAM 403, one or more network I / Fs 408, and a bus line 415.

[0039] <Functional configuration> 5 is a diagram illustrating an example of the functional configuration of an information processing device according to an embodiment. For example, the information processing device 100 implements an acquisition unit 501, a word embedding vector generation unit 502, a document vector generation unit 503, an extraction unit 504, a calculation unit 505, an output unit 506, a reception unit 507, a data management unit 508, and a storage unit 509 by executing a predetermined program on the CPU 401 in FIG. 4. Note that at least a portion of the above functional configurations may be implemented by hardware.

[0040] The acquisition unit 501 executes an acquisition process for acquiring multiple pieces of target document data to be searched. The acquisition unit 501 acquires target document data, which is document data of existing technologies such as patent documents, papers, and technical journals, from one or more external data servers 103 or the local data server 102, and stores the acquired data (acquired data 511) in the storage unit 509, etc. Note that the acquisition unit 501 may acquire feature amount data of the target document data instead of (or in addition to) the target document data, and store the data as acquired data 511 in the storage unit 509, etc.

[0041] The word embedding vector generation unit 502 executes a word embedding vector generation process to generate vectors of distributed representations (hereinafter referred to as word embedding vectors) representing multiple target document data sets to be searched. The word embedding vector generation unit 502 vectorizes multiple target document data sets using, for example, a method such as Word2Vec or Doc2Vec described in FIG. 3.

[0042] The document vector generation unit 503 executes a document vector generation process to generate vectors of distributed representations (hereinafter referred to as document vectors) representing technical document data received on the operation screen 200, as shown in Fig. 2(A). The document vector generation unit 503 vectorizes the technical document data using a method such as Word2Vec or Doc2Vec described in Fig. 3.

[0043] The extraction unit 504 performs an extraction process to extract a predetermined number of principal component vectors by performing principal component analysis on a plurality of known technology vectors. Note that principal component analysis is a multivariate analysis method that synthesizes a small number of uncorrelated variables called principal components that best represent overall variance from a large number of correlated variables, and is used to reduce the dimension of data.

[0044] The calculation unit 505 executes calculation processes such as addition and subtraction between vectors, calculation of similarity between vectors, etc. The calculation unit 505 also performs various calculations such as calculation of similarity or deviation between a plurality of composite vectors of a document vector and each of a plurality of known technology vectors representing existing technologies and the plurality of known technology vectors.

[0045] The output unit 506 executes an output process using a plurality of known technology vectors representing existing technologies and a document vector to output search results including information on one or more technologies that can expand a new technological field by combining the document vector with the technology represented by the known technology vector. For example, the output unit 506 uses the calculation unit 505 to calculate a plurality of composite vectors between the document vector and each of the plurality of known technology vectors representing existing technologies, and similarities between the plurality of known technology vectors. The output unit 506 also outputs search results including information on technologies represented by known technology vectors corresponding to one or more composite vectors among the plurality of composite vectors whose calculated similarities are lower than those of the other composite vectors. The output unit 506 may, for example, transmit (output) the output results to another device such as the operation terminal 101 for display, or may output the search results to the storage unit 509 for storage.

[0046] The reception unit 507 executes a reception process for receiving a setting operation by a user by displaying an operation screen 200 as shown in Fig. 2(A) on the operation terminal 101, etc. For example, the user receives a file set in the file setting field 202 on the operation screen 200 as technical document data.

[0047] The data management unit 508 stores and manages various data and information such as acquired data 511, word embedding vectors 512, document vectors 513, and known technology vector group 514 in, for example, the storage unit 509. Here, the acquired data 511 includes multiple target document data acquired by the acquisition unit 501, or data such as features corresponding to multiple target document data. The word embedding vectors 512 include word embedding vectors generated by the word embedding vector generation unit 502. The document vectors 513 include document vectors generated by the document vector generation unit 503.

[0048] The known technology vector group 514 includes a plurality of vectors obtained by vectorizing document data of existing technologies. The known technology vector group 514 may be generated using the acquired data 511 acquired by the acquisition unit 501, or may be a group of vectors generated in advance and stored in the storage unit 509.

[0049] The data management unit 508 may store the above-mentioned data in a storage device, a storage server, a cloud service, or the like external to the information processing device 100. Here, as an example, the following description will be given assuming that the data management unit 508 stores and manages the above-mentioned data in the storage unit 509.

[0050] The storage unit 509 is realized by, for example, the program executed by the CPU 401 in FIG. 4, the HD 404, the HDD controller 405, etc., and stores various information, data, programs, etc., including each piece of data managed by the data management unit 508.

[0051] The functional configuration of the information processing device 100 shown in Fig. 5 is an example. For example, at least some of the functional configurations of the information processing device 100 shown in Fig. 5 may be provided outside the information processing device 100. Furthermore, the functional configurations of the information processing device 100 shown in Fig. 5 may be distributed and provided in multiple computers 400.

[0052] <Processing flow> Next, the processing flow of the information providing method according to this embodiment will be described.

[0053] (Processing of information processing device) 6 is a sequence diagram showing an example of processing by the information processing device according to an embodiment. This processing shows an overview of processing that the information processing device 100 executes when a user sets technical document data or the like on the operation screen 200 as shown in FIG. 2(A) and selects the "Execute" button 203.

[0054] In steps S601 and S602, when the receiving unit 507 receives a setting operation by the user on the operation screen 200 as shown in Fig. 2(A), the receiving unit 507 notifies the acquiring unit 501 of the received setting information. Note that this setting information includes, for example, technical document data set by the user or link information for acquiring the technical document data.

[0055] In step S603, the acquisition unit 501 acquires multiple pieces of target document data to be searched, for example, from one or more external data servers 103 or the local data server 102. For example, the acquisition unit 501 acquires document data related to existing technologies, such as patent documents, papers, or technical journals. If the setting information received from the reception unit 507 includes link information for acquiring the technical document data and the index document data, the acquisition unit 501 also acquires the technical document data and the index document data using the link information.

[0056] In step S604, the acquisition unit 501 stores in the storage unit 509 the target document data acquired from one or more external data servers 103 or the local data server 102, and the acquired technical document data and index document data.

[0057] In step S605, the acquisition unit 501 requests the word embedding vector generation unit 502 to generate a word embedding vector.

[0058] In steps S606 and S607, the word embedding vector generation unit 502 retrieves the target document data retrieved by the retrieval unit 501 from the storage unit 509, and generates word embedding vectors representing the retrieved target document data. In steps S608 and S609, the word embedding vector generation unit 502 stores the generated word embedding vectors in the storage unit 509, and sends a completion notification to the retrieval unit 501, which is the requesting source, indicating that the generation of the word embedding vectors has been completed.

[0059] In step S610, the acquisition unit 501 requests the document vector generation unit 503 to generate document vectors of distributed representations representing the technical document data.

[0060] In steps S611 and S612, the document vector generation unit 503 retrieves the word embedding vectors stored in the storage unit 509 and uses the retrieved word embedding vectors to generate document vectors of distributed representations representing the technical document data. In steps S613 and S614, the document vector generation unit 503 stores the generated document vectors in the storage unit 509 and sends a completion notification to the requesting acquisition unit 501 indicating that the generation of the document vectors is complete.

[0061] In step S615, the acquisition unit 501 transmits a search request requesting output of search results to the output unit 506. In response to this, in step S616, the output unit 506 executes a search process, which will be described later.

[0062] In the above example, the acquisition unit 501 controls the word embedding vector generation unit 502, the document vector generation unit 503, the output unit 506, etc., but this is just an example. For example, the information processing device 100 may have a separate control unit that controls the acquisition unit 501, the word embedding vector generation unit 502, the document vector generation unit 503, the output unit 506, etc. Alternatively, in the information processing device 100, the reception unit 507 may control the acquisition unit 501, the word embedding vector generation unit 502, the document vector generation unit 503, the output unit 506, etc.

[0063] Next, examples of search processing executed by the information processing device 100 will be described by illustrating a plurality of embodiments.

[0064] [First embodiment] 7 is a flowchart showing an example of search processing according to the first embodiment. Fig. 7(A) shows an example of search processing executed by the information processing device 100 in step S616 of Fig. 6, for example.

[0065] In steps S701 and S702, the calculation unit 505 acquires the known technology vector group 514, the document vector 513, and the like from the storage unit 509.

[0066] In step S703, the calculation unit 505 executes the processes of steps S704 and S705 for each of the acquired known technology vectors. In step S704, the calculation unit 505 generates a composite vector by combining (adding) the known technology vector and the document vector. In addition, in step S705, the calculation unit 505 calculates the average distance between the generated composite vector and the group of known technology vectors.

[0067] In step S706, the calculation unit 505 extracts one or more composite vectors whose average distance is greater than the other composite vectors. For example, the calculation unit 505 extracts a predetermined number of composite vectors in descending order of average distance and notifies the output unit 506.

[0068] In step S706, the output unit 506 outputs the search results for the new area based on the composite vectors extracted by the calculation unit 505. For example, the output unit 506 displays on the operation terminal 101 or the like a result display screen that lists information such as technical fields or technical document data corresponding to each of the predetermined number of composite vectors notified by the calculation unit 505 in descending order of average distance.

[0069] Fig. 7(B) shows another example of the search process executed by the information processing device 100, for example, in step S616 in Fig. 6. Among the processes shown in Fig. 7(B), the processes of steps S701, S702, and S707 are similar to the processes shown in Fig. 7(A), so the following description will focus on the differences from the processes shown in Fig. 7(A).

[0070] In step S711, the calculation unit 505 executes the processes of steps S712 and S713 for each of the acquired known-technology vectors. In step S712, the calculation unit 505 generates a composite vector by combining (adding) the known-technology vector and the document vector. In step S713, the calculation unit 505 calculates the average angle between the generated composite vector and the group of known-technology vectors.

[0071] In step S714, the calculation unit 505 extracts one or more resultant vectors whose average angles are larger than the other resultant vectors. For example, the calculation unit 505 extracts a predetermined number of resultant vectors in descending order of average angles and notifies the output unit 506.

[0072] 7A and 7B are examples of values ​​indicating the similarity (or deviation) between the composite vector and the known technology vector group. The values ​​indicating the similarity (or deviation) between the composite vector and the known technology vector group are not limited to the average distance and average angle, and may be values ​​such as cosine similarity.

[0073] The calculation unit 505 according to the first embodiment may calculate the similarity (or deviation) between the composite vector and the group of known technology vectors in each process of Figures 7(A) and (B), and extract one or more composite vectors that are not similar (deviant) from the group of known technology vectors.

[0074] By performing the processing shown in Figures 6 and 7 above, the information processing device 100 can provide an information processing device that can search for technical documents or technical fields that can expand unexplored technical areas by combining a specified technology with a plurality of existing technical documents.

[0075] [Second embodiment] 8 is a flowchart showing an example of search processing according to the second embodiment. This processing shows another example of the search processing executed by the information processing device 100 in step S616 of FIG.

[0076] In step S801, the acquisition unit 501 acquires a known technology vector and a document vector through the acquisition process shown in steps S701 and S702 in FIG.

[0077] In steps S802 and S803, the extraction unit 504 performs a principal component analysis on the group of known technology vectors (plurality of known technology vectors) and executes an extraction process to extract a predetermined number of principal component vectors.

[0078] In step S804, the calculation unit 505 executes the processes of steps S805 and S806 for each known technology vector. In step S805, the calculation unit 505 generates a composite vector by combining (adding) the known technology vector and the document vector. In step S806, the calculation unit 505 calculates the similarity (or deviation) between the generated composite vector and the principal component vector. Here, the similarity (or deviation) may be the average distance or average angle described with reference to FIGS. 7(A) and 7(B), or another similarity (for example, cosine similarity, etc.).

[0079] In step S807, the calculation unit 505 extracts one or more composite vectors whose similarity is lower than that of the other composite vectors (or whose deviation is higher than that of the other composite vectors). For example, the calculation unit 505 extracts a predetermined number of composite vectors in descending order of similarity and notifies the output unit 506.

[0080] In step S808, the output unit 506 executes an output process to output the search results for the new area based on the composite vectors extracted by the calculation unit 505. As an example, the output unit 506 displays on the operation terminal 101 or the like a result display screen that displays a list of information such as technical fields or document data corresponding to each of the predetermined number of composite vectors notified by the calculation unit 505 in order of highest similarity.

[0081] Through the above processing, the information processing device 100 can provide an information processing device that can search for technical documents or technical fields that can expand unexplored technical areas by combining a specified technology with multiple existing technical documents.

[0082] [Third embodiment] 9 is a flowchart showing an example of output processing according to the third embodiment. This processing shows another example of the output processing executed by the information processing device 100 in step S808 in FIG.

[0083] In step S901, the calculation unit 505 selects, for example, the composite vector with the lowest similarity from one or more composite vectors extracted in step S807 of FIG.

[0084] In step S902, the calculation unit 505 acquires an index word vector group. Here, the index word vector group is a vectorization of index words such as "fast," "clear," "safe," "beautiful," and "simple." Note that the index word vector group may be acquired from the storage unit 509 of the information processing device 100 or from an external server device such as the external data server 103.

[0085] In step S903, the calculation unit 505 executes the process of step S904 for each index word vector. In step S904, the calculation unit 505 calculates the average distance (an example of similarity) between the index word vector and the composite vector selected in step S901.

[0086] In step S905, the output unit 506 outputs search results including the word corresponding to the index word vector with the smallest average distance calculated in steps S904 and S904. For example, if the word corresponding to the index word vector with the smallest average distance is "conservation," the output unit 506 may display search results representing the unexplored technical field "conservation" 212 on the operation terminal 101, etc., as shown in FIG. 2(B).

[0087] [Fourth embodiment] Fig. 10 is a flowchart showing an example of search processing according to the fourth embodiment. This processing shows another example of search processing executed by the information processing device 100, for example, in step S616 in Fig. 6. Note that the processing of steps S801 to S803 is the same as the processing of steps S801 to S803 described in Fig. 8, and therefore description thereof will be omitted here.

[0088] In step S1001, the calculation unit 505 creates a complementary space of the extracted principal component vectors.

[0089] In step S1002, the calculation unit 505 executes the processes of steps S1003 to S1005 for each known technology vector. In step S1003, the calculation unit 505 generates a composite vector by combining (adding) the known technology vector and the document vector. In step S1004, the calculation unit 505 projects the generated composite vector into a complementary space. In step S1005, the calculation unit 505 calculates the vector length of the composite vector projected into the complementary space.

[0090] In step S1006, the calculation unit 505 extracts one or more composite vectors whose calculated vector length is longer than the other composite vectors.

[0091] In step S1007, the output unit 506 executes an output process to output the search results for the new area based on the one or more composite vectors extracted by the calculation unit 505. As an example, the output unit 506 displays on the operation terminal 101 or the like a result display screen that displays a list of information such as technical fields or document data corresponding to each of the one or more composite vectors notified from the calculation unit 505 in descending order of similarity.

[0092] Through the above processing, the information processing device 100 can provide an information processing device that can search for information on one or more technologies that can expand an unexplored technological field by combining them with a specified technology from multiple existing technical documents.

[0093] [Fifth embodiment] 2(B), for example, it is assumed that "storability" 212 is obtained as a search result for a technology that can expand an unexplored technological field by combining it with the set technology "back monitor" 211 through the search process described in Figures 7, 8, and 10. In this case, the information processing device 100 may be capable of further searching for a combination with an existing technology (for example, "back monitor + video recording" 213) that can expand the searched unexplored technological field "storability" 212 by combining it with the set technology "back monitor" 211.

[0094] 11 is a sequence diagram showing an outline of search processing according to the fifth embodiment. This processing shows another example of the search processing executed by the information processing device 100 in step S616 of FIG.

[0095] In step S1101, the information processing device 100 executes, for example, any one of the search processes described with reference to FIGS.

[0096] In step S1102, the calculation unit 505 determines an index vector representing the new region found in the search process of step S1101. This index vector is a vector of distributed representation representing an index (e.g., "preservation") representing the new region found in the search process of step S1101, and can be created, for example, by the same method as for document vectors. The calculation unit 505 may acquire the index vector representing the new region from multiple index vectors stored in advance in the storage unit 509, or may create the index vector representing the new region using the document vector generation unit 503.

[0097] In step S1104, the calculation unit 505 executes the processes of steps S1105 and S1106 for each known technology vector. In step S1105, the calculation unit 505 combines the known vector and the document vector to generate a combined vector. In step S1106, the similarity between the generated combined vector and the index vector determined in step S1102 is calculated.

[0098] In step S1107, the calculation unit 505 extracts one or more composite vectors whose similarity is higher than that of the other composite vectors, and notifies the output unit 506 of the extracted composite vectors.

[0099] In step S1108, the output unit 506 outputs search results including information about technologies corresponding to one or more composite vectors extracted by the calculation unit 505. For example, if the extraction unit 504 extracts the composite vector "back monitor + video shooting," the output unit 506 may display, on the result display screen, for example, vector 213 representing "back monitor + video shooting," as shown in FIG. 2(B). Alternatively, the output unit 506 may display, on the result display screen, a list of document data corresponding to the multiple composite vectors extracted by the calculation unit 505, or character strings indicating technical fields, in order of similarity.

[0100] [Sixth embodiment] The search process described in the fifth embodiment can be performed using multiple index vectors, such as a first index vector representing a new area searched for in the search process and a second index vector representing an index set by the user.

[0101] Fig. 12 is a flowchart showing an example of search processing according to the sixth embodiment. This processing shows another example of the search processing executed by the information processing device 100 in step S616 in Fig. 6. Note that detailed description of processing similar to that in Fig. 11 will be omitted here.

[0102] In step S1201, the information processing device 100 executes, for example, any of the search processes described with reference to FIGS.

[0103] In step S1202, the calculation unit 505 determines a first index vector representing the new region found in the search process of step S1101.

[0104] In step S1203, the calculation unit 505 obtains (or creates) a second index vector representing the index set by the user. The second index vector is an index vector of the distributed representation created by the document vector generation unit 503 when the index is set by, for example, the user's input operation on the setting screen.

[0105] In steps S1204 and S1205, if weighting of the first and second index vectors has been set, the calculation unit 505 weights the first and second index vectors.

[0106] In step S1206, the calculation unit 505 performs the processes of steps S1107 and S1108 on each of the known technical vectors. In step S1105, the calculation unit 505 combines the known vector and the document vector to generate a composite vector. In step S1106, the calculation unit 505 calculates the similarity between the generated composite vector and the first and second index vectors. For example, if the first and second index vectors are weighted, the calculation unit 505 calculates the similarity between the composite vector and a vector obtained by weighting the first and second index vectors. Furthermore, if the first and second index vectors are not weighted, the calculation unit 505 calculates the similarity between the composite vector and a vector obtained by averaging the first and second index vectors.

[0107] In step S1209, calculation unit 505 extracts one or more composite vectors whose similarity is higher than that of the other composite vectors, and notifies output unit 506 of the extracted composite vectors.

[0108] In step S1210, the output unit 506 outputs the search results including the information on the technology corresponding to the one or more composite vectors extracted by the calculation unit 505.

[0109] [Seventh embodiment] In the seventh embodiment, an example of processing will be described in which an index approximate word group is used to specify (limit) expressions (words, sentences, etc.) of similar targets when verbalizing search results.

[0110] 13 is a flowchart showing an example of search processing according to the seventh embodiment. This processing shows another example of the search processing executed by the information processing device 100 in step S616 of FIG.

[0111] In steps S1301 and S1302, the calculation unit 505 acquires the known technology vector group 514, the document vector 513, and the like from the storage unit 509.

[0112] In step S1303, the calculation unit 505 acquires a group of index approximate words from, for example, the storage unit 509. This group of index approximate words is used to limit the words used in the search results. If the search results are approximated to words using the entire word embedding vector, they may be converted into unique expressions in technical documents or words that do not make sense as single words. Therefore, by using this group of index approximate words, the words that can be used to output the search results are limited.

[0113] In step S1304, the calculation unit 505 executes the processes of steps S1305 and S1306 for each of the acquired known technology vectors. In step S1305, the calculation unit 505 generates a composite vector by combining (adding) the known technology vector and the document vector. In addition, in step S1306, the calculation unit 505 calculates the degree of dissociation (or similarity) between the generated composite vector and the group of known technology vectors.

[0114] In step S1307, calculation unit 505 extracts one or more composite vectors whose degree of dissociation is greater than (or whose degree of similarity is smaller than) the other composite vectors.

[0115] In step S1308, the calculation unit 505 (or the output unit 506) approximates one or more extracted composite vectors to words (or sentences) using the index approximate word group. In step S1309, the output unit 506 displays the search results for the new area using the approximated words (or sentences).

[0116] Through the above processing, the information processing device 100 can output search results using more appropriate words.

[0117] [Eighth embodiment] The information processing device 100 may have a function of downloading target document data displayed on the result display screen in response to a download operation on the result display screen displayed by the output unit 506. Fig. 14 is a sequence diagram showing an example of document data download processing according to the eighth embodiment.

[0118] In step S1401, the information processing device 100 may execute the search process shown in FIGS. 11 and 12, for example, and then execute the processes in step S1402 and thereafter.

[0119] In step S1402, the output unit 506 displays, as an example, a result display screen 1510 as shown in Fig. 15(B). In the example of Fig. 15(B), the result display screen 1510 displays a list of multiple relevant documents (target document data) 1511 extracted by the search process in step S1401 in descending order of relevance (for example, similarity). The user can request downloading of document data of the selected document name by selecting a display element 1512 displaying the document name from the multiple relevant documents 1511.

[0120] In steps S1403 and S1404, upon receiving a document data download request from the user, the receiving unit 507 transmits to the output unit 506 a download request including document information of the requested document data (hereinafter referred to as requested document data).

[0121] If the requested document data requested in the download request is a free document that can be used free of charge, the output unit 506 executes process 1410 for the free document. If the requested document data requested in the download request is a paid document that can be used for a fee, the output unit 506 executes process 1420 for the paid document. If the document data requested in the download request includes both free and paid documents, the output unit 506 executes process 1410 for the free document and process 1420 for the paid document.

[0122] If the requested document data requested in the download request includes a free document, the output unit 506 acquires the requested document data from the storage unit 509 or the like in step S1411.

[0123] If the document data requested in the download request includes a paid document, output unit 506 executes the processes of steps S1421 to S1423.

[0124] In step S1421, the output unit 506 checks whether the user has a usage right, such as a fixed-price usage right, to use the requested paid document. If the user does not have the usage right, the output unit 506 executes the processes of steps S1422 and S1423. On the other hand, if the user has the usage right, the output unit 506 executes the process of step S1423.

[0125] In step S1422, the output unit 506 performs the billing process using, for example, the reception unit 507. For example, the reception unit 507 displays a purchase screen for document data and receives payment operations by the user.

[0126] In step S1423, the output unit 506 acquires the request document data from the storage unit 509 or the like.

[0127] In step S1431, the output unit 506 transmits the acquired request document data to, for example, the operation terminal 101 or the like.

[0128] Through the above process, the user can easily obtain document data searched by the information processing system 1.

[0129] <Example of display screen> Next, an example of a display screen that the information processing device 100 displays on the operation terminal 101 or the like will be described.

[0130] (Example of the results display screen) Fig. 15 is a diagram showing an example of a result display screen according to an embodiment. Fig. 15(A) shows an example of a result display screen that the information processing device 100 displays on the operation terminal 101 or the like in the first to fifth embodiments, for example. In the example of Fig. 15(A), a result display screen 1500 displays a list of multiple relevant documents (for example, text data names, etc.) 1501 extracted by the search process in descending order of novelty (for example, descending order of similarity).

[0131] 15(B) shows another example of a result display screen that the information processing device 100 displays on the operation terminal 101 or the like. In the example of FIG. 15(B), as described above, a result display screen 1510 displays a list of multiple relevant documents (e.g., text data names, etc.) 1511 extracted by the search process in descending order of relevance (e.g., similarity). A user can request downloading of document data of a selected document name by selecting a display element 1512 displaying a document name from the multiple relevant documents 1511.

[0132] FIG. 15(C) shows another example of a result display screen displayed by the information processing device 100 on the operation terminal 101 or the like. The result display screen 1520 visually displays a combination of technologies, such as "back monitor + video recording," in a graph with the index "storability" on the horizontal axis and the index "continuous operation" on the vertical axis, to indicate the similarity between the two indexes. In this way, the information processing device 100 may perform a search process for multiple indexes. Furthermore, the information processing device 100 may display, for example, search results such as those shown in FIG. 2(B) on the result display screen. In this way, the result display screen output by the output unit 506 can be output in various formats.

[0133] (Example of settings screen) Fig. 16 is a diagram showing an example of a setting screen according to an embodiment. The accepting unit 507 may display, for example, a setting screen 1600 as shown in Fig. 16 and accept various setting information through user operations.

[0134] 16, a setting screen 1600 has a setting field 1601 for setting technical document data to be searched, as well as a setting field 1602 for setting an exclusive known technical document group. For example, if a user wants to prevent unexpected and unrelated technical fields from being output in the search results, the user can exclude the specified range from the search results by specifying a file or dataset that specifies the area. Alternatively, the information processing device 100 may output search results within the range of an area specified by the user.

[0135] The setting screen 1600 also has a setting field 1603 where the number of existing technologies to be combined can be set. For example, when "3" is set in the setting field 1603 where the number of existing technologies to be combined can be set, the information processing device 100 performs evaluation by combining up to three vectors, including the set technical document data and two known technology vectors. The setting screen 1600 also has a setting field 1604 where the number of indicators to be detected can be set.

[0136] As described above, according to each embodiment of the present invention, it is possible to provide an information processing device 100 that can search for information on one or more technologies that can expand unexplored technological fields by combining them with a specified technology from multiple existing technical documents.

[0137] <Supplementary information> Each function of each embodiment described above can be realized by one or more processing circuits. Here, the term "processing circuit" in this specification includes a processor programmed to perform each function by software, such as a processor implemented by an electronic circuit, as well as devices such as an ASIC (Application Specific Integrated Circuit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), and conventional circuit modules designed to perform each of the functions described above.

[0138] Furthermore, the devices described in the examples represent only one of multiple computing environments for implementing the embodiments disclosed herein. In one embodiment, information processing device 100 includes multiple computing devices, such as a server cluster. The multiple computing devices are configured to communicate with each other via any type of communication link, including a network or shared memory, and perform the processes disclosed herein. Furthermore, each element of information processing device 100 may be integrated into a single information processing device or may be separated into multiple information processing devices.

[0139] 5 can be configured to share, for example, various combinations of the processes of the information processing device 100 shown in FIGS. 6 to 14. For example, at least a portion of the processes executed by the information processing device 100 may be executed by the operation terminal 101. Furthermore, at least a portion of the processes executed by the information processing device 100 may be executed by, for example, an external server device, a cloud service, or the like. [Explanation of symbols]

[0140] 1. Information Processing Systems 100 Information processing device 502 Word Embedding Vector Generation Unit 503 Document Vector Generation Unit 504 Extraction part 505 Arithmetic unit 506 Output section 507 Reception [Prior art documents] [Patent documents]

[0141] [Patent Document 1] Japanese Patent Application Laid-Open No. 2001-265808

Claims

1. a reception unit that receives a setting operation for setting technical document data; a document vector generation unit that generates document vectors of distributed representations representing the technical document data; a calculation unit that calculates the similarity between a plurality of composite vectors obtained by combining each of a plurality of known technology vectors representing existing technologies with the document vector, and the known technology vector; an output unit that outputs search results including information about technologies represented by the known technology vectors corresponding to one or more composite vectors among the plurality of composite vectors whose similarity is lower than that of the other composite vectors; An information processing device having the above.

2. a reception unit that receives a setting operation for setting technical document data; a document vector generation unit that generates document vectors of distributed representations representing the technical document data; an extraction unit that performs principal component analysis on a plurality of known technology vectors representing existing technologies to extract a predetermined number of principal component vectors; a calculation unit that calculates the similarity between a plurality of composite vectors obtained by combining the document vector and each of the known technology vectors and the principal component vector; an output unit that outputs search results including information on technologies represented by the known technology vectors corresponding to one or more composite vectors among the plurality of composite vectors whose similarity is lower than that of the other composite vectors; An information processing device having the above.

3. The information processing device according to claim 1 , wherein the output unit displays the search results in a list in order of similarity.

4. a reception unit that receives a setting operation for setting technical document data; a word embedding vector generation unit that generates word embedding vectors, which are vectors of distributed representations representing a plurality of target document data sets to be searched, using the plurality of target document data sets; a document vector generation unit that generates document vectors of distributed representations representing the technical document data; a calculation unit that creates a complementary space of the principal component vectors of the word embedding vector, projects a plurality of composite vectors of the document vector and each of known technology vectors representing existing technologies onto the complementary space, and calculates the vector length in the complementary space; an output unit that outputs search results including information about technologies represented by the known technology vectors corresponding to one or more composite vectors among the plurality of composite vectors, the vector length of which in the complementary space is longer than the other composite vectors; An information processing device having the above.

5. The information processing device according to claim 1 , wherein a technical area specified by a file or a data set is excluded from the search results.

6. The information processing device according to claim 1 , wherein the output unit outputs the search results by approximating existing words or sentences using embedding vectors of the words.

7. The information processing device A reception process for receiving a setting operation for setting technical document data; a document vector generation process for generating document vectors of distributed representations representing the technical document data; A computation process for calculating the similarity between a plurality of composite vectors obtained by combining each of a plurality of known technology vectors representing existing technologies with the document vector, and the known technology vector; and an output process for outputting search results including information on technologies represented by the known technology vectors corresponding to one or more composite vectors among the plurality of composite vectors, the similarity of which is lower than that of the other composite vectors; A method of providing information.

8. The information processing device A reception process for receiving a setting operation for setting technical document data; a document vector generation process for generating document vectors of distributed representations representing the technical document data; an extraction process of extracting a predetermined number of principal component vectors by performing principal component analysis on a plurality of known technology vectors representing existing technologies; a computation process for calculating the similarity between a plurality of composite vectors obtained by combining the document vector and each of the known technology vectors and the principal component vector; and an output process for outputting search results including information on technologies represented by the known technology vectors corresponding to one or more composite vectors among the plurality of composite vectors, the similarity of which is lower than that of the other composite vectors; A method of providing information.

9. A reception process for receiving a setting operation for setting technical document data; a word embedding vector generation process for generating word embedding vectors, which are vectors of distributed representations representing a plurality of target document data sets to be searched, using the plurality of target document data sets; a document vector generation process for generating document vectors of distributed representations representing the technical document data; A computation process of creating a complementary space of the principal component vectors of the word embedding vector, projecting a plurality of composite vectors of the document vector and each of known technology vectors representing existing technologies onto the complementary space, and calculating the vector length in the complementary space; and an output process for outputting search results including information on technologies represented by the known technology vectors corresponding to one or more composite vectors among the plurality of composite vectors, the vector length of which in the complementary space is longer than the other composite vectors; A method of providing information.

10. A program that causes a computer to execute the information providing method according to any one of claims 7 to 9.

Citation Information

Patent Citations

  • System and method for information retrieval

    JP2001265808A

  • Idea generation support device and idea generation support method

    JP2005070823A

  • Technique / intellectual property evaluating device, and technique / intellectual property evaluating method

    JP2005326897A

  • Apparatus and program for supporting systematization of idea

    JP2011204009A

  • Information recommendation device, information recommendation method, and program

    JP2011210090A