Information processing system

The information processing system addresses cumbersome data set evaluation by using similarity-based data search, labeling, and display tools to efficiently identify desired information within data sets.

JP2025130417APending Publication Date: 2025-09-08RICOH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024027569
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-27
Publication Date
2025-09-08

AI Technical Summary

Technical Problem

Conventional data classification systems require cumbersome processes to determine if a data set contains desired information, necessitating manual referencing of data sets.

Method used

An information processing system equipped with a data search unit for similarity-based data retrieval, a label assignment unit for data labeling, and a display information generation unit to associate user-assigned names with labels, facilitating efficient data set evaluation.

Benefits of technology

The system assists in quickly determining whether a data set contains desired information, enhancing user efficiency in data retrieval and classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025130417000001_ABST
    Figure 2025130417000001_ABST
Patent Text Reader

Abstract

To assist in determining whether a data set contains desired information.SOLUTION: An information processing system comprises: a data retrieval unit for retrieving a plurality of pieces of data based on similarity with input information; a labeling unit for assigning a label to a data set based on words contained in the set; and a display information generating unit for generating display information that displays a name given by a user to the set in association with the label.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing system. [Background technology]

[0002] Researchers and people working in and outside the office (hereinafter referred to as "users") sometimes want to collect and utilize data that is useful for their work. In order to efficiently utilize the collected data, it is possible to classify the data into highly related data sets using virtual storage locations such as folders. In this case, each data set (storage location) can be given a name according to its classification.

[0003] Patent Document 1 discloses a configuration for automatically classifying images by classifying scanned images into preset tags. Summary of the Invention [Problem to be solved by the invention]

[0004] However, in conventional techniques, determining whether a certain data set contains the information the user is looking for requires cumbersome work such as referencing data belonging to the data set.

[0005] The present invention has been made in view of the above points, and has as its object to support a decision as to whether or not a certain data set contains desired information. [Means for solving the problem]

[0006] In order to solve the above problem, the information processing system has a data search unit that searches for multiple data based on similarity with input information, a label assignment unit that assigns a label to the set of data based on words contained in the set, and a display information generation unit that generates display information that associates the name assigned to the set by the user with the label. [Effects of the Invention]

[0007] It can assist in determining whether a data set contains desired information. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 illustrates an example of a configuration of an information processing system according to a first embodiment. [Figure 2] 1 is a diagram illustrating an example of a hardware configuration of an information processing device 10 according to a first embodiment. [Figure 3] FIG. 2 is a diagram illustrating an example of a functional configuration of an information processing system according to a first embodiment. [Figure 4] 10 is a flowchart illustrating an example of a processing procedure of an information collection process. [Figure 5] FIG. 10 is a diagram illustrating an example of a collection condition input screen. [Figure 6] FIG. 2 is a diagram illustrating an example of the configuration of a document vector storage unit 141. [Figure 7] 10 is a flowchart illustrating an example of a procedure for document collection result output processing. [Figure 8] FIG. 2 is a diagram illustrating an example of the configuration of a document information storage unit 22. [Figure 9] FIG. 10 is a diagram showing an example of a sorting result of document information. [Figure 10] FIG. 10 is a diagram showing a display example of a search result screen. [Figure 11] FIG. 10 is a diagram illustrating an example of a query relation diagram. [Figure 12] FIG. 2 is a diagram illustrating an example of the configuration of a document-related storage unit 142. [Figure 13] FIG. 10 is a diagram showing a display example of a document details screen. [Figure 14] FIG. 10 is a diagram illustrating an example of a document relationship diagram. [Figure 15] 10 is a flowchart illustrating an example of a processing procedure for generating a workspace. [Figure 16] FIG. 2 is a diagram illustrating an example of the configuration of a workspace storage unit 25. [Figure 17] FIG. 10 is a diagram showing a second display example of the search result screen. [Figure 18] 10 is a flowchart illustrating an example of a processing procedure for an expert collection result output process. [Figure 19] FIG. 10 is a diagram illustrating an example of a sorting result of experts. [Figure 20] 10 is a flowchart illustrating an example of a processing procedure for workspace collection result output processing. [Figure 21] FIG. 10 is a diagram illustrating an example of a sorting result of a workspace. [Figure 22] FIG. 10 is a diagram illustrating a display example of a workspace collection result screen. [Figure 23] FIG. 10 is a diagram illustrating a display example of a workspace details screen. [Figure 24] 10 is a flowchart illustrating an example of a processing procedure executed by an information processing system according to a second embodiment. [Figure 25] FIG. 10 is a diagram illustrating a display example of a workspace list screen. [Figure 26] FIG. 10 is a diagram for explaining the desirable number of labels. [Figure 27] 13 is a flowchart illustrating an example of a processing procedure executed by an information processing system according to a third embodiment. [Figure 28] 13 is a flowchart illustrating an example of a processing procedure executed by an information processing system according to a fourth embodiment. [Figure 29] 13 is a flowchart illustrating an example of a processing procedure executed in response to a workspace list request in the fourth embodiment. [Figure 30] 13 is a flowchart illustrating an example of a processing procedure executed by an information processing system according to a fifth embodiment. [Figure 31] 13 is a flowchart illustrating an example of a processing procedure executed by an information processing system according to a sixth embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Fig. 1 is a diagram showing an example of the configuration of an information processing system in a first embodiment. In Fig. 1, the information processing system includes an information management device 20, an information processing device 10, and one or more user terminals 30. The information processing device 10 is connected to the information management device 20 via a network N1. The user terminal 30 is connected to the information management device 20 via a network N2, and is connected to the information processing device 10 via a network N3.

[0010] The user terminal 30 is a terminal used by a user who desires to collect (access) certain information. For example, a PC (Personal Computer), a tablet terminal, a smartphone, or the like may be used as the user terminal 30. In this embodiment, document information, expert information, and workspaces are given as examples of types of information that a user desires to collect.

[0011] Document information is information including attribute information or bibliographic information related to electronic data in which a document is recorded (hereinafter referred to as "document data"). A document is a collection of one or more words or sentences (and, of course, may include alphanumeric characters and other multilingual characters). Document data may be in any format that can represent a sentence. For example, document data may be data that represents a document in text format, or data in a format specialized for a specific application. Alternatively, document data may be data that represents words or sentences themselves or concepts corresponding to words or sentences using images, audio, or video (video). In other words, document data may be image data, audio data, or video data. Furthermore, the storage format of document data is not limited to a specific one. For example, document data may be stored in a file, as a database record, or in another format.

[0012] Expert information is information about a person (hereinafter referred to as an "expert") who is presumed to know certain information (or be knowledgeable about certain information).

[0013] A workspace is information that indicates the results of information collection (search results) conducted in the past using an information processing system, or information that has been edited from the collection results. Alternatively, a workspace can be said to be an example of a collection (data set) of all or part of the data (document data) of the search results. In this embodiment, the information that a user wishes to collect is called "knowledge."

[0014] When document information relating to certain knowledge is collected, a user can obtain the desired knowledge by, for example, viewing document data relating to the document information.

[0015] When expert information relating to an expert who is knowledgeable about a certain knowledge is collected, a user can obtain desired knowledge from the expert by, for example, accessing the expert.

[0016] When a workspace related to certain knowledge (such as the results of past collection of information by other users or edited data thereof) is collected, a user can obtain desired knowledge based on the workspace.

[0017] The information management device 20 is one or more computers that store information to be collected (document information, expert information, and workspaces).

[0018] The information processing device 10 is one or more computers that collect information that matches a collection condition from the information management device 20 based on the information collection condition input by a user.

[0019] The information management device 20 and the information processing device 10 may be realized using the same computer. In this case, the network N1 corresponds to a signal line such as a bus within the computer that constitutes the information management device 20 and the information processing device 10. Alternatively, each user terminal 30 may also function as the information processing device 10. In this case, the network N3 corresponds to a signal line such as a bus within the user terminal 30.

[0020] The scene (situation) in which the information processing system is used is not limited to a specific form, but may be used within a company, for example. That is, each employee of a company (including not only companies but also government agencies, various organizations, unions, etc., and not only full-time employees but also temporary workers, part-time workers, casual workers, etc.) may be a user (in this embodiment, each employee of a company is described as a user, but this is not limited to this, and the information processing system can also be applied to a case where the information processing system is used by general users).

[0021] In this case, the information management device 20 is a group of computers that manage various types of information within the company. For example, the information management device 20 manages document information related to various document data created within the company, information related to the organizational structure of the company, information related to each employee within the company, and workspaces resulting from information collection within the company. The information management device 20 may also manage business-related electronic communications (email, chat, etc.) between employees within the company. In this case, the network N2 corresponds to, for example, a wide area network (WAN) or local area network (LAN) within the company.

[0022] The information processing device 10 may be installed within a company, or may be installed outside the company (in a cloud environment (e.g., a data center) connected to an in-company network via the Internet). When the information processing device 10 is installed within a company, the network N1 and the network N3 correspond to, for example, a wide area network (WAN) or local area network (LAN) within the company. When the information processing device 10 is installed within a company, the network N1 and the network N3 correspond to, for example, the Internet. Note that the information processing device 10 may collect information desired by the user from information made public outside the company.

[0023] 2 is a diagram showing an example of the hardware configuration of the information processing device 10 according to the first embodiment. The information processing device 10 in FIG. 2 includes a drive device 100, an auxiliary storage device 102, a memory device 103, a processor 104, and an interface device 105, which are all interconnected via a bus B.

[0024] A program for realizing processing in the information processing device 10 is provided by a recording medium 101 such as a CD-ROM. When the recording medium 101 storing the program is set in the drive device 100, the program is installed from the recording medium 101 to the auxiliary storage device 102 via the drive device 100. However, the program does not necessarily have to be installed from the recording medium 101, but may be downloaded from another computer via a network. The auxiliary storage device 102 stores the installed program as well as necessary files, data, etc.

[0025] When an instruction to start a program is received, the memory device 103 reads and stores the program from the auxiliary storage device 102. The processor 104 is a CPU or a GPU (Graphics Processing Unit), or a CPU and a GPU, and executes functions related to the information processing device 10 in accordance with the program stored in the memory device 103. The interface device 105 is used as an interface for connecting to a network.

[0026] The information management device 20 and the user terminal 30 may also have the same hardware configuration as that shown in FIG.

[0027] Fig. 3 is a diagram showing an example of the functional configuration of the information processing system according to the first embodiment. In Fig. 3, the user terminal 30 has a display control unit 31. The display control unit 31 is realized by processing that one or more programs (e.g., a web browser program) installed in the user terminal 30 cause a processor of the user terminal 30 to execute.

[0028] The display control unit 31 displays a screen based on display information transmitted from the information processing device 10, and transmits to the information processing device 10 a request in response to an input on the screen.

[0029] The information management device 20 has a document management unit 21. The document management unit 21 is realized by processing in which one or more programs installed in the information management device 20 are executed by a processor of the information management device 20. The information management device 20 also uses a document information storage unit 22, an employee information storage unit 23, an organization information storage unit 24, a workspace storage unit 25, and the like. Each of these storage units can be realized, for example, by using an auxiliary storage device of the information management device 20, or a storage device connectable to the information management device 20 via a network.

[0030] The document management unit 21 performs registration, update, deletion, etc. on a plurality of pieces of document information stored in the document information storage unit 22 .

[0031] The employee information storage unit 23 stores attribute information and the like (hereinafter referred to as "employee information") of each employee of a company (hereinafter referred to as "company X") that uses the information management device 20.

[0032] The organizational information storage unit 24 stores information (hereinafter referred to as "organizational information") that represents the organizational structure of company X. For example, the organizational information may be information that represents the organizational structure in the form of a graph in which each organization is a node and hierarchical relationships (parent-child relationships) between organizations are branches.

[0033] The workspace storage unit 25 stores information about workspaces. For example, as described above, a workspace is information (hereinafter referred to as "collection results" or "search results") that accepts edits by a user to a collection result of certain information (e.g., document information) or a workspace, and reflects (stores) the content of the edits in the workspace storage unit 25. Therefore, information about a certain workspace is, for example, information that associates the workspace with document information included in the collection result that corresponds to the workspace.

[0034] The information processing device 10 includes a reception unit 121, a vector conversion unit 122, a comparison unit 123, a data search unit 124, a classification unit 125, a label assignment unit 126, a relational diagram generation unit 127, an expert collection unit 128, a workspace collection unit 129, a display information generation unit 130, a display information transmission unit 131, a workspace generation unit 132, and a workspace editing unit 133. Each of these units is realized by a process executed by the processor 104 of one or more programs installed in the information processing device 10. The information processing device 10 also uses a document vector storage unit 141, a document relation storage unit 142, and the like. Each of these storage units can be realized using, for example, the auxiliary storage device 102, or a storage device connectable to the information processing device 10 via a network.

[0035] The reception unit 121 receives (accepts) a request to collect information desired by a user from the user terminal 30. The request to collect information includes conditions (collection conditions) related to the collection of information. The collection conditions include a type of information to be collected (hereinafter referred to as "information type") and a character string expressing the information to be collected in natural language (hereinafter referred to as "query"). The query is an example of input information.

[0036] In this embodiment, the options for the information type are, for example, "document," "expert," and "workspace." "Document" is an information type corresponding to document information. "Expert" is an information type corresponding to expert information. "Workspace" is an information type corresponding to a workspace.

[0037] A query is, for example, a set of one or more words. A query may be a list of one or more words, or may have the form of one or more sentences.

[0038] The vector conversion unit 122 analyzes a query included in the collection conditions and document data related to each piece of document information stored in the document information storage unit 22, and converts the query or document data into vector-format data (hereinafter simply referred to as a "vector"). A vector is also called a distributed representation or an embedded representation, and is a representation corresponding to the meaning contained in the source data (query, document data, etc.). For example, the vector conversion unit 122 generates a vector using natural language processing such as BERT. The BERT model may be switched using user attributes. The vector conversion unit 122 generates vectors for each piece of document data in advance and stores them in the document vector storage unit 141. Hereinafter, a vector based on a query will be referred to as a "query vector," and a vector based on document data will be referred to as a "document vector."

[0039] The comparison unit 123 compares the query vector with each document vector and evaluates the similarity of each document vector with the query vector. In this embodiment, the index for evaluating the similarity is called "similarity."

[0040] The comparison unit 123 also calculates the similarity between document vectors for all pairs of two document vectors, and records the similarity for each pair of document vectors in the document association storage unit 142.

[0041] The data search unit 124 extracts (collects) document information (document data) related to the query based on the similarity for each document vector, which is the result of the comparison between the query vector and the document vector by the comparison unit 123 (i.e., based on the similarity between the query and the document data).

[0042] The "comparison" process performed by the comparison unit 123 may be referred to as "search," and the data search unit 124 may treat the above-mentioned process as the search results by the comparison unit 123. In this case, the collection of information may be referred to as information search or simply search.

[0043] The classification unit 125 classifies the document information (document data) extracted by the data search unit 124 based on each document vector. Clustering, for example, is used for classification. A group of document data after classification is called a "class."

[0044] The labeling unit 126 assigns labels to classes and workspaces. The labeling unit 126 also assigns labels to each piece of document data in advance based on the content of each piece of document data (words contained in each piece of document data). The results of assigning labels to each piece of document data are recorded in the document information storage unit 22. In this embodiment, a label refers to a character string (e.g., "word") that (succinctly) indicates the characteristics of an object to which a label is to be assigned.

[0045] The relationship diagram generation unit 127 generates a relationship diagram, which is a graphic showing the relationship between document data, classes, and queries, based on the classification results by the classification unit 125 and the labeling results by the labeling unit 126. The relationship diagram generation unit 127 also generates a relationship diagram showing the relationship between certain document data and other document data, based on each document vector.

[0046] The expert collection unit 128 extracts (collects) people (employees, external experts, etc.) associated with document data related to the query as experts based on the comparison result by the comparison unit 123. A person associated with certain document data is, for example, a person who created or updated the document data.

[0047] The workspace collection unit 129 extracts (collects) workspaces related to document data related to the query based on the comparison result by the comparison unit 123. A workspace related to certain document data refers to, for example, a collection result including document information related to the document data or a workspace corresponding to the edited data thereof.

[0048] The display information generation unit 130 generates display information to be displayed on the user terminal 30. For example, the display information generation unit 130 generates display information related to the processing results of the data search unit 124, the expert collection unit 128, and the workspace collection unit 129, or generates display information that displays the association diagram generated by the association diagram generation unit 127. For example, if the display control unit 31 of the user terminal 30 is realized by a web browser, a web page is an example of the display information. However, the display information may be generated in other formats.

[0049] The display information transmitting unit 131 transmits the display information generated by the display information generating unit 130 to the user terminal 30.

[0050] In response to a user's instruction regarding the document information collection result, workspace generation unit 132 generates a workspace related to the collection result and stores the workspace in workspace storage unit 25.

[0051] The workspace editing unit 133 accepts edits made by the user to the workspace, and reflects the edited content in the workspace storage unit 25 .

[0052] 3 is merely an example. The device where each unit is located may be changed to any one of the user terminal 30, the information processing device 10, or the information management device 20, as appropriate.

[0053] The processing procedure executed by the information processing system will be described below: Fig. 4 is a flowchart for explaining an example of the processing procedure of the information collection processing.

[0054] In step S101, the display control unit 31 of the user terminal 30 accepts input of collection conditions from the user via a collection condition input screen displayed on the display device of the user terminal 30.

[0055] FIG. 5 is a diagram showing an example of a collection condition input screen. As shown in FIG. 5, the collection condition input screen 510 includes an information type selection area 511, a query input area 512, an execute button 513, and the like. The information type selection area 511 is an area for accepting the selection of an information type. In this embodiment, the options for the information type are "document," "expert," and "workspace," and therefore the information type selection area 511 may be a list box including options corresponding to "document," "expert," and "workspace." The example in FIG. 5 shows an example in which "document" has been selected.

[0056] The query input area 512 is an area for receiving a query input. The query may be input using a keyboard or the like of the user terminal 30 (including direct input using a touch panel), or may be input by voice via a microphone of the user terminal 30.

[0057] The execution button 513 is a button for accepting an instruction to execute information collection (execute search).

[0058] The collection condition input screen 510 may be displayed on the user terminal 30, for example, in response to a user logging in to the information processing device 10. Hereinafter, a user who inputs collection conditions (search conditions) will be referred to as a "logged-in user."

[0059] After an information type is selected and a query is entered, when the login user presses the execute button 513, the display control unit 31 sends an information collection request to the information processing device 10, which includes the selected information type and the entered query as information collection conditions.

[0060] When the reception unit 121 of the information processing device 10 receives an information collection request, the vector conversion unit 122 converts the query (hereinafter referred to as the "target query") included in the information collection request (hereinafter referred to as the "target collection request") into a query vector (S102).

[0061] Next, the comparison unit 123 compares the query vector with the document vector corresponding to each piece of document data related to the document information managed by the information management device 20, and calculates the similarity between the query vector and the document vector (S103). The document vector corresponding to each piece of document data managed by the information management device 20 is stored in the document vector storage unit 141.

[0062] FIG. 6 is a diagram illustrating an example of the configuration of the document vector storage unit 141. As shown in FIG. 6, the document vector storage unit 141 stores a document ID, a document name, and a document vector for each piece of document data. The document ID is identification information for document information related to the document data, and associates the document information in the information management device 20 with the document vector in the document vector storage unit 141. The document name is the name or title of the document data. For example, if the document data is saved in a file format, the file name may be used as the document name. Like the query vector, the document vector is a vector representation (for example, a distributed representation or an embedded representation) according to the meaning of the content of the document data.

[0063] The similarity between a query vector and a document vector can be calculated using the angle (cosine similarity) or distance between the query vector and the document vector, similar to the calculation of the similarity between general vectors. For example, when using cosine similarity, the cosine similarity between vector a and vector b can be calculated based on the following formula:

[0064]

number

[0065] Next, the information processing device 10 branches the process depending on the information type of the target collection condition (hereinafter referred to as "target information type") (S105). If the target information type is "document", the information processing device 10 executes document collection result (document search result) output process (S106). If the target information type is "expert", the information processing device 10 executes expert collection result (expert search result) output process (S107). If the target information type is "workspace", the information processing device 10 executes workspace collection result (workspace search result) output process (S108).

[0066] Next, details of step S106 will be described. Fig. 7 is a flowchart for explaining an example of a processing procedure of the document collection result output processing.

[0067] In step S201, the data search unit 124 acquires (extracts) document information of the top N document data from the document information storage unit 22 based on the document IDs of the top N document vectors in terms of similarity.

[0068] Fig. 8 is a diagram showing an example of the configuration of the document information storage unit 22. As shown in Fig. 8, the document information storage unit 22 stores one or more records including a document ID, a document name, a creator, an update history, a file path, a summary, access control information, a label list, etc. One record corresponds to one piece of document information.

[0069] The document ID and document name are as described above. The document ID and document name for the same document data are the same in the document information storage unit 22 and the document vector storage unit 141.

[0070] The creator is the identification information of the creator of the document data. The update history is information including the date of update and the identification information of the person who updated the document data for each update. In this embodiment, the identification information of the creator or updater of the document data is assumed to be an employee ID of Company X. The file path is the path name of the file storing the document data. The summary is a summary of the contents of the document data (e.g., a summary). The access control information is information for restricting access to the document information to a predetermined range of users. In other words, the access control information is information indicating whether each user has access authority. For example, the access control information may include information indicating users or groups with read authority and information indicating users or groups with write authority. A group refers to a set of one or more users. The label list is a list of labels (hereinafter referred to as "document labels") assigned to the document data by the label assignment unit 126. Words with relatively large TF-IDF values ​​from among the words included in the document data may be used as document labels.

[0071] In step S201, document information for which the logged-in user has access rights is obtained from the top N pieces of document information (note that, as will be described later, information may also be obtained to display that the logged-in user does not have access rights to document information for which the logged-in user does not have access rights).

[0072] Next, the data search unit 124 sorts (arranges) the acquired document information in descending order of similarity (S202).

[0073] Fig. 9 is a diagram showing an example of the sorting result of document information, in which document names and similarities are sorted in descending order of similarity.

[0074] Next, the display information generating unit 130 generates display information for displaying the sorted results as the document information collection results (search results) (S203).

[0075] The display information generating unit 130 generates display information based on the creator, update history, file path, summary, label list, etc. of the document information that the login user has permission to view from among the top N document data.

[0076] Next, the display information sending unit 131 and the display control unit 31 of the user terminal 30 execute a process of outputting the display information (S204). Specifically, the display information sending unit 131 sends the display information to the user terminal 30. The display control unit 31 of the user terminal 30 displays a search result screen as a result of document collection based on the display information.

[0077] 10 is a diagram showing a display example of the search result screen 520. As shown in FIG.

[0078] The information collection condition display area 521 is an area that displays the target collection conditions, and includes an information type display area 5211 and a query display area 5212. The information type display area 5211 is an area that displays the target information type. The query display area 5212 is an area that displays the target query. The information type display area 5211 and the query display area 5212 may be operable. In this case, the information type and part of the query may be changed via the information type display area 5211 and the query display area 5212, and the execute button 5213 may be pressed, thereby re-executing step S101 and subsequent steps in FIG. 4 .

[0079] The search result display area 522 is an area where the creator, updater, file path, summary, label list, etc. are displayed for each of the top N document information items. Note that the updater may be, for example, the updater involved in the last update in the update history.

[0080] The logged-in user can refer to the search result screen 520 to check a list of document information collected according to the target collection conditions.

[0081] The search result screen 520 also includes a query relationship diagram button 523. When the query relationship diagram button 523 is pressed, the user terminal 30 transmits a request corresponding to the query relationship diagram button 523 to the information processing device 10. When the classification unit 125 of the information processing device 10 receives the request ("Query relationship diagram" in S205 of FIG. 7), it classifies the document vectors with the top N similarities into multiple classes by clustering (S206). The clustering may be performed using, for example, the k-means method or any other known method.

[0082] Next, the labeling unit 126 assigns a label to each class (S207). For example, the labeling unit 126 may assign one or more words with relatively high TF-IDF values ​​in a set of document data belonging to a certain class as the label of the class. Alternatively, the labeling unit 126 may assign one or more document labels with relatively high appearance frequencies in a list of document labels of each document data belonging to the class as the label of the class.

[0083] Next, the relationship diagram generation unit 127 generates a relationship diagram (hereinafter referred to as a "query relationship diagram") showing the relationship between the target query and the top N pieces of document information (S208). Next, the display information generation unit 130 and the user terminal 30 execute a query relationship diagram output process. Specifically, the display information generation unit 130 transmits display information of the query relationship diagram to the user terminal 30. The display control unit 31 of the user terminal 30 displays the query relationship diagram based on the display information.

[0084] FIG. 11 is a diagram showing an example of a query relationship diagram. As shown in FIG. 11, the query relationship diagram is a graphic diagram in the form of a graph in which the target query, each class, and N pieces of document information are nodes. The target query and each class are connected by a branch, and each piece of document information is connected by a branch to the class to which it belongs. In other words, in FIG. 11, the rounded rectangle directly connected to the target query is the node corresponding to the class. The character string in the node is the label of the class corresponding to the node.

[0085] The oval nodes connected to the rounded rectangles are nodes corresponding to the document information classified into the class corresponding to the rectangle. The character string within the node is the document name of the document information corresponding to the node. By referring to the query relationship diagram, users can get an overview of the relationship between the collected document information group and the target query.

[0086] Alternatively, when a details button 524 corresponding to any document information is pressed on the search result screen 520 (FIG. 10), the display control unit 31 of the user terminal 30 transmits a request corresponding to the details button 524 to the information processing device 10. The request includes, for example, the document ID of the document information.

[0087] When the display information generation unit 130 of the information processing device 10 receives the request ("Details" in S205), it executes detailed information output processing for document information (hereinafter referred to as "target document information") related to the document ID (hereinafter referred to as "target document ID") included in the request (S210). Specifically, the display information generation unit 130 refers to the document related storage unit 142 and identifies the document ID of document information (hereinafter referred to as "related document information") related to document data having the top M similarities to document data related to the target document information.

[0088] Fig. 12 is a diagram showing an example of the configuration of the document relation storage unit 142. As shown in Fig. 12, the document relation storage unit 142 is a storage unit in a matrix format in which the document IDs of all document data are arranged in the row and column directions. All document data refers to all document data whose document information is stored in the document information storage unit 22. The value of an element in a certain row and a certain column is the similarity between the document vector of the document data associated with the document ID in that row and the document vector associated with the document ID in that column.

[0089] The similarity between the document data stored in the document relation storage unit 142 is calculated in advance by, for example, the comparison unit 123. The comparison unit 123 can acquire the document vector of each document data by referring to the document vector storage unit 141 (FIG. 6), and can calculate the similarity between the document data using the document vector.

[0090] For example, the display information generating unit 130 can sort the similarities in the row of the target document ID in descending order and identify the document IDs of the top M related document information.

[0091] The display information generation unit 130 acquires the target document information and each related document information from the document information storage unit 22 (FIG. 8) of the information management device 20 based on the respective document IDs, and generates display information for the document details screen based on the acquired document information. When the display information transmission unit 131 transmits the display information to the user terminal 30, the display control unit 31 of the user terminal 30 displays the document details screen based on the display information.

[0092] 13 is a diagram showing a display example of a document details screen 530. As shown in FIG. 13, a document details screen 530 includes a target document display area 531, a related document display area 532, and the like.

[0093] The target document display area 531 is an area where the target document information is displayed. In the target document display area 531, in addition to the items displayed on the search result screen 520 (FIG. 10), the update history and summary of the target document information are displayed.

[0094] The related document display area 532 is an area where related document information is displayed. The related document display area 532 includes a details button 5321 for each piece of related document information. When the details button 5321 corresponding to any piece of related document information is pressed, the related document information is treated as target document information and detailed information output processing similar to that of step S210 is executed. As a result, the user can recursively collect document information related to the target query (in a chain reaction).

[0095] The document details screen 530 also includes a document relationship diagram button 533. When the document relationship diagram button 533 is pressed, the user terminal 30 transmits a request corresponding to the document relationship diagram button 533 to the information processing device 10. When the information processing device 10 receives the request ("document relationship diagram" in S211 of FIG. 7), it executes document relationship diagram output processing (S212). In the document relationship diagram output processing, the target query is replaced with the target document data, and the top N document information is replaced with M related document information, and processing similar to steps S206 to S209 is executed.

[0096] Specifically, the classification unit 125 classifies document vectors related to the M pieces of related document information into multiple classes by clustering. Then, the label assignment unit 126 assigns a label to each class. Then, the relationship diagram generation unit 127 generates a relationship diagram (hereinafter referred to as a "document relationship diagram") showing the relationship between the target document information and the M pieces of related document information. The display information generation unit 130 generates display information for the document relationship diagram and transmits the display information to the user terminal 30. The display control unit 31 of the user terminal 30 displays the document relationship diagram based on the display information.

[0097] Fig. 14 is a diagram showing an example of a document relationship diagram. As shown in Fig. 14, the document relationship diagram is a graphic diagram in the form of a graph in which target document information, each class, and M pieces of related document information are nodes. The target document information and each class are connected by branches, and each piece of related document information is connected by a branch to the class to which it belongs. In other words, in Fig. 14, the rounded rectangle directly connected to the target document information is a node corresponding to a class. The character string in the node is the label of the class corresponding to the node.

[0098] By visualizing the relationship structure between documents as shown in FIG. 11 or FIG. 14, the user can intuitively grasp the relationship between the information they want to collect and the document information.

[0099] The oval nodes connected to the rounded rectangles are nodes corresponding to related document information classified into the class corresponding to the rectangles. The character strings within the nodes are the document names of the related document information corresponding to the nodes. By referring to the document relationship diagram, users can get an overview of the relationship between the target document information and the related document information.

[0100] Alternatively, when a link to any document name is selected on the search result screen 520 (FIG. 10) ("Document Link" in S205 of FIG. 7), when a link to any document name is selected on the document details screen 530 (FIG. 13) ("Document Link" in S211), or when a link to any document name is selected on the query relationship diagram (FIG. 11) or the document relationship diagram (FIG. 14) ("Document Link" in S213), the user terminal 30, the information processing device 10, and the information management device 20 execute document data output processing (S214). Specifically, the display control unit 31 of the user terminal 30 transmits to the information processing device 10 the document ID of the document information for which the link to the document name was clicked (hereinafter referred to as the "target document ID"). The display information transmission unit 131 of the information processing device 10 transmits to the user terminal 30 a redirect command including a URL (Uniform Resource Locator) or the like for requesting the information management device 20 to refer to the document data associated with the target document ID. When the display control unit 31 of the user terminal 30 accesses the URL in accordance with the redirect command, the document management unit 21 of the information management device 20 transmits the document data associated with the target document ID to the user terminal 30. Upon receiving the document data, the display control unit 31 of the user terminal 30 displays the document data, allowing the user to check the contents of the document data.

[0101] The logged-in user can associate the document information collection results with the target query and save them as a workspace. Saving the collection results as a workspace can be likened to saving the collection results as a bookmark. In this case, the logged-in user selects, from the selection components 525 arranged for each document information on the search result screen 520 ( FIG. 10 ), a selection component 525 corresponding to the document information they wish to associate with the target query and save in the workspace. For example, the logged-in user selects one or more document information items corresponding to the information they desire to save in the workspace. All document information items included in the collection results may be selected. When the workspace creation button 526 is pressed with one or more selection components 525 selected, the display control unit 31 of the user terminal 30 displays a screen for accepting input of a name to be assigned to the workspace (hereinafter referred to as the “workspace name”). When the logged-in user inputs the workspace name via the screen, the display control unit 31 transmits a workspace creation request to the information processing device 10, including the document IDs of the document information items corresponding to the selected selection components 525, the workspace name, and the target query. Upon receiving the generation request, the information processing device 10 executes, for example, the processing procedure shown in FIG.

[0102] FIG. 15 is a flowchart illustrating an example of a processing procedure for generating a workspace.

[0103] In step S251, the labeling unit 126 extracts a relatively important portion (a predetermined number) of words from a set of words included in the document data related to one or more selected pieces of document information as labels for the workspace. The label extraction method may be the same as that described above.

[0104] In step S252, the classification unit 125 classifies the document vectors of each selected document information (hereinafter referred to as "selected document information") into multiple classes (hereinafter referred to as "belonging classes") by clustering. The classification method into classes may be the same as that described above (for example, S206 in FIG. 7).

[0105] Next, the labeling unit 126 assigns a label to each class (S253). The method of assigning a label to each class may be the same as that described above (for example, S207 in FIG. 7).

[0106] Next, the workspace generating unit 132 stores the workspace that associates the selected document information with the target query in the workspace storage unit 25 of the information management device 20 (S254).

[0107] Fig. 16 is a diagram showing an example of the configuration of the workspace storage unit 25. As shown in Fig. 16, the workspace storage unit 25 stores, for each workspace, a workspace including a workspace ID, workspace name, label, creator, updater, query, number of uses, evaluation score, belonging data ID, belonging data path, belonging class label, etc.

[0108] The workspace ID is identification information of the workspace, and is assigned to the workspace by the workspace generation unit 132 in step S254, for example. The workspace name is the name of the workspace entered by the user, as described above. The creator is identification information (user ID, name, etc.) of the creator of the workspace. Here, the identification information of the logged-in user is stored as the creator. The updater is identification information (user ID, name, etc.) of the person who updated the workspace when it was updated. In other words, the workspace can be updated. The query is the query (target query) entered when collecting the document information that formed the workspace. Therefore, the query can also be said to be information indicating the perspective on which the workspace is based on the collection of document information. The number of uses is the number of times the workspace has been used (referenced). The evaluation score is the evaluation value entered by a user who referenced the workspace. For example, the evaluation score is the average value of a five-point evaluation. The belonging data ID is the document ID of each piece of document information belonging to the workspace. The belonging data path is the file path of the document data related to each piece of document information. The class label is a label for each class generated in step 153. The same class label is saved for document information classified into the same class in the workspace.

[0109] Note that multiple pieces of document information belonging to the same workspace may be classified into units (hereinafter referred to as "folders") that can be arbitrarily configured by the user, separate from classes. In this case, folders may be able to form a hierarchical structure similar to folders in a general OS. That is, document information in one workspace may be classified into a hierarchical structure. Folders for a workspace may be created in the processing procedure of FIG. 15 or at a user-specified timing after the processing procedure of FIG. 15. In either case, the workspace generation unit 132 receives from the user a specification of the parent of a folder to be created in the hierarchical structure, the folder name of the folder, and document information to be classified into the folder. The parent of the folder to be created is the workspace if the folder is a folder directly under the workspace. If the folder to be created is a child of another folder, the other folder is the parent. The folder name may be set by the user. This allows the workspace to be classified according to the user's perspective or convenience, separate from classes. Note that document information belonging to the same class may be allowed or prohibited to be classified into different folders.

[0110] When classification by folders is introduced, each record in the workspace storage unit 25 may further include the items "folder name" and "parent" for each data ID to which it belongs (i.e., for each document information belonging to a workspace). "Folder name" is the folder name of the folder to which the document information related to the data ID belongs. "Parent" is the workspace name if the parent of the folder is a workspace, and is the folder name of the other folder if the parent of the folder is another folder.

[0111] As described above, a workspace is information associated with document information. Therefore, in the document information collection results, for document information that has already been saved (associated) with a workspace, workspaces associated with the document information may also be collected. In this case, the search result screen 520 displayed in step S204 of FIG. 7 may have a configuration such as that shown in FIG. 17.

[0112] Fig. 17 is a diagram showing a second display example of the search result screen. In Fig. 17, the same parts as in Fig. 10 are given the same reference numerals, and their explanation will be omitted.

[0113] The search result screen 520 shown in Fig. 17 further includes a related workspace display area 527. The related workspace display area 527 is an area where the workspace names of workspaces related to each piece of document information are displayed. A workspace related to a certain piece of document information is a workspace that includes the document ID of the document information as its belonging data ID. Each workspace name may be provided with information (e.g., a link) to guide the user to the workspace.

[0114] In this way, a link to a workspace is displayed for each piece of collected document information, allowing the user to obtain other document information related to the collected document information based on the workspace.

[0115] Next, a detailed description will be given of step S107 in Fig. 4. Fig. 18 is a flowchart for explaining an example of a processing procedure for expert collection result output processing.

[0116] In step S301, the expert collection unit 128 collects (extracts) identification information of experts associated with document information related to the top N document vectors extracted in step S104 of FIG. 4 (hereinafter referred to as "target document information"), for example, by referring to the document information storage unit 22 (FIG. 8). Identification information of an expert associated with certain document information is identification information of a person who is presumed to be an expert on the document information. In this embodiment, the creator or updater of the document information is presumed to be an expert on the document information. This is because it is considered highly likely that the creator or updater is familiar with the content contained in the document data related to the document information or knows the content of the document data. Therefore, the expert collection unit 128 refers to the document information storage unit 22 (FIG. 8) and collects the employee ID of the creator or updater as identification information of the expert for each target document information.

[0117] Next, the expert collection unit 128 calculates the relevance with the target query for each collected employee ID (i.e., for each expert) (S302). The relevance with the target query refers to an index indicating the strength of the relevance with the target query. The relevance with the target query for a certain expert may be a value based on the similarity calculated in step S103 of FIG. 4 for document information created or updated by the expert among the top N document information. In this case, for example, the average or sum of the similarities may be used as the relevance. Furthermore, taking into account the accessibility (ease of access) from the user who input the target query (i.e., the logged-in user), an index indicating the proximity to the logged-in user may be added to the relevance of each expert. The index may be evaluated based on the distance between the department to which the logged-in user belongs and the department to which the expert belongs. The distance between departments can be evaluated based on organizational information. For example, the distance between department A and department B may be a value based on the sum of the number of levels from a common upper organization (e.g., a business division or headquarters) between department A and department B to department A and department B. In this case, the smaller the sum (i.e., the closer the distance), the larger the value added to the relevance degree. Note that organizational information can be acquired from the organizational information storage unit 24. Alternatively, the index may be evaluated based on the amount of communication within the company between the logged-in user and the expert. For example, the index may be calculated based on email exchanges, chat exchanges, the number of times the logged-in user has attended the same meeting, etc. In this case, the more communication an expert has with the logged-in user, the larger the relevance degree value may be. The relevance degree may also be calculated using other methods.

[0118] Next, the expert collection unit 128 sorts the identification information of the experts in descending order of the degree of association (S303).

[0119] Fig. 19 is a diagram showing an example of the sorting results of experts. Fig. 19 shows an example in which the identification information of experts is sorted in descending order of relevance. For convenience, in Fig. 19, the names of the experts are used as identification information.

[0120] In addition, if the number of experts collected exceeds a threshold (here, M), the expert collection unit 128 may extract the top M experts in terms of relevance, and only the extracted experts may be processed in subsequent steps.

[0121] Next, the expert collection unit 128 acquires the employee information of the expert (hereinafter referred to as "expert information") from the employee information storage unit 23 of the information management device 20 based on the identification information of the expert (S304). The employee information storage unit 23 stores, for each employee, employee information that can be shared within the company, such as name, department, and contact information (telephone number, email address, etc.).

[0122] Next, the display information generating unit 130 generates display information for displaying the sorting result as the collection result of the expert information (S305).

[0123] Next, the display information sending unit 131 and the display control unit 31 of the user terminal 30 execute a process of outputting the display information (S306). Specifically, the display information sending unit 131 sends the display information to the user terminal 30. The display control unit 31 of the user terminal 30 displays an expert collection result screen based on the display information. The expert collection screen has, for example, the same configuration as the search result screen (FIG. 10), and the search result display area 522 is a screen that includes a list of the expert information sorted in step S303.

[0124] When one piece of expert information is selected on the expert collection screen and a command to display detailed information is input, the display control unit 31 of the user terminal 30 transmits a detailed information display request, including the employee ID associated with the selected expert information, to the information processing device 10. Upon receiving the detailed information display request ("detailed information" in S307), the information processing device 10 executes detailed information output processing for the expert associated with the employee ID included in the detailed information display request (S308). For example, the display information generation unit 130 acquires various information associated with the employee ID from the employee information storage unit 23 or another database in the information management device 20 (hereinafter referred to as "detailed information") and generates display information for a screen (hereinafter referred to as an "expert details screen") that displays the acquired detailed information. When the display information transmission unit 131 transmits the display information to the user terminal 30, the display control unit 31 of the user terminal 30 displays the expert details screen based on the display information. The logged-in user can obtain more detailed information about the expert by referring to the expert details screen.

[0125] Alternatively, when one piece of expert information is selected on the expert collection screen and an instruction to display the retained information is input, the display control unit 31 of the user terminal 30 transmits a retained information display request including the employee ID associated with the selected expert information to the information processing device 10. Upon receiving the retained information display request ("Retained Information" in S307), the information processing device 10 executes a retained information output process for the expert associated with the employee ID included in the retained information display request (S309). For example, the display information generation unit 130 acquires a list of document information including the employee ID as a creator or updater from the document information storage unit 22 (FIG. 8). That is, the acquired document information is considered to be document information related to document data including information held by the expert. The display information generation unit 130 generates display information for a screen (hereinafter referred to as a "retained information screen") that displays the list of acquired document information. When the display information transmission unit 131 transmits the display information to the user terminal 30, the display control unit 31 of the user terminal 30 displays the retained information screen based on the display information. The logged-in user can check the information held by the expert by referring to the held information screen. Note that workspaces created or updated by the expert (Figure 16) may also be displayed on the held information screen.

[0126] Alternatively, when one piece of expert information is selected on the expert collection screen and an instruction to display an access route is input, the display control unit 31 of the user terminal 30 transmits an access route display request including the employee ID associated with the selected expert information to the information processing device 10. Upon receiving the access route display request ("Access Route" in S307), the information processing device 10 executes an access route output process for the expert associated with the employee ID included in the access route display request (S310). The access route is information indicating a route for a logged-in user to access the expert. For example, in a graph (e.g., a tree structure) representing the organizational structure of a company, the access route may be a route from the department to which the logged-in user belongs to to the department to which the expert belongs. In this case, the display information generation unit 130 can identify the access route based on the organizational information. Alternatively, the access route may be a path of human relationships from the logged-in user to the expert. In this case, the display information generation unit 130 may recursively search for acquaintances of the logged-in user from, for example, email history, chat history, meeting history, etc. stored in a database within the company, and acquire, as an access route, a list of acquaintances searched until the expert appears as an acquaintance. The display information generation unit 130 generates display information for a screen (hereinafter referred to as an "access route screen") that displays the acquired access route. When the display information transmission unit 131 transmits the display information to the user terminal 30, the display control unit 31 of the user terminal 30 displays the access route screen based on the display information.

[0127] For example, by displaying the "expert name" in place of the document name of "Space Business Development.pdf," which is located in the center of the relationship diagram shown in the example of document relationship diagram in Figure 14 above (for example, "Employee TS," which has the highest degree of relationship), users can intuitively grasp the access route to the expert from the relationship diagram between the information held by each person and the correlation diagram between the information held by each person, making it easier to search for knowledge and experts. By referring to the access route screen, logged-in users can obtain clues for accessing the expert.

[0128] Next, a detailed description will be given of step S108 in Fig. 4. Fig. 20 is a flowchart for explaining an example of a processing procedure for workspace collection result output processing.

[0129] In step S401, the workspace collection unit 129 collects workspaces related to the document information (hereinafter referred to as "target document information") related to the top N document vectors extracted in step S104 of Fig. 4 from the workspace storage unit 25 (Fig. 16). The identification information of a workspace related to certain document information is a workspace that includes the document ID of the document information as the belonging data ID.

[0130] Next, the workspace collection unit 129 calculates the relevance with the target query for each collected workspace (S402). The relevance with the target query is an index that indicates the strength of the relevance with the target query. The relevance of a certain workspace with the target query may be a value based on the similarity calculated in step S103 of FIG. 4 for document information related to the workspace among the top N pieces of document information. In this case, for example, the average or total of the similarities may be used as the relevance.

[0131] Next, the workspace collection unit 129 sorts the workspaces in descending order of relevance (S403).

[0132] Fig. 21 is a diagram showing an example of the sorting result of workspaces, in which workspaces are sorted in descending order of relevance.

[0133] In addition, if workspaces have been collected in excess of a threshold (here, M workspaces), the workspace collection unit 129 may extract the top M workspaces in terms of relevance, and only use the extracted workspaces as the processing targets for subsequent steps.

[0134] Next, the display information generating unit 130 generates display information for displaying the sorted results as the workspace collection results (S404).

[0135] Next, the display information sending unit 131 and the display control unit 31 of the user terminal 30 execute a process of outputting the display information (S405). Specifically, the display information sending unit 131 sends the display information to the user terminal 30. The display control unit 31 of the user terminal 30 displays a workspace collection result screen based on the display information.

[0136] 22 is a diagram showing a display example of a workspace collection result screen 540. As shown in FIG.

[0137] The information collection condition display area 541 is an area for displaying target collection conditions, and includes an information type display area 5411 and a query display area 5412. The functions of the information type display area 5411 and the query display area 5412 are the same as those of the information type display area 5211 and the query display area 5212 on the search result screen 520 (FIG. 10).

[0138] The search result display area 542 is an area where a list of sorted workspaces is displayed.

[0139] The logged-in user can refer to workspace collection result screen 540 to check a list of workspaces collected according to the target collection conditions.

[0140] When a details button 524 corresponding to any workspace is pressed on the workspace collection result screen 540, the display control unit 31 of the user terminal 30 transmits a request corresponding to the details button 524 to the information processing device 10. The request includes, for example, the workspace ID of the workspace.

[0141] When the display information generation unit 130 of the information processing device 10 receives the request, it executes detailed information output processing for the workspace (hereinafter referred to as the "target workspace") related to the workspace ID (hereinafter referred to as the "target workspace ID") included in the request (S406). Specifically, the display information generation unit 130 references the workspace storage unit 25 and generates display information for a screen (hereinafter referred to as the "workspace details screen") that displays more detailed information about the target workspace than the content displayed on the workspace collection result screen 540. When the display information transmission unit 131 transmits the display information to the user terminal 30, the display control unit 31 of the user terminal 30 displays the workspace details screen based on the display information.

[0142] 23 is a diagram showing a display example of a workspace details screen 550. As shown in FIG. 23, a workspace details screen 550 includes a basic information display area 551, a configuration display area 552, an associated document display area 553, and the like.

[0143] Basic information display area 551 is an area that includes the content displayed on workspace collection result screen 540 for the target workspace, an edit button 5511, and an evaluation button 5512.

[0144] The configuration display area 552 is an area that contains information indicating the relationship between the document information group belonging to the target workspace (FIG. 16), which can be identified based on the class label and data ID of the target workspace, and the class into which the document information group is classified. FIG. 23 shows an example in which three classes belong to the target workspace.

[0145] The belonging document display area 553 is an area that includes a list of document information that belongs to the class (hereinafter referred to as the "target class") selected in the configuration display area 552. In Fig. 23, "No access right" is displayed for the third document information. "No access right" indicates that the logged-in user does not have access rights to the document information.

[0146] The logged-in user can edit a workspace via the workspace details screen 550. For example, any document information belonging to the target workspace can be deleted from the target workspace, or certain document information can be added to the target workspace. After performing such an editing operation, when the logged-in user presses the edit button 5511, the user terminal 30 transmits the edited content to the information processing device 10. When the workspace editing unit 133 of the information processing device 10 receives the edited content ("Edit" in S407), it reflects the edited content in the record corresponding to the target workspace in the workspace storage unit 25 (FIG. 16) (S408).

[0147] Alternatively, when the rating button 5512 is pressed on the workspace details screen 550, the display control unit 31 of the user terminal 30 displays a screen for receiving input of a rating point. When a rating point of 0 to 5 is input on this screen, the display control unit 31 of the user terminal 30 transmits the input rating point to the information processing device 10. When the workspace generation unit 132 of the information processing device 10 receives the rating point ("Rating" in S407), it updates the number of uses and the rating point of the record corresponding to the target workspace in the workspace storage unit 25 (FIG. 16) (S409). Specifically, the workspace generation unit 132 adds 1 to the number of uses. Regarding the rating point, if the number of uses before the update is x1, the number of uses after the update is x2, and the rating point before the update is y1, the workspace generation unit 132 calculates the updated rating point y2 as follows: y2=y1×x1÷x2 Alternatively, when a link to any document name is selected in the associated document display area 553 on the workspace details screen 550, the user terminal 30, the information processing device 10, and the information management device 20 execute document data output processing (S410). Note that the document data output processing is as described in step S214 of Fig. 7. Therefore, as a result of the document data output processing, the user can confirm the contents of the document data related to the document name.

[0148] As described above, according to the first embodiment, the name of a workspace, which is a data set, can be displayed in association with the label assigned to that workspace (FIG. 23). Here, the label of a certain workspace is assigned based on the content of the data belonging to that workspace. In other words, the label of a certain workspace can be said to be information that succinctly indicates the content of the data belonging to that workspace. Therefore, it is possible to assist in determining whether a certain data set (workspace) contains desired information.

[0149] Next, a second embodiment will be described. In the second embodiment, differences from the first embodiment will be described. Therefore, points not specifically mentioned may be the same as in the first embodiment. In the second and subsequent embodiments, it is assumed that one or more workspaces have already been generated.

[0150] FIG. 24 is a flowchart illustrating an example of a processing procedure executed by the information processing system according to the second embodiment.

[0151] In step S501, the reception unit 121 of the information processing device 10 receives a workspace list request from the display control unit 31 of the user terminal 30. The workspace list request is a request to display a list of all workspaces. The workspace list request may be input, for example, by selecting "workspace" as the information type in the information type selection area 511 of the collection condition input screen 510 (FIG. 5) and leaving the query input area 512 blank (i.e., without inputting a query) and pressing the execute button 513. Alternatively, the workspace list request may be input using another screen.

[0152] Next, a loop process including steps S502 and S503 is executed for each workspace stored in workspace storage unit 25. The workspace that is the processing target in this loop process is hereinafter referred to as the "target workspace."

[0153] In step S502, the labeling unit 126 calculates the TF-IDF value of each word in a set of words included in the document data relating to all document information belonging to the target workspace.

[0154] Next, the label assignment unit 126 extracts some of the words with the highest TF-IDF values ​​(for example, up to the Kth word) as labels for the target workspace, and assigns the labels to the target workspace (S503). Assigning a label to the target workspace means recording the label in the "label" of the record corresponding to the target workspace in the workspace storage unit 25 (FIG. 16).

[0155] When the loop process is completed for all workspaces, the display information generating unit 130 generates display information that displays the workspace name of each workspace in association with the label assigned to that workspace (S504).

[0156] Next, the display information sending unit 131 and the display control unit 31 of the user terminal 30 execute a process of outputting the display information (S505). Specifically, the display information sending unit 131 sends the display information to the user terminal 30. The display control unit 31 of the user terminal 30 displays a workspace list screen based on the display information.

[0157] FIG. 25 is a diagram showing an example of the workspace list screen. As shown in FIG. 25, the workspace list screen 610 includes, for each workspace, a workspace name and a label associated with each other. Each label is displayed surrounded by a line (this is also the case in the first embodiment). That is, the display information generating unit 130 generates display information so that the label is displayed surrounded by a line. The "line" includes curved and straight lines, as well as boundary lines separating areas of different colors. This makes it easier to visually distinguish the difference between the label and the workspace name or document name, and makes the label more noticeable than the workspace name or document name. As a result, when a user glances at the search results, the search results are more easily understood, which is expected to improve operability, such as adjusting search results.

[0158] 25 shows an example in which multiple labels are displayed for each workspace, but it is also possible to display only one label (for example, the word with the highest TF-IDF value). However, it is preferable to display multiple labels, and in that case, three labels would be preferable.

[0159] Fig. 26 is a diagram for explaining the desirable number of labels. In Fig. 26, the rounded rectangles represent labels, and the ovals represent keywords related to the images that users associate with the labels.

[0160] For example, as shown in Figure 26(1), if only one label, "brain," is given, people will recall keywords such as "emotion," "memory / creation," and "consciousness" as images of the content of documents related to that workspace.

[0161] In contrast, if there are two labels, "brain" and "exercise," the resulting images are related to the body, such as "brain activity," "sleep," "concentration," and "stress relief," as shown in Figure 26(2). This shows a dramatic change in direction compared to when there is only one label. Furthermore, the image based on the label "brain" alone is clearly different from the images of "robot," "deep learning," and "artificial life" that are based on two labels, "brain" and "machine." In other words, the amount of information increases dramatically between when there is only one label and when there are two or more labels, and the images that emerge visually are also significantly different.

[0162] As such, there is a demand for a UI that can accurately convey an image visually. If the number of labels is too large, it will become noise and be influenced by each user's experience, and there is a possibility that the correct image will not be conveyed accurately. The image shown by the labels needs to accurately convey information as the image that the majority of people have. It is desirable that the number of labels be a maximum of four or less. If there are more than that, the image will differ from person to person. The desirable number of labels is three, which is thought to be able to convey an accurate image concisely. Note that this point is also true in the first embodiment.

[0163] As described above, according to the second embodiment, similarly to the first embodiment, it is possible to display the name of a workspace, which is a data set, in association with the label assigned to the workspace, thereby helping to determine whether a certain data set (workspace) contains desired information.

[0164] Next, a third embodiment will be described. In the third embodiment, differences from the above-described embodiments will be described. Therefore, unless otherwise specified, the third embodiment may be the same as the above-described embodiments.

[0165] Fig. 27 is a flowchart for explaining an example of a processing procedure executed by an information processing system according to the third embodiment. In Fig. 27, the same steps as those in Fig. 24 are assigned the same step numbers, and their explanations will be omitted. In Fig. 27, step S502 is replaced with step S502a.

[0166] In step S502a, the labeling unit 126 calculates the TF-IDF value of each word in a set of words contained in document data relating to all document information belonging to the target workspace and words contained in the names of units for classifying document data belonging to the target workspace (hereinafter referred to as "classification unit names"). That is, the TF-IDF value of each word is calculated for a set obtained by adding words contained in each classification unit name to the set of words contained in all document data belonging to the target workspace. Here, the classification unit name is, for example, a folder name. The workspace name of the target workspace may be included in the classification unit name.

[0167] As described above, according to the third embodiment, the label displayed in association with the workspace name can be selected from among words constituting the folder name or workspace name arbitrarily assigned by the user, thereby increasing the likelihood that the label displayed reflects the user's intention.

[0168] Next, a fourth embodiment will be described. In the fourth embodiment, differences from the above-described embodiments will be described. Therefore, unless otherwise specified, the fourth embodiment may be the same as the above-described embodiments.

[0169] Fig. 28 is a flowchart illustrating an example of a processing procedure executed by an information processing system according to the fourth embodiment. In Fig. 28, the same steps as those in Fig. 24 are assigned the same step numbers, and descriptions thereof will be omitted.

[0170] 28, when the configuration of any of the existing workspaces stored in workspace storage unit 25 (FIG. 16) is edited by workspace editing unit 133 (Yes in S601), steps S502 and S503 are executed for that workspace (hereinafter referred to as the "target workspace"). As a result, the label stored for the target workspace in workspace storage unit 25 is updated (changed).

[0171] Here, editing the configuration of a workspace refers to a change to the workspace that may change the labels of the workspace. Therefore, adding new document information (i.e., document data) to a workspace or deleting any document information (i.e., document data) belonging to a workspace corresponds to editing the workspace. In this case, the word group of the document data that forms the parent set of the labels of the workspace changes.

[0172] Furthermore, when the fourth embodiment is combined with the third embodiment, changing the folder structure within a workspace (adding a folder, deleting a folder, changing a folder name) or changing the workspace name also corresponds to editing the workspace.

[0173] In the fourth embodiment, the processing procedure executed in response to a workspace list request differs from that in the above-described embodiments.

[0174] Fig. 29 is a flowchart for explaining an example of a processing procedure executed in response to a workspace list request in the fourth embodiment. In Fig. 29, the same steps as in Fig. 24 are assigned the same step numbers, and their explanations will be omitted. In Fig. 29, step S602 is executed instead of the loop processing including steps S502 and S503.

[0175] In step S602, labeling unit 126 acquires the workspace names and labels of all workspaces stored in workspace storage unit 25 (FIG. 16).

[0176] In the following steps S504 and S505, display information is generated based on the acquired workspace name and label, and workspace list screen 610 (FIG. 25) is displayed based on the display information.

[0177] As described above, according to the fourth embodiment, it is not necessary to execute processing to assign a label each time the workspace list screen 610 (FIG. 25) is displayed, thereby making it possible to improve the efficiency of the display processing of the workspace list screen 610.

[0178] Next, a fifth embodiment will be described. In the fifth embodiment, differences from the above-described embodiments will be described. Therefore, unless otherwise specified, the fifth embodiment may be the same as the above-described embodiments.

[0179] Fig. 30 is a flowchart illustrating an example of a processing procedure executed by an information processing system according to the fifth embodiment. In Fig. 30, the same steps as those in Fig. 24 are assigned the same step numbers, and descriptions thereof will be omitted.

[0180] In Fig. 30, when the content of document data relating to any document information stored in the document information storage unit 22 (Fig. 8) has been edited (i.e., when a word constituting the document data has been changed) (Yes in S603), steps S604 and thereafter are executed. Editing of document data may be detected by polling the storage unit (file system, etc.) in which the document data is stored. Hereinafter, edited document data will be referred to as "target document data."

[0181] In step S604, the workspace collection unit 129 identifies the workspace to which the target document data belongs by referring to the workspace storage unit 25. Specifically, the workspace collection unit 129 identifies a workspace including an associated data path that matches the file path of the target document data as the workspace to which the target document data belongs.

[0182] Next, a loop process including steps S502 and S503 is executed for each identified workspace, and as a result, the label of that workspace may be updated.

[0183] In the fifth embodiment, the processing procedure in response to a workspace list request may be the same as that in the fourth embodiment.

[0184] As described above, according to the fifth embodiment, it is not necessary to execute processing to assign a label each time the workspace list screen 610 (FIG. 25) is displayed, thereby making it possible to improve the efficiency of the display processing of the workspace list screen 610.

[0185] Next, a sixth embodiment will be described. In the sixth embodiment, differences from the above-described embodiments will be described. Therefore, unless otherwise specified, the sixth embodiment may be the same as the above-described embodiments.

[0186] FIG. 31 is a flowchart illustrating an example of a processing procedure executed by the information processing system according to the sixth embodiment.

[0187] In step S605, the information processing device 10 executes a workspace search process based on the query. This search process is a process procedure that follows steps S101 to S105 in Fig. 4 and continues to step S401 in Fig. 20, among the process procedures that are executed when "workspace" is selected as the information type on the collection condition input screen 510 (Fig. 5), and the execute button 513 with a query input is pressed. Therefore, in step S605, workspaces to which the document data with the top N similarities to the query belong are searched for.

[0188] Next, a loop process including steps S502 and S503a is executed for each workspace found in the loop process. The workspace that is the processing target in the loop process is hereinafter referred to as the "target workspace."

[0189] In step S502, the labeling unit 126 calculates the TF-IDF value of each word in a set of words included in document data related to all document information belonging to the target workspace, as described in Fig. 24. Note that the words included in the document data (i.e., the words whose TF-IDF values ​​are calculated in step S502) will be referred to as "document words" hereinafter.

[0190] Next, the labeling unit 126 extracts some of the document words with the highest TF-IDF values ​​(for example, up to the Kth) from among the document words that include any of the words that make up the query, as labels for the target workspace, and assigns the labels to the target workspace (S503). That is, the labeling unit 126 determines whether each document word includes any of the words that make up the query (hereinafter referred to as "query word") in descending order of TF-IDF value, and when K document words that are determined to include the query word are found, extracts the K document words as labels.

[0191] When the loop processing is completed for all workspaces, steps S504 and S505 are executed as described with reference to FIG.

[0192] As described above, according to the sixth embodiment, by narrowing down the labels displayed for each workspace to those that match the query, which is the search statement, it is possible to make it easier for the user to find a workspace that matches their search intent.

[0193] The functions of the above-described embodiments can be realized by one or more processing circuits. Here, the term "processing circuit" in this specification includes a processor programmed to execute each function by software, such as a processor implemented by an electronic circuit, as well as devices such as an ASIC (Application Specific Integrated Circuit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), and conventional circuit modules designed to execute each of the above-described functions.

[0194] Although the embodiments of the present invention have been described in detail above, the present invention is not limited to such specific embodiments, and various modifications and variations are possible within the scope of the gist of the present invention as described in the claims.

[0195] For example, aspects of the present invention are as follows. <1> a data search unit that searches for a plurality of data based on similarity with input information; a labeling unit that assigns a label to the set of data based on words contained in the set; a display information generating unit that generates display information that displays a name given by a user to the set in association with the label; An information processing system comprising: <2> a data search unit that searches for a plurality of data based on similarity with input information; a labeling unit that assigns a label to each of the data based on a word contained in the data; a display information generating unit that generates display information for each of the data items, the display information displaying the name of the data item and the label assigned to the data item in association with each other; An information processing system comprising: <3> the label assignment unit assigns a plurality of the labels; Characterized by <1> or <2> The information processing system described herein. <4> The labeling unit selects words to be used as the labels based on TF-IDF values ​​of the words included in the data. Characterized by <1> ~ <3> Any of the information processing systems described above. <5> the display information generation unit generates the display information so that the label is displayed surrounded by a line. Characterized by <1> ~ <4> Any of the information processing systems described above. <6> a classification unit that classifies the plurality of data into a plurality of classes; and the labeling unit assigns a label to each of the classes based on a word contained in the data belonging to the class; an association diagram generating unit that generates an association diagram showing the association between the label assigned to the class and the name of data belonging to the class; the relationship graph includes the labels as nodes; the display information generating unit generates display information for displaying the related diagram. Characterized by <1> ~ <5> Any of the information processing systems described above. <7> a receiving unit that receives a request to display a list of data sets, each of which includes one or more pieces of data; a labeling unit that calculates a TF-IDF value of each word included in data belonging to each of the datasets in response to the request, and assigns a label to each of the datasets based on the TF-IDF value; a display information generation unit that generates display information including the label for each of the data sets; An information processing system comprising: <8> The labeling unit calculates, for each of the datasets, a TF-IDF value between each word contained in data belonging to the dataset and each word contained in a name of a unit for classifying data in the dataset. Characterized by <7> The information processing system described herein. <9> a reception unit that receives an edit request for a data set, each of which includes one or more pieces of data; a labeling unit that calculates a TF-IDF value of each word included in data belonging to each of the datasets changed in response to the editing request, and assigns a label to each of the datasets based on the TF-IDF value; a display information generation unit that generates display information including the label for each of the data sets; An information processing system comprising: <10> a data search unit that searches a plurality of data based on similarity with input information; a labeling unit that calculates, for each dataset to which any of the plurality of datasets belongs, a TF-IDF value of each word included in the dataset, among datasets each of which includes one or more datasets, and assigns a label to the dataset based on the TF-IDF value and the input information; a display information generation unit that generates display information including the label for each of the data sets; An information processing system comprising: [Explanation of symbols]

[0196] 10. Information processing equipment 20 Information management device 21 Document Management Department 22 Document information storage unit 23 Employee information storage unit 24 Organization information storage section 25 Workspace storage 26 Conference information storage unit 30 User terminals 31 Display control unit 40 conferencing devices 100 Drive device 101 Recording media 102 Auxiliary storage device 103 Memory Device 104 processors 105 Interface Device 121 Reception 122 Vector conversion section 123 Comparison Section 124 Data Search Section 125 Classification Department 126 Labeling Unit 127 Related Diagram Generation Unit 128 Expert Collection Department 129 Workspace Collection Department 130 Display information generation section 131 Display information transmission unit 132 Workspace Generation Unit 133 Workspace Editorial Department 141 Document Vector Storage Unit 142 Document-related storage unit B Bus [Prior art documents] [Patent documents]

[0197] [Patent Document 1] Japanese Patent Publication No. 2022-20412

Claims

1. a data search unit that searches for a plurality of data based on similarity with input information; a labeling unit that assigns a label to the set of data based on words contained in the set; a display information generating unit that generates display information that displays a name given by a user to the set in association with the label; An information processing system comprising:

2. a data search unit that searches for a plurality of data based on similarity with input information; a labeling unit that assigns a label to each of the data based on a word contained in the data; a display information generating unit that generates display information for each of the data items, the display information displaying the name of the data item and the label assigned to the data item in association with each other; An information processing system comprising:

3. the label assignment unit assigns a plurality of the labels; 3. The information processing system according to claim 1 or 2.

4. the labeling unit selects words to be used as the labels based on TF-IDF values ​​of the words included in the data.

3. The information processing system according to claim 1 or 2.

5. the display information generation unit generates the display information so that the label is displayed surrounded by a line.

3. The information processing system according to claim 1 or 2.

6. a classification unit that classifies the plurality of data into a plurality of classes; and the labeling unit assigns a label to each of the classes based on a word contained in the data belonging to the class; an association diagram generating unit that generates an association diagram showing the association between the label assigned to the class and the name of data belonging to the class; the relationship graph includes the labels as nodes; the display information generating unit generates display information for displaying the related diagram.

3. The information processing system according to claim 1 or 2.

7. a receiving unit that receives a request for displaying a list of data sets, each of which includes one or more pieces of data; a labeling unit that calculates a TF-IDF value of each word included in data belonging to each of the datasets in response to the request, and assigns a label to each of the datasets based on the TF-IDF value; a display information generation unit that generates display information including the label for each of the data sets; An information processing system comprising:

8. the labeling unit calculates, for each of the datasets, a TF-IDF value between each word contained in data belonging to the dataset and each word contained in a name of a unit for classifying data in the dataset; 8. The information processing system according to claim 7.

9. a receiving unit that receives an edit request for a data set, each of which includes one or more pieces of data; a labeling unit that calculates a TF-IDF value of each word included in data belonging to each of the datasets changed in response to the editing request, and assigns a label to each of the datasets based on the TF-IDF value; a display information generation unit that generates display information including the label for each of the data sets; An information processing system comprising:

10. a data search unit that searches for a plurality of data based on similarity with input information; a labeling unit that calculates, for each dataset to which any of the plurality of datasets belongs, a TF-IDF value of each word included in the dataset, among datasets each including one or more datasets, and assigns a label to the dataset based on the TF-IDF value and the input information; a display information generation unit that generates display information including the label for each of the data sets; An information processing system comprising:

Citation Information

Patent Citations

  • Image processing device and information processing system

    JP2022020412A