Information processing device, information processing method, program, and information management system

The information processing apparatus addresses the inefficiency of requiring pre-set storage conditions by calculating variation information from document data to automatically determine storage location reconfiguration, enhancing the management of document data.

JP2025083026APending Publication Date: 2025-05-30RICOH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023196667
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Conventional techniques require users to set storage conditions in advance to determine whether to reconfigure storage locations, making it cumbersome and inefficient.

Method used

An information processing apparatus that calculates variation information from document data using vector information and determines whether to reorganize storage locations based on this information without pre-set storage conditions.

Benefits of technology

Enables automatic determination of storage location reconfiguration, eliminating the need for users to set conditions in advance and improving efficiency in managing document data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025083026000001_ABST
    Figure 2025083026000001_ABST
Patent Text Reader

Abstract

To provide an information processing device to determine whether to reconfigure a storage place without setting a storage condition and so on for the storage place in advance.SOLUTION: The present invention pertains to an information management device to store multiple document data associated with workspace information in a storage place and to manage it. The information processing device has a dispersion information control unit to calculate the dispersion information of the multiple document data stored in the storage place using vector information converted from the document data, and a determination unit to determine whether to reconfigure the storage place based on the dispersion information.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, an information processing method, a program, and an information management system.

Background Art

[0002] Regarding various operations, there are times when one wants to collect and utilize various document data such as materials useful for one's own operations, human resource information, or organizational knowledge. For example, if document data related to a development theme is stored in a database called a workspace, the document data used for development can be acquired from the workspace. When a user collects document data and stores it in the workspace, an appropriate storage location (for example, a folder) is assigned to the collected and added document data. However, depending on the type and quantity of the added document data, it may be necessary to review the overall folder configuration.

[0003] A technique for determining a storage destination folder based on a file name is known (see, for example, Patent Document 1). Patent Document 1 discloses a technique for presenting an appropriate folder storage location and determining whether to modify a folder hierarchy structure based on an added file name and folder storage conditions.

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in the conventional technique, unless the user has previously set storage conditions and the like for the storage location, a determination as to whether to reconfigure the storage location has not been made. In order to automatically determine whether to reconfigure the storage location, it has been necessary for the user to set storage conditions in advance, which has been troublesome.

[0005] In view of the above problems, the present invention provides a technique for determining whether to perform reconfiguration of a storage location without previously setting storage conditions and the like for the storage location.

Means for Solving the Problems

[0006] In view of the above problems, the present invention is an information processing apparatus that stores and manages a plurality of document data associated with workspace information in a storage location, and uses vector information converted from the document data to calculate variation information of the plurality of document data stored in the storage location, and a variation information control unit, and a determination unit that determines whether to perform reorganization of the storage location based on the variation information.

Effects of the Invention

[0007] The present invention can determine whether to perform reorganization of the storage location without previously setting storage conditions and the like for the storage location.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

[0009] Hereinafter, as an example of embodiments for carrying out the present invention, an information processing apparatus and an information processing method performed by the information processing apparatus will be described.

[0010] [First Embodiment] [Example of Judging the Necessity of Folder Reorganization] In this embodiment, an information processing apparatus that automatically determines the timing for reorganizing folders in a workspace will be described. The workspace is a database of the collection results (search results) of data collected using the information processing apparatus in the past and the data obtained by editing the collection results.

[0011] When a user collects and utilizes document data useful for business, the information processing apparatus recommends or actually adds document data to be added to the workspace. Document data gradually accumulates in the workspace, but it is not easy for the user to determine the criteria for reviewing the folder configuration of the workspace or adding new folders. This is because the user does not always grasp the overall picture of the workspace. Since there is no room to make the above determination during normal use, the user simply adds folders continuously, and the folder structure may gradually become distorted.

[0012] Therefore, in this embodiment, the information processing apparatus determines the timing for performing folder reorganization in consideration of the semantic dispersion (variation) of folders in the workspace.

[0013] FIG. 1 is a diagram for explaining an outline of a process for determining the timing of folder reorganization. (1) The information processing apparatus generates a document vector from the document data included in each folder and calculates the variation of the document vectors (calculates the dispersion). When the dispersion is large, it indicates that the folder contains diverse information, and when the dispersion is small, it indicates that the folder is composed only of specialized information. (2) Although details will be described later, the information processing apparatus 50 compares the magnitude of the variation of information for each folder with a threshold value or the like to determine the necessity (timing) for performing reorganization.

[0014] Therefore, according to this embodiment, it is possible to determine an appropriate timing for performing folder reorganization without setting the storage conditions for folders in advance.

[0015] <Regarding terms> Document data refers to the electronic data in which a document is recorded (hereinafter referred to as "document data"). A document is a collection of one or more words or sentences (and of course, may include alphanumeric characters and other languages). Document data may be in any form as long as it can represent sentences. For example, document data may be data that represents a document in text form, or data in a form specialized for a specific application. Or, document data may be data that represents a word or sentence itself, or a concept corresponding to a word or sentence, by means of an image, voice, or video (moving picture), etc. That is, document data may be image data, voice data, or video data. Furthermore, the storage format of document data is not limited to a specific one. For example, document data may be stored and saved in a file, may be stored as a record in a database, or may be stored in other forms. Document data may be, for example, in-house documents (PDF files, application files), minutes of in-house meetings, etc., which are used for business, or may be personal document data such as web pages and hobby images.

[0016] The storage location is information for accessing the stored document data. The storage location is, for example, a folder, a directory, etc. A folder or a directory is a place for classifying and storing electronic data in a computer. The storage location may also be a URL or a path.

[0017] The variation information is an index indicating to what extent a folder contains diverse information. Diverse information is information indicating to what extent a plurality of document data are diverse. In the present embodiment, the variation information may simply be referred to as variation. The variation information is obtained, for example, by obtaining a dispersion-covariance matrix from the document vectors into which the document data are converted, and performing principal component analysis on the dispersion-covariance matrix.

[0018] A workspace is a collection result (search result) of document data that has been processed using an information processing apparatus in the past, or a database in which the collection result has been edited. Workspace information is information regarding the workspace, and includes information regarding the folder structure and the document data stored in each folder.

[0019] Vector information is a quantity represented as a combination of a plurality of different values that are independent of each other. In the present embodiment, a document vector converted from document data will be described as an example.

[0020] Reconfiguration means relocating a plurality of document data included in one or more folders to the same number of folders, or to more or fewer folders than the original number, while keeping the number of folders the same. In the present embodiment, it is considered that the number of folders rarely decreases.

[0021] <System Configuration Example> In the present embodiment, as an example, a form in which a service providing company cooperates with a partner system to provide a customer with a document data search service via a network will be described. A link to the document data that matches the search is saved in the workspace. The service providing company manages the workspace in which the user saves the link to the document data, the folders in the workspace, and the link to the document data. Since the management of the link is considered in terms of security and the data capacity required for storage, the service providing company may manage the document data itself.

[0022] FIG. 2 is an example of a schematic configuration diagram of the information management system 100. The service providing company 103 can communicate with the customer system 101 and the partner system 102. In the service providing company 103, the information processing apparatus 50 described later performs folder reconfiguration and the like in the workspace described in the present embodiment.

[0023] The customer system 101 is an in-house system used by users (such as employees) within the customer. The customer system 101 includes a cloud system 121, edge devices 113, various databases 112, etc. The customer system 101 manages various data in the database 112. The data can be of various types, for example, human resources DB, meeting information, minutes of meetings, images, audio, documents, etc., as long as it is electronic data. Among these data, for the document data that the customer system 101 has requested management of, the service provider company 103 holds a link to the document data.

[0024] In addition, the customer system 101 has a user terminal 30. The user terminal 30 is a terminal operated by the users of the customer system 101 and can be connected to the information processing device 50. The user can display and view the document data managed by the information processing device 50 on the user terminal 30. The user can obtain document data such as document files, expert searches, etc., and workspace information from the information processing device 50 and utilize them in their work.

[0025] The linked system 102 is a system that cooperates regarding the document data provided by the information processing device 50 to the customer system 101. The linked system 102 uses, for example, external cooperation services 115 such as a remote conferencing system, and various databases 116 (expert DB, external organization DB, thesis DB, public document DB, library information, etc.) to provide the document data required by the customer. The information processing device 50 searches for the document data requested for search from the customer system 101 in the linked system 102, obtains the document data that matches the search, and provides it to the customer system 101. The user can search for the necessary document data from a group of data such as in-house documents like pdf and office files, and minutes of in-house meetings, and organize the document data.

[0026] FIG. 3 shows a system configuration example of a cloud-based information management system 100 configured based on the schematic configuration of FIG. 2. In FIG. 3, the information processing apparatus 50 of the service providing company 103 is provided as a cloud service. Cloud computing refers to a form in which resources on a network are used without awareness of specific hardware resources. Therefore, although the information processing apparatus 50 mainly exists on the Internet, it may also exist on the premises of the service providing company 103. Cloud computing can be in the form of SaaS (Software as a Service), PaaS (Platform as a Service), or IaaS (Infrastructure as a Service), and any form is acceptable. "On-premises" refers to a form in which hardware such as server network equipment and software are owned and operated by the company itself.

[0027] The customer system 101 includes a user terminal 30, a cloud system 121, a storage 122, and the like. The cloud system 121 communicates with the information processing apparatus 50 using a linked account with the information processing apparatus 50. The information processing apparatus 50 receives document data from the customer system 101 and converts it into irreversible data. The received document data is discarded. The information processing apparatus 50 retains only the converted irreversible data.

[0028] On the user terminal 30, a web browser is executed, and the user connects the web browser to the information processing apparatus 50 and requests the information processing apparatus 50 to search for document data. The information processing apparatus 50 returns a search result of the document data in response to the search request from the user terminal 30. Also, the user terminal 30 can be connected to the cloud system 121, and the user can upload and download document data to and from the cloud system 121. The user terminal 30 can download document data from the cloud system 121 based on the search result.

[0029] FIG. 4 shows a system configuration example of an on-premises information management system 100 configured based on the schematic configuration of FIG. 2. In FIG. 4, the information processing apparatus 50 of the service providing company 103 is provided as an on-premises server. The on-premises server may be, for example, within the LAN of the customer system 101.

[0030] Note that since the system configuration of FIG. 4 only changes the information processing apparatus 50 from a cloud service to an on-premises server, it is the same for the information processing apparatus 50 to hold converted irreversible data, for the user terminal 30 to search for document data with respect to the information processing apparatus 50, and for the user terminal 30 to upload or download document data.

[0031] FIG. 5 shows a system configuration example of a cloud / on-premises hybrid information management system 100 configured based on the schematic configuration of FIG. 2. In FIG. 5, the information processing apparatus 50 of the service providing company 103 is provided by being divided into two systems, a cloud service 50b and an on-premises server 50a.

[0032] Similar to FIG. 3, the cloud system 121 of the customer system 101 transmits the document data within the customer to the on-premises server 50a. The on-premises server 50a transmits the document data within the customer to the cloud service 50b. The cloud service 50b holds only the converted irreversible data, and the on-premises server 50a holds the converted irreversible data and the original data.

[0033] In the system configuration of FIG. 5, the cloud system 121 of the customer system 101 communicates with the on-premises server 50a using a linked account. The user terminal 30 communicates with the cloud service 50b using a web browser and requests the information processing apparatus 50 to search for document data. The information processing apparatus 50 returns a search result of the document data in response to the search request from the user terminal 30. Also, the user terminal 30 communicates with the cloud system 121 using a web browser to upload or download document data. The user terminal 30 can download the document data from the cloud system 121 based on the search result.

[0034] FIG. 6 shows a system configuration example of an installation application-form information management system 100 configured based on the schematic configuration of FIG. 2. The functions of the information processing apparatus 50 are installed in the user terminal 30 as an application. The cloud system 121 holds the document data of the customer system 101. The user operates the user terminal 30 to upload the document data to the cloud system 121 or download it from the cloud system 121. Further, the application can convert the document data of the customer system 101 into converted irreversible data and holds the converted irreversible data. The application can search for the converted irreversible data. Also, the user terminal 30 can access the cloud system 121 with the user's account and acquire the document data that matches the search from the cloud system 121.

[0035] <Hardware Configuration Example> With reference to FIG. 7, the hardware configurations of the information processing apparatus 50 and the user terminal 30 according to the present embodiment will be described.

[0036] <<Information Processing Apparatus and Terminal Apparatus>> FIG. 7 is a diagram showing an example of the hardware configurations of the information processing apparatus 50 and the user terminal 30 according to the present embodiment. As shown in FIG. 7, the information processing apparatus 50 and the user terminal 30 are constructed by a computer 500. The computer 500 includes a CPU 501, a ROM 502, a RAM 503, an HD (Hard Disk) 504, an HDD (Hard Disk Drive) controller 505, a display 506, an external device connection I / F (Interface) 508, a network I / F 509, a bus line 510, a keyboard 511, a pointing device 512, an optical drive 514, and a media I / F 516.

[0037] Among these, the CPU 501 controls the operations of the information processing apparatus 50 and the entire user terminal 30. The ROM 502 stores programs used for driving the CPU 501 such as the IPL. The RAM 503 is used as a work area for the CPU 501. The HD 504 stores various data such as programs. The HDD controller 505 controls the reading and writing of various data to and from the HD 504 according to the control of the CPU 501. The display 506 displays various information such as a cursor, menu, window, characters, or images. The external device connection I / F 508 is an interface for connecting various external devices. The external devices in this case are, for example, a USB (Universal Serial Bus) memory, a printer, and the like. The network I / F 509 is an interface for performing data communication using the network N2. The bus line 510 is an address bus, a data bus, etc. for electrically connecting each component such as the CPU 501 shown in FIG. 7.

[0038] Also, the keyboard 511 is a type of input means having a plurality of keys used for inputting characters, numerical values, or various instructions. The pointing device 512 is a type of input means for selecting and executing various instructions, selecting a processing target, moving a cursor, and the like. The optical drive 514 controls the reading and writing of various data to and from the optical storage medium 513 as an example of a removable recording medium. Note that the optical storage medium 513 may be a CD, DVD, Blu-Ray (registered trademark), or the like. The media I / F 516 controls the reading and writing (storage) of data to and from the recording medium 515 such as a flash memory.

[0039] <Regarding functions> <<Information processing apparatus>> FIG. 8 is an example of a functional block diagram of the information processing apparatus 50 according to the present embodiment. The information processing apparatus 50 includes an input unit 51, a search unit 52, a vector conversion unit 53, a semantic comparison unit 54, a reconstruction execution unit 55, a labeling control unit 56, a workspace control unit 57, a variation information control unit 58, a setting unit 59, a display screen generation unit 60, a communication unit 61, and a determination unit 62. Each of these units included in the information processing apparatus 50 is a function or a means for realizing the operation of any of the components shown in FIG. 7 by the instruction from the CPU 501 according to the program expanded from the HD 504 onto the RAM 503.

[0040] The input unit 51 receives character input from a keyboard or the like. The input unit 51 may receive voice input from a microphone or the like and convert it into text data on the own device or a predetermined server.

[0041] The search unit 52 receives a search query from the input unit 51 or the communication unit 61 and executes a search for document data. After generating the search result, the search unit 52 causes the display screen generation unit 60 to display it.

[0042] The vector conversion unit 53 vectorizes the document data to be searched and the input query into document vectors. By vectorizing, it becomes easier for the semantic comparison unit 54 and the like to calculate the similarity (such as Cos similarity or vector space distance) between the document data and the input query. For vectorization of document data, existing tools and methods can be used. For example, BERT (registered trademark. Bidirectional Encoder Representations from Transformers) is a deep learning model for natural language processing developed by Google and can convert natural language document data into document vectors. In addition, methods such as Doc2Vec may be used for vectorization of natural language.

[0043] The semantic comparison unit 54 compares the document vectors of the queries generated by the vector conversion unit 53 with the document vectors of the document data to be searched, thereby comparing the semantic similarity. The semantic comparison unit 54 can also compare the similarity between document data. As methods for comparing vectors, there are cosine similarity and a method using the distance in the vector space.

[0044] The reconstruction execution unit 55 performs clustering of the document data in consideration of the semantic proximity of the document data. As the clustering algorithm, any method such as k-means, hierarchical clustering, Ward method, EM algorithm, etc. may be used. A plurality of clusters are generated by clustering, and one cluster corresponds to one folder. The document data belonging to the cluster is stored in the corresponding folder.

[0045] The labeling control unit 56 attaches a representative label to the folder. In addition, the labeling control unit 56 can also attach labels to the clustered document data or the document data with high similarity. Examples of the algorithm used for labeling include TF-IDF.

[0046] The workspace control unit 57 creates, edits, manages, etc. the workspace based on the workspace information.

[0047] The variation information control unit 58 is a means for calculating the variation information of a plurality of document data stored in a folder (storage location) using the vector information converted from the document data. Specifically, the variation information control unit 58 utilizes the conversion of the document data in the folder into document vectors to calculate the semantic dispersion (variation) of each folder in the workspace in order to determine the necessity of performing folder reconstruction.

[0048] The setting unit 59 accepts various settings from the user, such as a threshold for determining whether to reconstruct the folder, and stores the setting information such as the threshold (see FIG. 31).

[0049] The display screen generation unit 60 causes various control results to be displayed on the user terminal 30 or the like. The display screen generation unit 60 provides this web app (screen information) to the user terminal 30 regarding the web app executed by the user terminal 30. When each screen is provided by a web page or a web app (the user terminal 30 executes a web browser), the screen information of various screens is described by HTML, XML, CSS (Cascade Style Sheet), JavaScript (registered trademark), and the like. When the user terminal 30 executes a native app installed thereon, the screen information is information displayed not by a web app, and the configuration of the screen is pre-owned by the native app. Further, the display screen generation unit 60 presents to the user the reason for reaching the folder reorganization.

[0050] The communication unit 61 transmits and receives various data (or information) to and from other terminals, devices, or systems (in this embodiment, the user terminal 30) via a communication network.

[0051] The determination unit 62 is a means for determining whether to perform the reorganization of the storage location based on the variation information. Specifically, the determination unit 62 determines whether to reorganize the folder. The determination unit 62 makes determinations such as comparing the variations between folders, comparing the variations between folder hierarchies, batch execution based on a threshold without comparison, or automatic execution based on the dispersion (variation) calculated by the variation information control unit 58. Further, the determination unit 62 determines whether to reorganize the folder by combining not only the variation information but also the number of document data in the folder and the variation information.

[0052] <<User Terminal>> The user terminal 30 includes a communication unit 31, a display control unit 32, and an operation reception unit 33. Each of these units is a function or a means for realizing the operation by an instruction from the CPU 501 according to a program developed from the HD 504 onto the RAM 503 among the components shown in FIG. 7.

[0053] The communication unit 31 transmits and receives various data (or information) to and from other terminals, devices, or systems via a communication network. The operation reception unit 33 receives various inputs from a user using a keyboard, a mouse, or a touch panel. The display control unit 32 causes the display 506 to display various images and screens.

[0054] <Calculation of Variation> The semantic dispersion (variation) of folders within a workspace will be described. In the present embodiment, as an example, the semantic dispersion of folders will be described as the dispersion of document vectors converted from the documents included in each folder. A large dispersion indicates that the folder contains diverse information within the workspace, and a small dispersion indicates that the folder is composed of only specialized information.

[0055]

Number

[0056]

Number

[0057]

Number

[0058] Next, the variation information control unit 58 expresses the variation of the document vectors by performing principal component analysis on the dispersion-covariance matrix document. For example, the eigenvalues and eigenvectors of the dispersion-covariance matrix can be obtained by principal component analysis.

[0059] The variation information control unit 58 calculates the variation using the contribution rate, using the eigenvalues and eigenvectors obtained from the dispersion-covariance matrix of the document vectors converted from the document data in the folder. In principal component analysis, the principal components are named the first principal component, the second principal component,..., the k-th principal component in descending order of the eigenvalues. The larger the eigenvalue, the better the principal component represents the original document data. The contribution rate is the eigenvalue of each principal component divided by the sum thereof. The variation information control unit 58 calculates, for example, the number of necessary eigenvalues (eigenvectors) (that is, the number of principal components) until a preset cumulative contribution rate is achieved as the variation of the workspace. The smaller the number of eigenvalues until the preset cumulative contribution rate is achieved, the smaller the variation, and the larger the number, the larger the variation.

[0060] The determination unit 62 determines the necessity of performing folder reorganization based on the magnitude of the variation for each folder in the workspace. Note that since there are multiple determination methods, the details will be described in each embodiment.

[0061] <Workspace Information> Referring to FIG. 9, the workspace information will be described. FIG. 9 is a diagram showing an example of the workspace information. The workspace information is composed of meta-information of the workspace itself (workspace name, creator, updater, workspace content, summary of the content, directory information, access right to the workspace, etc.) and information such as the document data included therein (link to the data, meta-information, text information of the document data, access right to the data, vector information, etc.). The workspace information may be constructed as a database or recorded in a json file or the like. Note that considering security, the meta-information, text information, etc. may be stored in the customer system 101.

[0062] The advantages of configuring a workspace include that the user can arbitrarily create and change the configuration of the workspace without depending on the location of the original data, etc., and accordingly, variation information can be calculated immediately even if the configuration is changed, etc. Therefore, it is also possible to associate the document vector of the document data with the original document data in advance.

[0063] Also, in this embodiment, since the variation information control unit 58 can calculate the variation using the document data itself rather than the document data name or tag information, more accurate semantic variation can be calculated. The fact that it is not necessary to set the conditions (attribute values) of each folder is also an advantage.

[0064] As shown in FIG. 9, the workspace information has, for each workspace, items such as a workspace ID, a workspace name, a label, a creator, an updater, an original search query, an evaluation score, included data, an included data label, and an included data path. · The workspace ID is identification information of the workspace. · The workspace name is the name of the workspace input by the user or automatically assigned. · The label is, for example, one or more words that are determined to be relatively important within the workspace using TF-IDF or the like in the set of document data belonging to the workspace. · The creator is identification information (such as a user ID or name) of the creator of the workspace. · The updater is identification information (such as a user ID or name) of the person who performed the update when the workspace is updated. · The original search query is the search query input in the collection of the document data that is the origin of the workspace. There may be multiple original search queries. Therefore, it can be said that the original search query is information indicating what kind of set of document data the workspace is based on. · The evaluation score is the value of the evaluation input by the user who referred to the workspace. For example, the average value of the numerical values in a five-level evaluation is the evaluation score. · The encapsulated data is the document data ID of each document data belonging to the workspace (it may also be a non-duplicate document data name). · The encapsulated data label is a label automatically assigned by the user or the system to the document data that is the encapsulated data. For example, the labeling control unit 56 or the user can assign a common category or the like to different document data. · The encapsulated data path is the file path of each document data for accessing each document data within the workspace. The path may be a URL.

[0065] <Operation or process> Next, with reference to FIG. 10, the process of determining the timing for folder reorganization during search execution will be described. FIG. 10 is a flowchart diagram illustrating an example of the process in which the information processing apparatus 50 reorganizes a folder based on the dispersion of document data within the folder. The process in FIG. 10 starts when document data is added to the workspace.

[0066] When the search unit 52 searches for document data, the workspace control unit 57 adds a link to the document data that matches the search to an appropriate location (folder) in an arbitrary workspace (S11). The destination folder for addition may be specified by the user or may be automatically determined. In the case of automatic determination, the workspace control unit 57 compares, for example, the input query with the search sentence that was the source of the workspace, and registers the document data that matches the search in the workspace. Also, the workspace control unit 57 determines the destination folder by judging the similarity between the document data that matches the search and the document data in each folder within the workspace. The workspace control unit 57 selects, for example, the folder with the smallest dispersion after saving, or the folder with the document data having the highest similarity, or the folder with the document data having the highest average similarity with a plurality of document data within the folder.

[0067] When document data is added to the workspace, the variation information control unit 58 repeats step S13 for the number of folders in the workspace (S12).

[0068] The variation information control unit 58 calculates the variation of each folder in the workspace where the document data is registered (S13). As described above, as an example, the vector conversion unit 53 converts the document data in the folder into a document vector. The variation information control unit 58 creates a scatter-covariance matrix based on the document vectors in the folder and performs principal component analysis. The variation information control unit 58 obtains the principal components of the scatter-covariance matrix and calculates the variation by counting the number of eigenvalues until the cumulative contribution rate becomes equal to or greater than the threshold value. Note that the vectorization of the document data may be performed in advance. The variation information of the folders to which no document data is added may be calculated in advance.

[0069] Next, the determination unit 62 repeats step S15 for the number of folders in the workspace (S14).

[0070] The determination unit 62 compares the variations between the folders (S15). For example, the determination unit 62 determines whether to perform at least the reorganization of the storage location where the variation information is larger based on the difference in the variation information between the folder to which the document data is added and other folders. That is, the determination unit 62 compares the variation of the folder to which the document data is added with the variation of the folder to which no document data is added. The comparison may be a pairwise comparison in which any two are taken out from all the folders in the workspace.

[0071] When the difference in the variation between all the folders is less than the threshold value (No in S16), the process proceeds to step S17. When there are folders whose difference in the variation is equal to the threshold value (Yes in S16), the process proceeds to step S18. In step S17, since there is no need to reorganize the folders, the process of FIG. 10 ends without reorganizing the folders.

[0072] In step S18, the reconstruction execution unit 55 reconstructs the folders in the workspace by clustering the document data included in the folders in the workspace (S17). The following folders can be the folders to be reconstructed. · The reconstruction execution unit 55 reconstructs all the folders in the workspace. For example, if there are three folders, folders A, B, and C are reconstructed into three or more folders. The reconstruction execution unit 55 increases the number of folders one by one until the difference in variation between all the folders becomes less than the threshold value. · For example, if document data is added to folder A and the difference in variation between folders A and B and between folders A and C is greater than or equal to the threshold value (A > B, A > C), the folder A with a large variation is reconstructed into two or more folders. · Reconstruct among multiple folders where the difference in variation between the folders is greater than or equal to the threshold value. For example, if document data is added to folder A and the difference in variation between folders A and B and between folders A and C is greater than or equal to the threshold value (A > B, A > C), at least one of reconstructing folders A and B into three or more folders or reconstructing folders A and C into three or more folders is performed.

[0073] The folder name is set by the labeling control unit 56 (for example, the word name with the largest TF-IDF in the folder, etc.).

[0074] Next, the display screen generation unit 60 creates a screen for displaying the folder reconstruction result, and the communication unit 61 transmits the screen information to the user terminal 30 (S18). The communication unit 31 of the user terminal 30 receives from the information processing device 50 that the determination unit 62 has determined to perform the folder reconstruction, and the display control unit 32 displays that the folder reconstruction is to be performed (folder reconstruction result). The folder reconstruction result is an inquiry as to whether the folder reconstruction may be performed, and further, it may include how the folder configuration will be by the folder reconstruction.

[0075] When an application is installed on the user terminal 30, the user terminal 30 creates and displays a screen for displaying the folder reorganization result.

[0076] <Specific example> Referring to FIG. 11, a specific example of folder reorganization will be described. FIG. 11 is a diagram for explaining a specific example of folder reorganization in a workspace having a plurality of folders.

[0077] First, assume that a workspace 130 named "Development of Function A" has the configuration shown in FIG. 11. The workspace 130 has a plan folder 132, a module a development folder 133, and a module b development folder 134 below the development folder 131 of Function A. The user is developing module a using the document data in the workspace 130 to realize Function A. This time, when developing the I / F of module a, the user searches for the document of the I / F specification and wants to add it to the workspace 130 named "Development of Function A".

[0078] The user operates the user terminal 30 to send a search request to the information processing apparatus 50 with "I / F specification of module a" as a query. The search unit 52 searches for document data that matches "I / F specification of module a" and returns the found document data named "Module a I / F Specification Document" to the user terminal 30. In addition, the search unit 52 adds a link to the document data that matches the search to the module a development folder 133 in the workspace 130. This is because, for example, the query "I / F specification of module a" or the document data name (for example, file name) "Module a I / F Specification Document" is similar to the encapsulated data "Development of module a" that the module a development folder 133 has.

[0079] FIG. 12 schematically shows the addition of a link of document data 135 named "I / F Specification of Module a" that matches the search to the module a development folder 133. The variation information control unit 58 calculates the variation of the module a development folder 133 to which the document data is added. When the difference between the variation of the module a development folder 133 and the variation of other folders exceeds the threshold value, the reconstruction execution unit 55 performs folder reconstruction. The reconstruction execution unit 55 may confirm the reconstruction execution determination with the user before reconstruction.

[0080] In this case, it is determined that the addition of the document data 135 named "Module a I / F Specification Document" to the module a development folder 133 has caused the variation between the module a development folder 133 and the module b development folder 134 to become too large compared to the planned folder 132 (assuming a brute-force comparison). The reconstruction execution unit 55 reconstructs the module a development folder 133 and the module b development folder 134. Since there are originally two folders, the reconstruction execution unit 55 specifies the number of clusters of k-means to be 3 or more.

[0081] FIG. 13 shows the folder configuration in the workspace 130 where the module a development folder 133 and the module b development folder 134 have been reconstructed. In FIG. 13, an I / F Specification folder 136 has been newly created. The folder name "I / F Specification" was determined as an important word by TF-IDF. The I / F Specification folder 136 is expected to be a folder containing document data common to the module a development folder 133 and the module b development folder 134.

[0082] Note that, instead of reconstructing the I / F specification folder 136 in parallel with the folders before reconstruction (module a development folder 133 and module b development folder 134) as shown in FIG. 13, the reconstruction execution unit 55 may create the I / F specification folder 136 in the module a development folder 133 and the module b development folder 134, respectively. In this case, the reconstruction execution unit 55 designates the number of clusters of k-means as 2 and reconstructs the module a development folder 133 and the module b development folder 134 separately.

[0083] <Example screen> Next, the screens displayed on the user terminal 30 will be described with reference to FIGS. 14 to 18. First, FIG. 14 shows a workspace details screen 150 displayed on the user terminal 30 before the "Module a I / F Specification Document" is added. The workspace details screen 150 has a bibliographic information column 151, a folder configuration column 152, and a file details column 153 within the folder.

[0084] In the bibliographic information column 151, the workspace name 154, the search sentence 155 that it was based on, the label 156, etc. are displayed from the items of the workspace information.

[0085] In the folder configuration column 152, the folder configuration within the workspace and the lists 157a to 157c of the document data within each folder are displayed. Also, in FIG. 14, the folder named "Development of Module a" is selected.

[0086] In the file details column 153 within the folder, details of the document data of the folder selected in the folder configuration column 152 are displayed. For example, in the file details column 153 within the folder, a thumbnail 161 of the document data, a file name 162, a label 163, a button 164 to open the folder, a download button 165, a details button 166, etc. are displayed. Note that in the file details column 153 within the folder, all the document data in the folder configuration column 152 are displayed by scrolling the screen.

[0087] FIG. 15 shows a message 170 that is displayed when the information processing apparatus 50 determines that the "Module a I / F Specification Document" matches the search and performs folder reorganization. In FIG. 15, on the workspace details screen 150, there is a message 170 that says "There is a better folder configuration plan for the 'Development of Function A'. Do you want to check the folder configuration plan?", a no button 172, and a yes button 173. When the user presses the yes button 172 (an example of a display component), the screen transitions to FIG. 16 (or FIG. 17 or FIG. 18 may also be used), and when the user presses the no button 173, the screen returns to FIG. 14.

[0088] As described above, when the determination unit 62 determines to perform folder reorganization, the display screen generation unit 60 displays a message 170 for confirming whether to perform folder reorganization and a display component (yes button) for receiving whether to perform folder reorganization.

[0089] FIG. 16 is a workspace details screen 180 that is displayed when the user presses the yes button 172. In FIG. 16, a newly created I / F specification folder 157d is added to the folder configuration column 152. When the user selects the I / F specification folder 157d, two document data 179a and 179b arranged in the I / F specification folder 157d are displayed in the in-folder file details column 153. One of the two document data 179a and 179b is the "Module a I / F Specification Document", and the other is the "Module b I / F Specification Document". It can be seen that before the reorganization, the "Module b I / F Specification Document" was in the module b development folder 134, but due to the reorganization, it was automatically arranged in the I / F specification folder 157d with less variation.

[0090] FIG. 17 shows an example of a display of a change comparison view before and after folder reorganization. FIG. 17 is displayed, for example, instead of FIG. 15 or following FIG. 15. In FIG. 17, a message 174 "Do you want to change the folder configuration as follows?", a folder configuration 175 before folder reorganization, and a folder configuration 176 after folder reorganization are shown. The user can check the folder configurations before and after folder reorganization and determine whether to perform folder reorganization. When the user presses the Yes button 177, the screen transitions to FIG. 16, and when the user presses the No button 178, the screen returns to FIG. 14. In this way, the display screen generation unit 60 displays the structure of the folder before reorganization and the structure of the folder after reorganization.

[0091] Furthermore, as shown in FIG. 18, the degree of variation may be indicated by color. FIG. 18 shows an example of a display of a change comparison view before and after folder reorganization. In FIG. 18, each folder is colored with a color corresponding to the magnitude of the variation. The folder configuration is the same as in FIG. 17.

[0092] In addition, a gradation 185 representing the relationship between the color and the magnitude of the variation is displayed. Since the user can determine which folder has a large variation by color, it becomes easier to determine whether to perform folder reorganization. Note that not only color but also, or instead of color, the magnitude of the variation may be displayed numerically. Also, instead of color, the magnitude of the variation may be displayed by the type and thickness of the frame line.

[0093] <Configuration change history> FIG. 19 shows a folder configuration history screen 300 displayed on the user terminal 30. When the user performs a predetermined operation from the workspace details screen 150 or other screens, the folder configuration history screen 300 is displayed. Each item of the folder configuration history screen 300 will be described. · The No item is a serial number assigned when folder reorganization is performed. · The Workspace item is the name of the reorganized workspace. · The item of the change content indicates the change that occurred in the folder due to the folder reorganization. The change content includes folder addition, folder change, folder hierarchy change, folder reduction, etc. Folder addition means that a new folder is added. Folder change means that multiple folders are reorganized (regardless of the increase or decrease in the number of folders). Folder hierarchy change means that a folder reorganization that changes the folder hierarchy has been performed. Folder reduction means that a folder has been reduced. By pressing the detailed link 181, for example, the folder configuration before and after the folder reorganization, and folders with large variations before the reorganization, etc. are displayed (see FIGS. 17 and 18). When the folder reorganization is thus performed, the change history of the folder configuration is displayed, and the change history includes any of folder addition, change, and hierarchy change. · The item of the execution date and time is the date and time when the folder reorganization was performed. · The item of the executor is the name of the user who permitted the folder reorganization proposed by the information processing apparatus 50. Automatic (system) indicates that the information processing apparatus 50 automatically performed the folder reorganization without asking the user, as will be described later.

[0094] <Main effects> According to this embodiment, the information processing apparatus 50 can determine whether to perform a folder reorganization without the user having to set storage conditions for the folder in advance. The user can perform the folder reorganization according to the proposal from the information processing apparatus 50.

[0095] [Second Embodiment] In this embodiment, an information processing apparatus 50 that collectively determines whether to perform a folder reorganization based on a threshold instead of the difference in variation between folders will be described.

[0096] In this embodiment, it will be described assuming that the hardware configuration diagram of FIG. 7 and the functional block diagram shown in FIG. 8 described in the first embodiment can be used.

[0097] FIG. 20 is a flowchart for explaining a process in which the information processing apparatus 50 reconstructs a folder based on the variation in document data within the folder. In the description of FIG. 20, the differences from FIG. 10 will mainly be explained. In FIG. 20, the variations between folders are not compared. Also, the determination method in step S24 is different from the determination method in step S16 of FIG. 10.

[0098] In step S24, the determination unit 62 determines whether there is a folder with a variation equal to or greater than a threshold value (S24). For example, if the variation information of the folder where the document data has been added exceeds the threshold value, the determination unit 62 determines that at least the reconstruction of the storage location where the variation information exceeds the threshold value is to be performed.

[0099] If the variations of all folders are all less than the threshold value (No in S24), the process proceeds to step S25. If there is a folder with a variation equal to or greater than the threshold value (Yes in S24), the process proceeds to step S26.

[0100] In step S26, the reconstruction execution unit 55 reconstructs the folder. The folders that can be the targets of reconstruction are as follows. · The reconstruction execution unit 55 reconstructs all folders in the workspace. · Only extract the folders with variations equal to or greater than the threshold value and reconstruct with these folders. For example, if the variation of folder B is equal to or greater than the threshold value, reconstruct folder B into two or more folders. Also, for example, if the variations of folders B and C are equal to or greater than the threshold value, reconstruct folders B and C into two or more respectively, or reconstruct folders B and C together into three or more folders.

[0101] The process in step S27 may be the same as the process in step S19 of FIG. 10.

[0102] <Main effects> According to the present embodiment, in addition to the effects of the first embodiment, if there is a folder with a large variation, the folder can be reconstructed without comparing the variations with other folders.

[0103] [Embodiment 3] In this embodiment, an information processing apparatus 50 that presents reasons leading to folder reorganization will be described.

[0104] FIG. 21 is an example of a functional block diagram of the information processing apparatus 50 and the user terminal 30 in this embodiment. In the description of FIG. 21, the differences from FIG. 8 will mainly be described. The information processing apparatus 50 in this embodiment newly has a word creation unit 63. The word creation unit 63 approximates the principal components to word vectors and creates words from the word vectors. Also, the word creation unit 63 compares the labels obtained based on the TF-IDF of the document data in the folder before and after reorganization, and acquires labels (words) that are different before and after reorganization.

[0105] FIG. 22 is a flowchart for explaining the process in which the information processing apparatus 50 reorganizes a folder based on the variation in the document data in the folder. Note that in the description of FIG. 22, the differences from FIG. 10 will mainly be described. In FIG. 22, the processes of steps S31 to S38 are the same as those of steps S11 to S18 in FIG. 10.

[0106] In step S39, the word creation unit 63 approximates the principal components of the folder (the one with the larger variation) having a difference in variation equal to or greater than the threshold value to word vectors before folder reorganization. For example, the developer prepares word vectors representing each word in advance. And, y = ax1 + bx2 + cx3 is set. This is with the vector representing the first principal component as y (word vector). Therefore, a to c are the eigenvectors corresponding to the eigenvalues of the first principal component. x1 to x3 are the original axes. Assume that y is a word indicating the factor of reorganization. The word creation unit 63 approximates y by a linear combination of the prepared word vectors. That is, the word creation unit 63 finds a linear combination of word vectors that minimizes the error with y. Alternatively, the word creation unit 63 finds a word vector with the highest similarity to y. The word creation unit 63 may perform the same processing for the second principal component and below. The similarity between vectors is determined by Cos similarity, vector space distance, etc.

[0107] Also, the word creation unit 63 compares the labels obtained by TF-IDF of the document data in the folder before and after the reconstruction. If there are different labels (words) before and after the reconstruction, the word may be regarded as a word indicating a factor for reconsideration. For example, for a folder with large variations, if the top words with the largest TF-IDF detected before the reconstruction are "development, test", and the top words with the largest TF-IDF detected in the folder added by the reconstruction are "test", then since "test" may have increased the variations, "test" is extracted.

[0108] Next, the display screen generation unit 60 generates a screen of the folder reconstruction result using the word that is the factor for the reconstruction, and the communication unit 61 transmits it to the user terminal 30 (S40). The communication unit 31 of the user terminal 30 displays the screen information, and the display control unit 32 displays the folder reconstruction result including the reason for reaching the folder reconstruction.

[0109] <Example of screen> FIG. 23 shows a message 194 including the reason for reaching the folder reconstruction, which is displayed on the user terminal 30 in step S40. In FIG. 23, the message 194 is pop-up displayed on the workspace details screen 150. The message 194 is "The amount of information contained in the folder of 'Module a development' of 'Development of Function A' is larger than that of other folders. There is a better folder configuration plan. Do you want to check the folder configuration plan?" That is, the message 194 in FIG. 23 shows which folder has large variations. The user can judge whether to reconstruct the folder considering the use of this folder, etc. When the user presses the yes button 195, it transitions to FIG. 16 (or FIG. 17 or FIG. 18 may also be used), and when the user presses the no button 196, it returns to FIG. 14. In this way, when the determination unit 62 determines to perform the reconstruction of the folder, the display screen generation unit 60 displays the reason for reaching the reconstruction.

[0110] FIG. 24 shows a message including words as reasons for folder reorganization, which is displayed by the user terminal 30 in step S40. FIG. 24 is another example of FIG. 23. Message 191 in FIG. 24 is "The amounts of 'design specifications' and 'I / F specifications' as information included in the folder of 'development of module a' for 'development of function A' are large, and it has become more complex than other folders. There is a better folder configuration plan. Do you want to check the folder configuration plan?" That is, "design specifications" and "I / F specifications" are words (words that have caused large variations) indicating factors for reorganization. In this way, the display screen generation unit 60 displays words that have brought about variation information equal to or greater than the threshold value as the reason for the reorganization.

[0111] Therefore, it is presumed that when a folder containing a large amount of document data including these words is created by folder reorganization. The user can determine whether to perform folder reorganization in consideration of whether to create a folder containing a large amount of document data including "design specifications" and "I / F specifications". When the user presses the yes button 192, the screen transitions to FIG. 16 (or FIG. 17 or FIG. 18 may also be used), and when the user presses the no button 193, the screen returns to FIG. 14.

[0112] <Main effects> According to the present embodiment, in addition to the effects of the first embodiment, the user can know the reason for reaching the folder reorganization, and it becomes easier to determine whether to perform the folder reorganization.

[0113] [Fourth Embodiment] In the present embodiment, an information processing apparatus 50 that determines whether to perform folder reorganization based on the difference in variation between folders with different hierarchies will be described. By doing so, when there is a difference in variation between folder hierarchies, it is possible to perform reorganization including the folder hierarchy.

[0114] In the present embodiment, it will be described assuming that the hardware configuration diagram of FIG. 7 and the functional block diagram of FIG. 8 described in the first embodiment can be used.

[0115] FIG. 25 is a flowchart for explaining the process of the information processing apparatus 50 reconfiguring the folder hierarchy based on the variation in document data within a folder. In the description of FIG. 25, the differences from FIG. 10 will mainly be explained. The processes of steps S41 to S44 may be the same as those of steps S11 to S14 in FIG. 10.

[0116] Next, in step S45, the determination unit 62 compares the variations between the folders including the hierarchy (S45). The comparison of the variations between the folders including the hierarchy will be explained with reference to FIG. 26.

[0117] The determination unit 62 determines whether there is a folder in which the difference in variation between folder hierarchies is equal to or greater than the threshold value (S46). That is, the determination unit 62 determines whether to perform the reconfiguration of a plurality of folders in the hierarchical structure based on the difference in variation information between one folder and the other folder in the hierarchical structure. If the difference in variation between all folder hierarchies is less than the threshold value (No in S46), the process proceeds to step S47. If there is a folder in which the difference in variation between folder hierarchies is equal to or greater than the threshold value (Yes in S46), the process proceeds to step S48.

[0118] In step S48, the reconfiguration execution unit 55 reconfigures the folders (S48). As the folder hierarchy to be reconfigured by the reconfiguration execution unit 55, the following folder hierarchies can be selected. · The reconfiguration execution unit 55 reconfigures all the folders within the workspace. · Only the folder hierarchy in which there is a folder with a difference in variation between folder hierarchies equal to or greater than the threshold value is reconfigured. For example, if the variation in the lower-level folders is larger, only the lower level is reconfigured. If the variation in the upper-level folders is larger, only the upper level is reconfigured (the lower level is connected to an arbitrary upper level) or the upper level and the lower level are combined and reconfigured into a folder with three or more levels. The subsequent processes may be the same as those in FIG. 10.

[0119] <Difference in variation of folder hierarchy> FIG. 26 is a diagram for explaining the difference in variation between folder hierarchies in a workspace. There are six folders 201 to 206 in this workspace, and folders 201 and 202, folders 203 and 204, and folders 205 and 206 respectively constitute hierarchies 207 to 209. For example, focusing on hierarchy 207, the variation information control unit 58 calculates the variations of folders 201 and 202 respectively. Therefore, the determination unit 62 compares the difference in variation between the two folders 201 and 202. The same applies to hierarchies 208 and 209.

[0120] When the difference in variation between the two folders 201 and 202 is equal to or greater than the threshold value, the reconstruction execution unit 55 may reconstruct only the folders 201 and 202 in hierarchy 207, or may reconstruct the folders 201 to 206. When the variation of folder 201 < the variation of folder 202, the reconstruction execution unit 55 may reconstruct only folder 202. When the variation of folder 201 > the variation of folder 202, the reconstruction execution unit 55 may reconstruct only folder 201, or may reconstruct folders 201 and 202 together. In the latter case, the use of hierarchical clustering in the fifth embodiment is preferable.

[0121] For example, when there are two folders 202 under folder 201, the determination unit 62 compares the difference in variation between folders 201 and 202 respectively, and when the difference in either one is equal to or greater than the threshold value, it determines to reconstruct the folders.

[0122] <Main effects> According to this embodiment, in addition to the effects of the first embodiment, when there is a difference in variation between folder hierarchies, the folder hierarchy including can be reconstructed.

[0123] [Fifth Embodiment] In this embodiment, an information processing apparatus 50 that clusters document data by hierarchical clustering will be described.

[0124] FIG. 27 is a flowchart for explaining the process of the information processing apparatus 50 reconfiguring a folder based on the variation in document data within the folder. Note that in the description of FIG. 27, the differences from FIG. 25 will mainly be explained. In FIG. 27, steps S61, S62, and S63 are added to FIG. 25.

[0125] When there is a folder hierarchy where the difference in the variation of the folder hierarchy is equal to or greater than the threshold value, in step S61, the reconstruction execution unit 55 performs hierarchical clustering (S61). That is, when the determination unit 62 determines to perform the reconstruction of a plurality of folders in the hierarchical structure, the reconstruction of the plurality of folders in the hierarchical structure can be performed by hierarchical clustering.

[0126] Literally, by performing hierarchical clustering, a plurality of folders having a hierarchy are reconfigured. Hierarchical clustering is a method of extracting the hierarchical structure of clusters, and it is a method of successively grouping the closest document data among the document data and gradually reducing the number of clusters. Therefore, by hierarchical clustering, it is possible to reconfigure the folder hierarchy so that folders containing document data with similar contents are gradually grouped into larger folders. There are various methods for hierarchical clustering, such as the centroid method, the group average method, and the Ward method, and any method can be used.

[0127] As the folder hierarchy to be reconfigured by the reconstruction execution unit 55, the following types of folder hierarchies can be selected. · The reconstruction execution unit 55 reconfigures all the folders within the workspace. · Only the folders where the difference in the variation between folder hierarchies is equal to or greater than the threshold value are reconfigured. For example, if the variation in the lower-level folders is larger, only the lower level is reconfigured. If the variation in the upper-level folders is larger, only the upper level is reconfigured (the lower level is connected to an arbitrary upper level) or the upper level and the lower level are reconfigured together. · The folder hierarchy where the difference in the variation between folder hierarchies is equal to or greater than the threshold value is reconfigured. For example, if the variation in the lower-level folders > the upper level (or vice versa), the upper level and the lower level are reconfigured together.

[0128] The reconstruction execution unit 55 reconstructs the folder configuration in the workspace using the result of hierarchical clustering (S62).

[0129] Next, the display screen generation unit 60 creates a screen for displaying the reconstruction result between folder hierarchies, and the communication unit 61 transmits the screen information to the user terminal 30 (S63).

[0130] <Main effects> According to the present embodiment, in addition to the effects of the fourth embodiment, when there are differences in the variations between folder hierarchies, a plurality of folders constituting the hierarchy can be reconstructed by hierarchical clustering, so that the folder hierarchy can be reconstructed when there is a hierarchical relationship in the document data within the folder.

[0131] [Sixth Embodiment] In the present embodiment, an information processing apparatus 50 that determines whether to reconstruct folders using both the difference in variations between folders and the number of document data within a folder will be described.

[0132] FIG. 28 is an example of a functional block diagram of the present embodiment. In the description of FIG. 28, mainly the differences from FIG. 8 will be described. The information processing apparatus 50 of the present embodiment newly has a document data number counting unit 64. The document data number counting unit 64 counts the number of document data within a folder for each folder.

[0133] FIG. 29 is a flowchart for explaining the process of the information processing apparatus 50 reconstructing folders based on the variation in document data within a folder and the number of document data. In the description of FIG. 29, mainly the differences from FIG. 10 will be described. In FIG. 29, step S73 is added among steps S71 to S76.

[0134] In step S73, the document data number counting unit 64 counts the number of document data within a folder for each folder (S73). Subsequently, steps S74 to S76 may be the same as steps S13 to S15 in FIG. 10.

[0135] Then, in step S77, the determination unit 62 determines whether there is a folder with a difference in variation between folders equal to or greater than a threshold value, or whether the number of document data in a folder is equal to or greater than the threshold value (S77). That is, the determination unit 62 determines whether to perform reorganization of at least the folder with a larger variation information or at least the folder with the number of document data equal to or greater than the threshold value based on the variation information and the number of document data stored in the folder. Instead of the difference in variation between folders, the variation of the folder may be compared with the threshold value.

[0136] If the difference in variation between all folders is less than the threshold value and the number of document data in the folder is less than the threshold value (No in S77), the process proceeds to step S78. If the difference in variation between folders is equal to or greater than the threshold value or the number of document data in the folder is equal to or greater than the threshold value (Yes in S77), the process proceeds to step S79.

[0137] In step S79, the reorganization execution unit 55 reorganizes the folders in the workspace by clustering the document data included in the folders in the workspace (S79). The following folders may be the folders to be reorganized. · The reorganization execution unit 55 reorganizes all the folders in the workspace. · For example, when document data is added to folder A and the difference in variation between folders A and B and between folders A and C is equal to or greater than the threshold value (A > B, A > C), the folder A with a large variation is reorganized into, for example, two or more folders. · Reorganize multiple folders with a difference in variation between folders equal to or greater than the threshold value. For example, when document data is added to folder A and the difference in variation between folders A and B and between folders A and C is equal to or greater than the threshold value (A > B, A > C), at least one of reorganizing folders A and B into three or more folders or reorganizing folders A and C into three or more folders is performed. · Reorganize only the folders with the number of document data in the folder equal to or greater than the threshold value. The process of step S80 may be the same as step S19 in FIG. 10.

[0138] <Main effects> According to the present embodiment, in addition to the effects of the first embodiment, since the folder can be reconfigured when the number of document data in the folder is large, the folder can be reconfigured even if the difference in variation is small.

[0139] [Seventh Embodiment] In the present embodiment, an information processing apparatus 50 that performs folder reorganization without notifying the user will be described by the user setting a high value for the threshold for determining whether to perform folder reorganization.

[0140] In the present embodiment, the hardware configuration diagram of FIG. 7 and the functional block diagram shown in FIG. 8 described in the first embodiment will be used for the description.

[0141] FIG. 30 is a flowchart for explaining the process of the information processing apparatus 50 reconfiguring a folder based on the variation of document data in the folder. In the description of FIG. 30, the differences from FIG. 20 will be mainly described. Steps S91 to S96 in FIG. 30 are the same as steps S21 to S26 in FIG. 20, but the determination method in step S94 is different from that in FIG. 20.

[0142] In step S94, the determination unit 62 determines whether the variation of each folder is equal to or greater than the automatic execution threshold (S94). If the variation of all folders is less than the automatic execution threshold (No in S94), the process proceeds to step S95. If the variation of a specific folder is equal to or greater than the automatic execution threshold (Yes in S94), the process proceeds to step S96.

[0143] Then, the reorganization execution unit 55 performs folder reorganization (S96). That is, when the determination unit 62 determines to perform folder reorganization, the reorganization execution unit 55 performs folder reorganization without querying the user. The folders to be reorganized may be the same as those in FIG. 20.

[0144] In addition, the display screen generation unit 60 does not generate a screen for the folder reorganization result. That is, since the variation within the folder was larger than the automatic execution threshold value, the folder could be reorganized without notifying the user terminal 30.

[0145] <Main effects> According to this embodiment, in addition to the effects of the first embodiment, when the variation of the folder is so large that it is recommended to be reorganized, the folder can be reorganized without notifying the user.

[0146] [Eighth Embodiment] In this embodiment, an information processing apparatus 50 in which a user can set details about parameters (for example, thresholds) for folder reorganization will be described.

[0147] FIG. 31 shows a folder reorganization setting screen 240 displayed on the user terminal 30. The setting items of the folder reorganization setting screen 240 will be described. · The item 241 for the scale of change is an item for the user to set the scale of change for folder reorganization (an example of scale-of-change information). The scale of change refers to whether it is a small-scale one that adds (or simply deletes) while maintaining the existing folders in the workspace, a large-scale one that creates a plurality of new folders from a plurality of folders in the workspace, and so on. There are options of large, medium, and small for the scale of change, which the user can select. For example, when "large" is selected, all the folders in the workspace are reorganized, and when "small" is selected, only the folders with large variations are reorganized. When "medium" is selected, for example, about the upper half of the folders with large variations are reorganized.

[0148] In this way, scale-of-change information regarding the scale of the folders to be reorganized in the workspace is preset, and the reorganization execution unit 55 performs the reorganization of the storage location according to the scale-of-change information.

[0149] · The item 242 of the variation difference detection threshold between folders is an item for the user to set the magnitude of the threshold value referred to in step S16 of FIG. 10 and the like. That is, the display screen generation unit 60 displays the folder reorganization setting screen 240 that accepts the setting of the threshold value, and the determination unit 62 compares the difference in variation information between the folder to which the document data is added and other folders with the threshold value, and performs reorganization of at least the folder with larger variation information.

[0150] · The item 243 of the variation difference detection threshold between folder hierarchies is an item for the user to set the magnitude of the threshold value referred to in step S46 of FIG. 25 and the like. That is, the display screen generation unit 60 displays the folder reorganization setting screen 240 that accepts the setting of the threshold value, and the determination unit 62 compares the difference in variation information between one folder and the other folder in the hierarchical structure with the threshold value, and determines whether to perform reorganization of a plurality of folders in the hierarchical structure.

[0151] · The item 244 of the variation difference detection threshold for a single folder is an item for the user to set the magnitude of the threshold value referred to in step S24 of FIG. 20 and the like. That is, the display screen generation unit 60 displays the folder reorganization setting screen 240 that accepts the setting of the threshold value, and the determination unit 62 compares the variation information of the folder to which the document data is added with the threshold value, and determines whether to perform reorganization of at least the storage location where the document data is added.

[0152] · The item 245 of the automatic reorganization setting / threshold value is an item for the user to set whether to automatically perform folder reorganization (enabled / disabled) as in the seventh embodiment (S77 in FIG. 29), and when performing automatically, the magnitude of the automatic execution threshold value. That is, the display screen generation unit 60 displays the folder reorganization setting screen 240 that accepts the setting of the threshold value for performing folder reorganization without asking the user, and the determination unit 62 compares the variation information with the threshold value and determines whether to perform folder reorganization.

[0153] <Main effects> According to this embodiment, the user can set the threshold value for determining whether to perform folder reorganization.

[0154] [Other application examples] The present invention is not limited to the specifically disclosed embodiments above, and various modifications and changes are possible without departing from the scope of the claims.

[0155] For example, in this embodiment, the folder of the document data associated with the workspace information is reconfigured, but the folder and the document data may not be managed in the workspace. For example, the folders and directories in the PC of an individual user may be reconfigured.

[0156] Also, in this embodiment, the information processing apparatus 50 performs folder reconfiguration, but the user terminal 30 may perform folder reconfiguration.

[0157] Also, the workspace may hold the document data body instead of being a link to the document data.

[0158] Also, in this embodiment, it is determined whether to perform folder reconfiguration when the document data that matches the search is added to the workspace, but the document data added to the workspace is not limited to the document data that matches the search. For example, it may be determined whether to perform folder reconfiguration when the user arbitrarily adds document data to the workspace.

[0159] Also, the folder reconfiguration may be automatically executed not at the time of adding document data but at a time zone such as at night when the processing load of the information processing apparatus 50 is small.

[0160] Also, the configuration examples such as FIG. 8 are divided according to the main functions in order to facilitate the understanding of the processing by the information processing apparatus 50. The present invention is not limited by the way and name of the division of the processing units. The processing of the information processing apparatus 50 can be further divided into more processing units according to the processing content. Also, one processing unit can be divided so as to include more processing.

[0161] Each function of the embodiment described above can be realized by one or more processing circuits. Here, the "processing circuit" in this specification means a processor programmed to execute each function by software, such as a processor implemented by an electronic circuit, an ASIC (Application Specific Integrated Circuit) designed to execute each function described above, a DSP (digital signal processor), an FPGA (field programmable gate array), or a device such as a conventional circuit module.

[0162] The device group described in the embodiment only shows one of a plurality of computing environments for implementing the embodiments disclosed in this specification. In one embodiment, the server 40 includes a plurality of computing devices such as a server cluster. The plurality of computing devices are configured to communicate with each other via any type of communication link including a network or a shared memory, and implement the processing disclosed in this specification.

[0163] Furthermore, the information processing device 50 can also combine the disclosed processing steps in various ways. Each element of the information processing device 50 may be integrated into one device or divided into a plurality of devices. Also, each process performed by the information processing device 50 may be performed by the terminal device 10.

[0164] <Appendix> [Appendix 1] An information processing device that stores and manages a plurality of document data associated with workspace information in a storage location, a variation information control unit that calculates variation information of the plurality of document data stored in the storage location using vector information converted from the document data; a determination unit that determines whether to perform reconstruction of the storage location based on the variation information; An information processing device having the above. [Appendix 2] The information processing apparatus according to Additional Note 1, wherein the determination unit determines whether to perform reconstruction of at least the storage location with larger variation information based on a difference in the variation information between the storage location where the document data is added and another storage location. [Additional Note 3] The information processing apparatus according to Additional Note 1, wherein the determination unit determines to perform reconstruction of at least the storage location where the variation information exceeds a threshold value when the variation information of the storage location where the document data is added exceeds the threshold value. [Additional Note 4] The information processing apparatus according to any one of Additional Notes 1 to 3, further comprising a display screen generation unit that displays a reason for reaching the reconstruction when the determination unit determines to perform reconstruction of the storage location. [Additional Note 5] The information processing apparatus according to Additional Note 4, wherein the display screen generation unit displays a word that has brought about variation information equal to or greater than a threshold value as the reason for reaching the reconstruction. [Additional Note 6] Change scale information regarding the scale of the storage location to be reconstructed within the workspace is preset, The information processing apparatus according to any one of Additional Notes 1 to 5, further comprising a reconstruction execution unit that performs reconstruction of the storage location according to the change scale information. [Additional Note 7] The information processing apparatus according to Additional Note 4 or 5, wherein the display screen generation unit displays the structure of the storage location before reconstruction and the structure of the storage location after reconstruction. [Additional Note 8] When the determination unit determines to perform reconstruction of the storage location, The information processing apparatus according to Additional Note 4 or 5, wherein the display screen generation unit displays a message for confirming whether to perform reconstruction of the storage location and a display component for receiving whether to perform reconstruction of the storage location. [Additional Note 9] The information processing apparatus further comprising a display screen generation unit that displays a setting screen for receiving setting of a threshold value, The determination unit compares the difference in the variation information between the storage location where the document data is added and other storage locations with the threshold value, and determines whether to perform reorganization of at least the storage location with larger variation information, as described in Supplementary Note 2. [Supplementary Note 10] The determination unit determines whether to perform reorganization of a plurality of the storage locations in the hierarchical structure based on the difference in the variation information between one of the storage locations in the hierarchical structure and the other storage location, as described in Supplementary Note 1. [Supplementary Note 11] It has a display screen generation unit that displays a setting screen for receiving setting of a threshold value. The determination unit compares the difference in the variation information between one of the storage locations in the hierarchical structure and the other storage location with the threshold value, and determines whether to perform reorganization of a plurality of the storage locations in the hierarchical structure, as described in Supplementary Note 10. [Supplementary Note 12] When the determination unit determines to perform reorganization of a plurality of the storage locations in the hierarchical structure. It has a reorganization execution unit that performs reorganization of a plurality of the storage locations in the hierarchical structure by hierarchical clustering, as described in Supplementary Note 11. [Supplementary Note 13] The determination unit determines whether to perform reorganization of at least the storage location with larger variation information or at least the storage location with the number of document data stored therein being equal to or greater than a threshold value, based on the variation information and the number of document data stored in the storage location, as described in Supplementary Note 2. [Supplementary Note 14] When the determination unit determines to perform reorganization of the storage location, it has a reorganization execution unit that performs reorganization of the storage location without inquiring the user, as described in Supplementary Note 3. [Supplementary Note 15] It has a display screen generation unit that displays a setting screen for receiving setting of a threshold value for performing reorganization of the storage location without inquiring the user. The determination unit compares the variation information with the threshold value to determine whether to perform the reconstruction of the storage location, according to the information processing apparatus described in Supplementary Note 14. [Supplementary Note 16] It has a display screen generation unit that displays a setting screen for receiving the setting of the threshold value. The determination unit compares the variation information of the storage location where the document data is added with the threshold value to determine whether to perform at least the reconstruction of the storage location where the document data is added, according to the information processing apparatus described in Supplementary Note 3. [Supplementary Note 17] When the reconstruction of the storage location is performed, it displays the change history of the configuration of the storage location. The change history includes any one of addition, change, and hierarchical change of the storage location, according to the information processing apparatus described in Supplementary Note 1.

Explanation of Signs

[0165] 30 User terminal 50 Information processing apparatus 100 Information management system

Prior Art Documents

Patent Documents

[0166]

Patent Document 1

Claims

1. An information processing apparatus that stores and manages a plurality of document data associated with workspace information in a storage location, comprising: a variation information control unit that calculates variation information of the plurality of document data stored in the storage location using vector information converted from the document data; a determination unit that determines whether to perform reorganization of the storage location based on the variation information; An information processing apparatus having the above.

2. The information processing apparatus according to claim 1, wherein the determination unit determines whether to perform reorganization of at least the storage location with larger variation information based on a difference in the variation information between the storage location where the document data is added and another storage location.

3. The information processing apparatus according to claim 1, wherein the determination unit determines to perform reorganization of at least the storage location where the variation information exceeds a threshold when the variation information of the storage location where the document data is added exceeds the threshold.

4. The information processing apparatus according to any one of claims 1 to 3, further comprising a display screen generation unit that displays a reason for the reorganization when the determination unit determines to perform reorganization of the storage location.

5. The information processing apparatus according to claim 4, wherein the display screen generation unit displays a word that has brought about variation information equal to or greater than a threshold as a reason for the reorganization.

6. Change scale information regarding the scale of the storage location to be reorganized within the workspace is preset, The information processing apparatus according to claim 1, further comprising a reorganization execution unit that performs reorganization of the storage location according to the change scale information.

7. The information processing apparatus according to claim 4, wherein the display screen generation unit displays the structure of the storage location before reorganization and the structure of the storage location after reorganization.

8. When the determination unit determines to perform reorganization of the storage location, The information processing apparatus according to claim 4, wherein the display screen generation unit displays a message for confirming whether to perform reorganization of the storage location and a display component for receiving whether to perform reorganization of the storage location.

9. having a display screen generation unit that displays a setting screen for receiving setting of a threshold, The determination unit compares the difference in the variation information between the storage location where the document data is added and other storage locations with the threshold value, and determines whether to perform reorganization of at least the storage location with larger variation information. The information processing apparatus according to claim 2.

10. The determination unit determines whether to perform reorganization of the plurality of storage locations in the hierarchical structure based on the difference in the variation information between one storage location and the other storage location in the hierarchical structure. The information processing apparatus according to claim 1.

11. It has a display screen generation unit that displays a setting screen for receiving setting of a threshold value. The determination unit compares the difference in the variation information between one storage location and the other storage location in the hierarchical structure with the threshold value, and determines whether to perform reorganization of the plurality of storage locations in the hierarchical structure. The information processing apparatus according to claim 10.

12. When the determination unit determines to perform reorganization of a plurality of storage locations in the hierarchical structure. The information processing apparatus according to claim 11, further comprising a reorganization execution unit that performs reorganization of a plurality of storage locations in the hierarchical structure by hierarchical clustering.

13. The determination unit determines whether to perform reorganization of at least the storage location with larger variation information or at least the storage location with the number of document data equal to or greater than a threshold value based on the variation information and the number of document data stored in the storage location. The information processing apparatus according to claim 2.

14. When the determination unit determines to perform reorganization of the storage location, the information processing apparatus according to claim 3, further comprising a reorganization execution unit that performs reorganization of the storage location without querying the user.

15. It has a display screen generation unit that displays a setting screen for receiving setting of a threshold value for performing reorganization of the storage location without querying the user. The determination unit compares the variation information with the threshold value, and determines whether to perform reorganization of the storage location. The information processing apparatus according to claim 14.

16. It has a display screen generation unit that displays a setting screen for receiving setting of a threshold value. The determination unit compares the variation information of the storage location where the document data is added with the threshold value, and determines whether to perform reorganization of at least the storage location where the document data is added. The information processing apparatus according to claim 3.

17. When the reconstruction of the storage location is performed, display the change history of the configuration of the storage location, The information processing apparatus according to claim 1, wherein the change history includes any one of addition, change, and hierarchical change of the storage location.

18. An information processing method performed by an information processing apparatus that stores and manages a plurality of document data associated with workspace information in a storage location, A process of calculating variation information of the plurality of document data stored in the storage location using vector information converted from the document data, A process of determining whether to perform reconstruction of the storage location based on the variation information, An information processing method for performing.

19. An information processing apparatus that stores and manages a plurality of document data associated with workspace information in a storage location, A variation information control unit that calculates variation information of the plurality of document data stored in the storage location using vector information converted from the document data, A determination unit that determines whether to perform reconstruction of the storage location based on the variation information, A program for functioning as.

20. An information management system having an information processing apparatus that stores and manages a plurality of document data associated with workspace information in a storage location and a user terminal, The information processing apparatus, A variation information control unit that calculates variation information of the plurality of document data stored in the storage location using vector information converted from the document data, A determination unit that determines whether to perform reconstruction of the storage location based on the variation information, The user terminal, A communication unit that receives from the information processing apparatus a notification that the determination unit has determined to perform reconstruction of the storage location, A display control unit that displays a message indicating that the reconstruction of the storage location is to be performed, An information management system having.

Citation Information

Patent Citations

  • Document management method, program, and system

    JP2004110445A