Information processing device, information processing method, and program
The information processing device automates content classification using metadata generation and semantic clustering to enhance search efficiency by reducing time spent searching for documents.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- COMPUTER ENGINEERING & CONSULTING LTD
- Filing Date
- 2025-10-02
- Publication Date
- 2026-05-08
AI Technical Summary
Existing content search systems are time-consuming due to reliance on personalized classification methods, making it difficult to efficiently find relevant documents.
An information processing device that assigns physical addresses to content items, generates metadata, creates a vector space representing content characteristics, and classifies content into clusters based on semantic similarity, allowing for efficient virtual classification and display of results.
Reduces the time required to search for content by implementing automated and semantic-based classification, enhancing search efficiency.
Smart Images

Figure 0007855780000001_ABST
Abstract
Description
[Technical Field]
[0001] This invention relates to an information processing device, an information processing method, and a program. [Background technology]
[0002] The search server uses the entered keywords to retrieve documents related to those keywords from its stored documents as search results. The search server then provides the search results to the user's terminal device, allowing the user to view the results on their terminal device. A technique is known for clustering search results based on their degree of similarity (see, for example, Patent Document 1). [Prior art documents] [Patent Documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2005-078245 [Overview of the Initiative] [Problems that the invention aims to solve]
[0004] When storing content, the classification method is dependent on the organization or the person in charge (personalized), which makes it time-consuming to find the necessary documents.
[0005] The object of the present invention is to provide an information processing device, an information processing method, and a program that can reduce the time required to search for content. [Means for solving the problem]
[0006] One aspect of the present invention is an information processing device comprising: a receiving unit for receiving content; a content processing unit that retrieves multiple content items from a content pool stored in which information identifying a physical address is assigned to each of the multiple content items received by the receiving unit, based on each of the multiple pieces of information identifying a physical address; generates metadata based on each of the retrieved multiple content items; generates a vector space representing the characteristics of the content based on the generated multiple pieces of metadata; generates multiple clusters from the multiple content items based on the generated vector space and the semantic similarity of the content, thereby virtually classifying the content; and a display control unit that causes a display unit to display the results of the virtual classification of the content by the content processing unit.
[0007] One aspect of the present invention is an information processing method performed by a computer, which receives a plurality of contents, retrieves a plurality of contents from a content pool in which information identifying a physical address is assigned to each of the received plurality of contents and stored, based on each of the pieces of information identifying a physical address, generates metadata based on each of the retrieved plurality of contents, generates a vector space representing the characteristics of the contents based on the generated plurality of metadata, generates a plurality of clusters from the plurality of contents based on the generated vector space and the semantic similarity of the contents, thereby virtually classifying the contents, and displays the result of the virtual classification of the contents on a display unit.
[0008] One aspect of the present invention is a program that causes a computer to receive multiple content items, retrieve multiple content items from a content pool in which information identifying a physical address is assigned to each of the received multiple content items and stored, retrieve multiple content items based on each of the information identifying a physical address, generate metadata based on each of the retrieved multiple content items, generate a vector space representing the characteristics of the content based on the generated multiple metadata, generate multiple clusters from the multiple content items based on the generated vector space and the semantic similarity of the content, thereby virtually classifying the content, and displaying the results of the virtual classification of the content on a display unit. [Effects of the Invention]
[0009] According to the present invention, the time required to search for content can be reduced. [Brief explanation of the drawing]
[0010] [Figure 1] This figure shows an example of an information processing device according to this embodiment. [Figure 2] This figure shows an example of user attribute information. [Figure 3] This figure shows an example of content metadata. [Figure 4] This is a diagram illustrating an example of the processing performed by the information processing device of this embodiment. [Figure 5] This is a diagram illustrating an example of the processing performed by the information processing device of this embodiment. [Figure 6] This is a diagram illustrating an example of the processing performed by the information processing device of this embodiment. [Figure 7] This is a diagram showing an example of a tendrogram. [Figure 8] This figure shows an example of creator attribute information. [Figure 9] This is a flowchart showing an example of the operation flow of the information processing device of this embodiment. [Figure 10]It is a flowchart showing an example of the operation flow of the information processing apparatus according to the present embodiment. [Figure 11] It is a diagram for explaining an example of the processing of the information processing apparatus according to the present embodiment. [Figure 12] It is a flowchart showing an example of the operation flow of the information processing apparatus according to the present embodiment. [Figure 13] It is a flowchart showing an example of the operation flow of the information processing apparatus according to the present embodiment. [Figure 14] It is a flowchart showing an example of the operation flow of the information processing apparatus according to the present embodiment. [Figure 15] It is a flowchart showing an example of the operation flow of the information processing apparatus according to the present embodiment. [Figure 16] It is a flowchart showing another example of the operation flow of the information processing apparatus according to the present embodiment. [Figure 17] It is a diagram showing an example of a content classification screen. [Figure 18] It is a diagram showing an example of an information processing apparatus according to a modified example of the embodiment. [Figure 19] It is a diagram showing an example of a hierarchical structure. [Figure 20] It is a flowchart showing an example of the operation flow of an information processing apparatus according to a modified example of the embodiment. [Figure 21] It is a flowchart showing an example of the operation flow of an information processing apparatus according to a modified example of the embodiment. [Figure 22] It is a flowchart showing an example of the operation flow of an information processing apparatus according to a modified example of the embodiment. [Figure 23] It is a flowchart showing an example of the operation flow of an information processing apparatus according to a modified example of the embodiment. [
Mode for Carrying Out the Invention
[0011] Hereinafter, an information processing apparatus, an information processing method, and a program according to the embodiment will be described with reference to the drawings. The embodiments described below are merely examples, and the embodiments to which the present invention is applied are not limited to the following embodiments. In all the figures used to illustrate the embodiments, components with the same function are given the same reference numerals, and repeated explanations are omitted. Furthermore, in this application, "based on XX" means "based on at least XX," and includes cases where it is based on another element in addition to XX. Also, "based on XX" is not limited to cases where XX is used directly, but also includes cases where it is based on something that has been calculated or processed from XX. "XX" is any element (for example, any information). In this application, "to acquire" is not limited to actively acquiring information by sending a transmission request, but may also include acquiring information by passively receiving information transmitted from another device. Furthermore, "to acquire" is not limited to directly acquiring the target information (information to be acquired) from an external source, but may also include acquiring the target information by generating it through calculations or processing of information obtained from an external source.
[0012] (Embodiment) (Information processing device) The information processing device 100 in this embodiment creates management rules for content such as documents and manages the content based on the created management rules. Furthermore, the information processing device 100 classifies the content being managed based on the management rules. This classification may be performed automatically. Figure 1 shows an example of the information processing device 100 of this embodiment. The information processing device 100 is implemented by a device such as a personal computer, server, smartphone, tablet computer, or industrial computer.
[0013] The information processing device 100 receives content. The information processing device 100 assigns information that identifies the physical address to the received content and stores the content with the physical address identification information in the content pool. Examples of information that identifies the physical address include the physical address and an ID equivalent to the physical address. The following explanation continues with an example of information that identifies the physical address, specifically the case where the physical address is applied. The information processing device 100 virtually classifies the content stored in the content pool and displays the results of this virtual classification on the display unit. An example of the results of this virtual classification is a classification tree that hierarchically represents the content. Furthermore, the information processing device 100 searches for content stored in the content pool based on user specifications, for example. The information processing device 100 displays the search results for content on the display unit.
[0014] The information processing device 100 includes a user interface (UI) 10, an account registration / update processing unit 20, a reception unit 25, a content pool 30, a content processing unit 70, and a display control unit 66. The content processing unit 70 comprises a virtual content pool construction processing unit 40, a hierarchical structure generation agent 50, a pseudo-metadata search agent 62, a natural language search agent 64, and a virtual content pool 80.
[0015] UI10 comprises an account registration / update UI12, a content pool UI14, a classification display UI16, and a natural language search UI18. An example of UI10 includes a display section. The virtual content pool construction processing unit 40 constructs the virtual content pool 80. The virtual content pool construction processing unit 40 comprises a content pool crawler 42, a metadata generation agent 44, and an LLM (Large-Scale Language Model) 4.
[0016] The virtual content pool 80 comprises a user attribute database (DB:Data Base) 81, a metadata DB 82, a usage log DB 83, a vector space generation agent 84, an embedded processing unit 85, and a content vector space storage unit 86.
[0017] The account registration / update UI12 is a user interface (UI) for user U to register and update their account. The content pool UI14 is a user interface (UI) that allows user U to view the content contained in the content pool 30 and perform predetermined actions. The classification display UI16 is a UI for classifying and displaying content. For example, a classification display UI may be created for each organization or user, and the displayed content may be made viewable based on the user's specifications. The natural language search UI18 is a user interface that receives a search prompt from the user U and performs a search based on the received search prompt.
[0018] The account registration / update processing unit 20 obtains an account registration request from the account registration / update UI 12. Based on the user identification information and other information included in the obtained account registration request, the account registration / update processing unit 20 registers user attribute information in the user attribute DB 81.
[0019] Figure 2 shows an example of user attribute information. As shown in Figure 2, an example of user attribute information includes an employee number as an example of user identification information, name, affiliation data, information indicating job title, and job type code. An example of affiliation data includes the name of the business unit, business unit, department, section, job description, and department code. Returning to Figure 1, we will continue the explanation.
[0020] After user attribute information is registered in the user attribute DB 81, the account registration / update processing unit 20, upon receiving an account update request from the account registration / update UI 12, updates the user attribute information registered in the user attribute DB 81 based on the user identification information and other data included in the received account update request.
[0021] The reception unit 25 receives content and stores the received content in the content pool 30. Here, the content may be created by the information processing device 100, or it may be created outside the information processing device 100 and input into the information processing device 100. When content is created outside the information processing device 100 and input to the information processing device 100, the content may also be transmitted to the information processing device 100 from a terminal device (not shown) connected to the information processing device 100 via a communication network. In this case, the communication unit (not shown) of the information processing device 100 receives the content transmitted by the terminal device, and the receiving unit 25 acquires the content received by the communication unit. Here, the communication network includes the internet, WAN (Wide Area Network), LAN (Local Area Network), public lines, provider equipment, dedicated lines, wireless base stations, etc. The receiving unit 25 may or may not save the received content in the content pool 30. The following explanation will continue with an example where the receiving unit 25 saves the received content in the content pool 30. The content pool 30 stores one or more content items in the information processing device 100. One or more content items can be displayed on the UI 10. The content pool 30 assigns a physical address to each of the one or more content items stored. Each of the one or more content items remains stored in the content pool 30 and is displayed on the UI 10. The content pool 30 notifies the virtual content pool construction processing unit 40 of the physical addresses assigned to the stored content items.
[0022] The virtual content pool construction processing unit 40 constructs the virtual content pool 80. The content pool crawler 42 receives physical addresses notified from the content pool 30. The content pool crawler 42 also checks the update status of the content in the content pool 30 and newly added content to collect (crawl) physical addresses. For example, the content pool crawler 42 polls the content pool 30 for updates and newly added content to collect physical addresses. In addition, the content pool crawler 42 collects physical addresses when information indicating updates and newly added content is notified (pushed) from the content pool 30 at regular intervals. The metadata generation agent 44 obtains a physical address from the content pool crawler 42 and identifies the content corresponding to that physical address from the content pool 30 based on the obtained physical address. The metadata generation agent 44 generates metadata from the identified content. For example, the metadata generation agent 44 uses an LLM (Large-Scale Language Model) 4 to generate metadata from the content. The metadata generation agent 44 stores the generated metadata in the metadata DB 82, associating it with content identification information.
[0023] Figure 3 shows an example of content metadata. In Figure 3, attribute items, descriptions, methods for automatically extracting metadata by an AI agent, character count, number of embedding dimensions, subspaces, chunk embedding / category, etc., are associated with each category. Here, the content metadata may be associated with some of the above information, for example, attribute items, descriptions, methods for automatically extracting metadata by an AI agent, and character count, for each category. As shown in Figure 3, an example of content metadata includes categories such as basic information, content information, permissions and access information, and tracking and management information.
[0024] Examples of basic information include title, creator, creation date, last updater, and last update date. All of these will be vectorized. Examples of content information include version, summary, keywords, category, theme, project status, intended use, content type, content format, and number of pages. Of these, the summary, keywords, category, theme, project status, intended use, content type, and content format are the ones that will be vectorized. Permission and access information includes access rights, confidentiality levels, approvers, and approval dates. Of these, access rights and confidentiality levels are the ones that will be vectorized. Tracking and management information includes status, storage location, associated content, expiration date (lifecycle), and content ID. Of these, status, expiration date (lifecycle), and content ID are subject to vectorization. The metadata generation agent 44 may also query the user for missing information and obtain it. For example, the metadata generation agent 44 outputs information to the display control unit 66 for querying missing information. The display control unit 66 obtains the information for querying missing information from the metadata generation agent 44 and displays the information for querying missing information on the UI 10. The metadata generation agent 44 obtains the information entered by the user in response to the information for querying missing information displayed on the UI 10. This allows the metadata generation agent 44 to complete the metadata. Return to Figure 1 and continue the explanation.
[0025] The user attribute DB81 stores the user attribute information registered by the account registration / update processing unit 20. The metadata DB82 stores the metadata that the metadata generation agent 44 has stored. The usage log DB83 stores the usage logs of user U. An example of a usage log is the search history, which includes information indicating the date and time of the search, the content ID (identification information), the search user's attributes, and content metadata information. If a natural language search was performed, the search history will also include information indicating the search conditions.
[0026] The vector space generation agent 84 retrieves user attribute information from the user attribute DB 81, metadata from the metadata DB 82, and usage logs from the usage log DB 83. Based on the retrieved metadata, the vector space generation agent 84 generates an N-dimensional (where N is an integer > 1) vector space as metadata that represents the characteristics of the content. For example, N may be the number of attribute items to be vectorized among the metadata included in the content. The N dimensions include the creator (affiliation, employee number, name, etc.) and a summary / abstract. An example of the creator is 10 dimensions, and an example of the summary / abstract is 768 dimensions (example of the number of dimensions of the embedded model) × 7 chunks (for 3000 characters), assuming 125 characters overlap in a 500-character chunk. The vector space generation agent 84 stores the generated N-dimensional vector space in the content vector space memory unit 86.
[0027] The content vector space storage unit 86 stores the N-dimensional vector space stored by the vector space generation agent 84. The embedding processing unit 85 retrieves content metadata from the vector space generation agent 84 and performs a process to embed it into the N-dimensional vector space stored in the content vector space storage unit 86.
[0028] After the embedding processing unit 85 has embedded the content into an N-dimensional vector space, the hierarchical structure generation agent 50, in the classification process, obtains metadata for the content to be classified. Here, the content to be classified may be all of the multiple content items embedded in the N-dimensional vector space, or it may be a portion of the multiple content items embedded in the N-dimensional vector space, for example, specified by the user. The content may be, for example, a document. Specifically, the hierarchical structure generation agent 50 obtains one or more pieces of metadata, such as title, summary, and keywords.
[0029] In the classification process, the hierarchical structure generation agent 50 creates a distance matrix for all vectorized metadata. Here, the distance matrix is a matrix that represents the distance between each element included in the dataset. Specifically, the hierarchical structure generation agent 50 derives multiple distance measures, such as cosine distance and bidirectional mean minimum cosine distance, based on a distance function. The hierarchical structure generation agent 50 derives the distance between content by weighting each of the derived distance measures as needed and combining them.
[0030] Here, we will explain the bidirectional average minimum cosine distance, dividing it into its basic concept and calculation method. [Basic concept] This evaluates the similarity between two sets of text (or documents). Each set is represented as a collection of vectors (embedding vectors). The main ideas are (1) through (3). (1) Calculate the distance between each vector in set A and the nearest vector in set B. (2) Calculate the distance between each vector in set B and the nearest vector in set A. (3) Derive the final similarity score by combining the minimum distances in both directions.
[0031] [Calculation method] Two text sets A = {a1, a2, ..., a n} and B={b1,b2,···,b m Let's explain what happens when} is present. 1. The average of the minimum distances from A to B is expressed by equation (1).
number
[0032] 2. The average of the minimum distances from B to A is expressed by equation (2).
number
[0033] In the classification process, the hierarchical structure generation agent 50 determines the optimal number of cluster divisions. Specifically, the hierarchical structure generation agent 50 performs clustering using hierarchical clustering methods such as Ward's method or group average method, and creates a tendrogram.
[0034] This explanation of Ward's algorithm will cover both its fundamental principles and its algorithm. [Basic principle] 1. Start with each data point as an independent cluster. 2. At each step, select and merge the two clusters that result in the smallest increase in variance when merged. 3. Repeat until all data points form a single cluster.
[0035] [algorithm] 1. Based on equation (5), cluster C k Calculate the sum of squared errors (ESS). [Number] In Equation (5), x i is a data point in cluster C k and μ k is the centroid of cluster C k . 2. Calculate the increase in ESS (ΔESS i , C j ) when two clusters (C ij ) are integrated based on Equation (6) [Number]
[0036] 3. Update the distance using the distance update formula When clusters K and L are integrated to form cluster M, the distance between another cluster J and M is represented by Equation (7). [Number] In Equation (7), nJ is the size (number of data points) of cluster J. The description of Ward's method ends here.
[0037] The group average method will be explained separately in terms of its basic principle and algorithm. [Basic Principle] 1. Start with each data point as an independent cluster 2. Define the distance between clusters as the "average distance of all data points" Specifically, the distance between cluster A (a set of points) and cluster B (a set of points) is calculated as the average of the distances between all points in cluster A and all points in cluster B.
[0038] [Algorithm] The distance D(C1, C2) between cluster C1 (size n1) and cluster C2 (size n2) is represented by Equation (8). [Number] Here, d(x, y) is the Euclidean distance between data point x and data point y.
[0039] As an example of a method for creating a hierarchical structure using the group average method, we will explain agglomerative clustering. 1. Start with each point as one cluster (number of clusters = N). 2. Calculate the average distance for all cluster pairs. 3. Combine the two clusters with the smallest average distance. 4. Update the distance matrix (recalculate the distance to the new cluster). 5. Repeat steps 2 through 4 until the number of clusters reaches 1.
[0040] The process is traced backward from the point where the number of clusters reaches 1, to the point where the number of clusters reaches a predetermined number CK (e.g., 5). The average silhouette score of the data points at each stage from cluster number 2 to CK is calculated, and the cluster partitioning is determined at the point with the highest average score. The data within the determined cluster is flattened, and the above process is repeated to determine the child clusters of this cluster. This process is repeated until the number of data points in a single cluster reaches a predetermined number (2 or more). Further details will be provided later. This concludes the explanation of the group average method.
[0041] In the classification process, the hierarchical structure generation agent 50 calculates a score using a clustering accuracy evaluation method such as the silhouette score for each cluster number, and identifies the cluster number CK based on the score calculation result. The following explanation continues with an example of a clustering accuracy evaluation method, specifically the case where the silhouette score is used. For example, the hierarchical structure generation agent 50 identifies the cluster number CK that has the highest silhouette score.
[0042] This explanation of silhouette scoring will cover both its basic concepts and its calculation method. [Basic concept] The silhouette score is an index used to evaluate the quality of clustering. It measures the appropriateness of clusters (cohesion within clusters and separation between clusters). The silhouette score is calculated for each data point. A silhouette score close to 1 indicates that the data points have been assigned to the appropriate clusters. A silhouette score close to 0 indicates that the data point is near the cluster boundary. A silhouette score close to -1 indicates a high probability that data points have been assigned to the wrong cluster.
[0043] [Calculation method] The silhouette score s(i) of a given data point i is given by equation (9). s(i)=(b(i)-a(i)) / max(a(i), b(i)) (9) Here, a(i) is the average distance (cohesion) between data point i and all other points in the same cluster, and b(i) is the average distance (separation) between data point i and all points in the nearest cluster (a cluster to which i does not belong).
[0044] This section provides specific examples of how to evaluate clustering using silhouette scores. As an example, it explains the cases where data points are in the correct cluster, in the wrong cluster, and in the intermediate position between two clusters. [If the data points are in the correct cluster] FIG. 4 is a diagram for explaining an example of the processing of the information processing apparatus according to the present embodiment. In FIG. 4, clusters A to C are shown. Cluster A includes five data points indicated by "x", cluster B includes five data points indicated by "x", and cluster C includes five data points indicated by "x". In this state, for the data point i included in cluster A, the above-described equation (9) holds. In the example shown in FIG. 4, s(i) is about 0.8. Since b(i) > a(i), S(i) is close to 1, indicating that the data point i is assigned to an appropriate cluster.
[0045] [When the data point is in the wrong cluster] FIG. 5 is a diagram for explaining an example of the processing of the information processing apparatus according to the present embodiment. In FIG. 5, clusters A to C are shown. Cluster A includes five data points indicated by "x", cluster B includes five data points indicated by "x", and cluster C includes five data points indicated by "x". In this state, for the data point i included in cluster B, the above-described equation (9) holds. In the example shown in FIG. 5, s(i) is about -0.8. Since b(i) < a(i), S(i) is close to -1, indicating that there is a high possibility that the data point i is assigned to the wrong cluster.
[0046] [When the data point is in the middle of two clusters] FIG. 6 is a diagram for explaining an example of the processing of the information processing apparatus according to the present embodiment. In FIG. 6, clusters A to C are shown. Cluster A includes five data points indicated by "x", cluster B includes five data points indicated by "x", and cluster C includes five data points indicated by "x". In this state, for the data point i located in the middle between cluster A and cluster B, the above-described equation (9) holds. In the example shown in Figure 6, s(i) is 0. Since b(i) ≈ a(i) (b(i) is approximately equal to a(i)), S(i) is close to 0 (zero), indicating that the data point is near the cluster boundary. This concludes our explanation of specific examples of how to evaluate clustering using silhouette scores.
[0047] Figure 7 shows an example of a tendrogram. A tendrogram is a method of hierarchically grouping data and visually representing the results in a tree-like structure. Figure 7 shows an example where the silhouette score is 0.375. Let's return to Figure 1 and continue the explanation.
[0048] In the classification process, the hierarchical structure generation agent 50 flattens subclusters with a cluster count of CK. Specifically, the hierarchical structure generation agent 50 selects, for example, clusters with a cluster count of CK within the clustering process and flattens the selected clusters. Because the hierarchical clustering process is order-dependent, the clusters are flattened temporarily in order to generate semantically appropriate clusters from among the generated subclusters. In the classification process, the hierarchical structure generation agent 50 recursively performs the process of determining the optimal number of cluster divisions for one or more subclusters, as described above.
[0049] In the classification process, the hierarchical structure generation agent 50 recursively generates child clusters using silhouette scores. By configuring in this way, the hierarchical structure is determined, and the content to be included in the clusters under it is determined. The hierarchical structure generation agent 50 recommends the optimal folder (lowest level). The hierarchical structure generation agent 50 displays a classification display UI 16 on UI 10 to recommend the optimal folder. For example, the hierarchical structure generation agent 50 outputs classification information to the display control unit 66 to recommend the optimal folder. The display control unit 66 obtains the classification information from the hierarchical structure generation agent 50 and displays the classification information on UI 10. The hierarchical structure generation agent 50 obtains information entered by the user in relation to the classification information.
[0050] The natural language search agent 64 obtains user identification information and search prompts from the natural language search UI 18 and outputs them to the pseudo-metadata search agent 62. The pseudo-metadata search agent 62 obtains user identification information and a search prompt from the natural language search agent 64. Based on the obtained user identification information and search prompt, the pseudo-metadata search agent 62 generates pseudo-metadata for searching the content stored in the content vector space storage unit 86.
[0051] The pseudo-metadata search agent 62 places the pseudo-metadata generated by LLM6 into an N-dimensional vector space in which multiple contents stored in the content vector space storage unit 86 are embedded, based on the pseudo-metadata it has generated. The pseudo-metadata search agent 62 lists a predetermined number of contents, starting with those closest to the placed pseudo-metadata, and derives their likelihood.
[0052] The natural language search agent 64 obtains a predetermined number of content items and information indicating the likelihood of each of those content items from the pseudo-metadata search agent 62. The natural language search agent 64 outputs the obtained predetermined number of content items and information indicating the likelihood of each of those content items to the display control unit 66. The display control unit 66 obtains a predetermined number of content items and information indicating the likelihood of each of those content items from the natural language search agent 64, and displays them on the natural language search UI 18.
[0053] In the information processing device 100, the vector space generation agent 84 may also derive the distance between contents based on the weight vector. For example, the vector space generation agent 84 generates the weights c of all attributes. i The distribution can be made equal. In this case, the vector space generation agent 84 analyzes the distribution state for each attribute and uses the following method to generate weights c i This determines the weight c. This makes it possible to define the distance between content by averaging the variability between each attribute. i It is assumed that it follows a normal distribution (mean μ, standard deviation σ). In this case, the mean of the variance for all attributes is given by equation (10), and the weight c for attribute i is given by i σ is expressed by equation (11). i This is the standard deviation of attribute i.
[0054]
number
number
[0055] In the information processing device 100, the vector space generation agent 84 may optimize the classification according to the creator's circumstances. For example, the vector space generation agent 84 generates the creator's attributes as M-dimensional (where M is an integer M>1) vector data. Figure 8 shows an example of creator attribute information. As shown in Figure 8, an example of creator attribute information includes an employee number as an example of creator identification information, a name, department data, information indicating job title, and a job type code. Examples of affiliation data include the name of the business unit, the name of the business division, the name of the department, the name of the section, the business content, the department code, the job title, and the job type code. For example, creator attribute information may be represented by employee number, job content, department code, information indicating the job title, and job type code.
[0056] The vector space generation agent 84 further adds the created M-dimensional vector data to generate an (M+N)-dimensional vector space as metadata representing the characteristics of the content. The vector space generation agent 84 stores the generated (M+N)-dimensional vector space in the content vector space storage unit 86. Returning to Figure 1, the explanation continues.
[0057] The embedding processing unit 85 embeds the content into the (N+M)-dimensional vector space stored in the content vector space storage unit 86. Here, the vector space generation agent 84 may create attributes of the creator, such as relationships between organizations (departments) and hierarchical relationships within departments. For example, the vector space generation agent 84 may use a graph neural network (GNN) to create relationships between organizations (departments) and hierarchical relationships within departments. By specifying the creator's attributes and cutting a cross-section of an N-dimensional space along the specified creator's attributes, it is possible to generate subclassifications of the classification. This method allows for the display of the most suitable classification for an organization or individual.
[0058] In the information processing device 100, the vector space generation agent 84 may acquire management rules. The vector space generation agent 84 quantifies the acquired management rules and generates an R-dimensional vector space (where R is an integer R > 1). For example, management rules may include confidentiality levels, access rights, compliance requirements, lifecycle settings, etc. The management rules may be set by user U. The vector space generation agent 84 further adds the created R-dimensional vector data to generate an (R+M+N)-dimensional vector space as metadata representing the characteristics of the content. The vector space generation agent 84 stores the generated (R+M+N)-dimensional vector space in the content vector space storage unit 86.
[0059] The embedding processing unit 85 embeds the content into the (R+N+M) dimension vector space stored in the content vector space storage unit 86. The embedding processing unit 85 acquires an image with content embedded in a (R+N+M)-dimensional vector space and outputs it to the display control unit 66. The display control unit 66 acquires the image with content embedded in a (R+N+M)-dimensional vector space from the embedding processing unit 85 and displays it on the UI 10. This makes it possible to visualize the distribution state in (R+N+M) dimensions, i.e., within a (R,N,M) vector space.
[0060] In the information processing device 100, the vector space generation agent 84 may optimize classification according to the time attribute. For example, the vector space generation agent 84 creates the time attribute as T-dimensional (T is an integer T>1) vector data. An example of a time attribute is a two-dimensional (attribute, time) attribute that specifies the creation date or expiration date and the time indicating the corresponding date and time. The vector space generation agent 84 further adds the created T-dimensional vector data to generate a (T+R+M+N)-dimensional vector space as metadata representing the characteristics of the content. The vector space generation agent 84 stores the generated (T+R+M+N)-dimensional vector space in the content vector space storage unit 86. The embedding processing unit 85 embeds the content into a (T+R+N+M) dimension vector space stored in the content vector space storage unit 86.
[0061] In the information processing device 100, the vector space generation agent 84 may create user attributes as U-dimensional vector data (where U is an integer U > 1). Since Figure 4 can be applied to user attributes, a detailed explanation is omitted. The vector space generation agent 84 further adds the created U-dimensional vector data to generate a (U + T + R + M + N)-dimensional vector space as metadata representing the characteristics of the content. The vector space generation agent 84 stores the generated (U + T + R + M + N)-dimensional vector space in the content vector space storage unit 86.
[0062] The embedding processing unit 85 embeds the content that user U has searched so far into a (U+T+R+N+M) dimension vector space stored in the content vector space storage unit 86. The vector space generation agent 84 calculates the distance between the content and specific content included in the search history. The pseudo-metadata search agent 62 retrieves search results based on the distance calculated by the vector space generation agent 84. For example, the pseudo-metadata search agent 62 may prioritize retrieving search results starting with those with the shortest distance.
[0063] Here, the vector space generation agent 84 may calculate the distance not only to specific content included in the search history, but also to all content included in the search history or to content included in the same classification. The pseudo-metadata search agent 62 may have a first trained model that derives content features based on one or more content combinations included in the content search results. The pseudo-metadata search agent 62 creates information for presenting the derived content features and outputs it to the display control unit 66. The display control unit 66 displays the information for presenting the content features from the pseudo-metadata search agent 62 on the UI 10.
[0064] The first trained model is created by machine learning the relationship between one or more content combinations included in the search results and the content features derived from those combinations. For example, the first trained model is created by machine learning with one or more content combinations included in the search results as explanatory variables and the content features derived from those combinations as the dependent variable.
[0065] Furthermore, the pseudo-metadata search agent 62 may have a second trained model that derives questions for narrowing down the content based on the characteristics of the derived content. The pseudo-metadata search agent 62 creates questions for narrowing down the derived content and outputs them to the display control unit 66. The display control unit 66 displays the questions for narrowing down the content from the pseudo-metadata search agent 62 on the UI 10. The second pre-trained model is created by machine learning the relationship between content features and the questions used to narrow down content based on those features. For example, the second pre-trained model is created by machine learning with content features as explanatory variables and the questions used to narrow down content based on those features as the dependent variable. This allows for the creation of information that presents the characteristics of the content to the user based on the search results, and enables the creation of questions to narrow down the most suitable content. Based on the answers, more relevant content can be presented.
[0066] In the information processing device 100, the embedding processing unit 85 may embed the content into a (T+R+N+M) dimension vector space stored in the content vector space storage unit 86, and present nearby content as similar documents to the user. Furthermore, the vector space generation agent 84 may present the reciprocal of the distance between contents as the similarity score to the user. In this case, if there is content with a distance of zero between contents, the vector space generation agent 84 creates information to notify the user U that there is a high possibility of it being a duplicate document and outputs it to the display control unit 66. The display control unit 66 may display the information from the vector space generation agent 84 indicating a high possibility of it being a duplicate document on the UI 10.
[0067] The vector space generation agent 84 may generate a (R+N)-dimensional vector space, a (T+N)-dimensional vector space, a (U+N)-dimensional vector space, a (T+M+N)-dimensional vector space, or a (U+M+N)-dimensional vector space as metadata representing the features of the content. In other words, the vector space generation agent 84 can generate a vector space of any combination of T, R, M, and U plus N.
[0068] In the embodiment described above, the functions of the information processing device 100 may be implemented in a distributed manner by multiple devices. In this case, the multiple devices that implement the functions of the information processing device 100 in a distributed manner may be configured to be directly connected to perform information input and output, or they may be configured to be connected via a communication network to send and receive information. Furthermore, among the multiple devices that implement the functions of the information processing device 100 in a distributed manner, there may be a mix of directly connected devices and devices connected via a network. All or part of the account registration / update processing unit 20, reception unit 25, content pool crawler 42, metadata generation agent 44, hierarchical structure generation agent 50, pseudo-metadata search agent 62, natural language search agent 64, display control unit 66, vector space generation agent 84, and embedded processing unit 85 are functional units (hereinafter referred to as software functional units) realized by a processor such as a CPU (Central Processing Unit) executing a program stored in a memory unit (not shown).
[0069] Furthermore, all or part of these functional units may be implemented by hardware such as LSI (Large Scale Integration), ASIC (Application Specific Integrated Circuit), or FPGA (Field-Programmable Gate Array), or by a combination of software functional units and hardware.
[0070] (Processing flow of the information processing device 100) The operation of the information processing device 100 will be described below with reference to Figures 9 to 18. Figure 9 is a flowchart showing an example of the operation flow of the information processing device 100 in this embodiment. The process by which the information processing device 100 stores content will be described with reference to Figures 1 and 9. (Step S1-1) The information processing device 100 performs the login process. For example, user U enters user identification information and a password on the initial screen (not shown) displayed on the UI 10. The information processing device 100 performs the login process based on the entered user identification information and password.
[0071] (Step S2-1) The information processing device 100 creates content. For example, user U creates content using the information processing device 100. Content may also be input to the information processing device 100 from an external device. For example, content may be transmitted to the information processing device 100 from a terminal device (not shown) connected to the information processing device 100 via a communication network. In this case, the information processing device 100 receives and acquires the content transmitted by the terminal device. (Step S3-1) The content pool UI14 detects when the "Save As" button (not shown) is pressed.
[0072] (Step S4-1) The reception unit 25 receives content and saves the received content to the content pool 30. (Step S5-1) The content pool 30 notifies the content pool crawler 42 of the physical addresses assigned to the stored content. (Step S6-1) The content pool crawler 42 accepts the physical address notified by the content pool 30.
[0073] (Step S7-1) The metadata generation agent 44 obtains a physical address from the content pool crawler 42. Based on the obtained physical address, the metadata generation agent 44 identifies the content corresponding to the physical address from the content stored in the content pool 30. The metadata generation agent 44 generates metadata from the identified content. The metadata generation agent 44 stores the generated metadata in the metadata DB 82, associating it with content identification information.
[0074] (Step S8-1) The hierarchical structure generation agent 50 performs classification processing on the content stored in the content vector space storage unit 86, specifically on the content to be classified from among the content included in the clusters generated by the vector space generation agent 84. (Step S9-1) The hierarchical structure generation agent 50 analyzes metadata for each piece of content included in the clusters generated by the classification process, using LLM5. Based on the metadata analysis results, the hierarchical structure generation agent 50 recommends the optimal folder (lowest level). The hierarchical structure generation agent 50 displays the classification display UI 16 for recommending the optimal folder on UI 10.
[0075] (Step S10-1) The hierarchical structure generation agent 50 determines whether or not user U has performed an action to respond to a recommendation on the classification display UI 16. (Step S11-1) If the hierarchical structure generation agent 50 determines that user U has not taken any action to respond to the recommendation on the classification display UI 16, it displays the classification display UI 16 on UI 10 to recommend the next most suitable folder. Proceed to step S10-1. (Step S12-1) When the hierarchical structure generation agent 50 determines that user U has responded to a recommendation on the classification display UI 16, it opens the recommended folder on the classification display UI 16 and displays the file name of the corresponding content.
[0076] In the operation flow of the information processing device 100 shown in Figure 9, the following processing may be performed. For example, in step S6-1, the content pool crawler 42 may accept physical addresses notified by any means, not limited to physical addresses notified by the content pool 30. Furthermore, in step S7-1, if the metadata generation agent 44 finds any missing information in the metadata generated from the identified content, it may ask user U questions to fill in the gaps. This allows the metadata generation agent 44 to complete the metadata.
[0077] Figure 10 is a flowchart showing an example of the operation flow of the information processing device 100 in this embodiment. The process by which the information processing device 100 generates a hierarchical structure of classification will be explained with reference to Figures 1 and 10. (Step S1-2) The hierarchical structure generation agent 50 collects documents to be classified. (Step S2-2) The hierarchical structure generation agent 50 automatically generates metadata from each of the multiple documents to be classified that it has collected.
[0078] (Step S3-2) The hierarchical structure generation agent 50 vectorizes all of the automatically generated metadata. (Step S4-2) The hierarchical structure generation agent 50 generates a distance matrix for all vectorized metadata.
[0079] (Step S5-2) The hierarchical structure generation agent 50 performs a process to determine the optimal number of cluster partitions CK. (Step S6-2) The hierarchical structure generation agent 50 flattens the clusters of CKs.
[0080] (Step S7-2) The hierarchical structure generation agent 50 determines whether the number of documents included in the cluster meets a pre-specified condition. If the number of documents included in the cluster meets the pre-specified condition, the process terminates; otherwise, it proceeds to step S5-2.
[0081] Figure 11 is a diagram illustrating an example of the processing of the information processing device 100 in this embodiment. Referring to Figure 11, the process by which the hierarchical structure generation agent 50 generates a distance matrix for all vectorized metadata will be explained. The set of documents to be vectorized includes Document A and Document B.
[0082] (Steps S1-3) The hierarchical structure generation agent 50 automatically generates metadata from both Document A and Document B. Specifically, the hierarchical structure generation agent 50 automatically generates metadata from both Document A and Document B, such as title, summary / abstract, and keywords.
[0083] (Step S2-3) The hierarchical structure generation agent 50 vectorizes all the metadata it generates. Specifically, the hierarchical structure generation agent 50 creates title vector A and title vector B by vectorizing the titles automatically generated from document A and document B, respectively. The hierarchical structure generation agent 50 creates summary vector A and summary vector B by vectorizing the summaries automatically generated from document A and document B, respectively. The hierarchical structure generation agent 50 creates keyword vector A and keyword vector B by vectorizing the keywords automatically generated from document A and document B, respectively.
[0084] (Step S3-3) The hierarchical structure generation agent 50 derives a distance measure for each piece of metadata based on all the vectorized metadata. Specifically, the hierarchical structure generation agent 50 derives the cosine distance based on title vector A and title vector B. The hierarchical structure generation agent 50 derives the cosine distance based on summary vector A and summary vector B. The hierarchical structure generation agent 50 derives the bidirectional mean minimum cosine distance based on keyword vector A and keyword vector B.
[0085] (Step S4-3) The hierarchical structure generation agent 50 derives the distance between content by weighting and combining each of the derived distance measures as needed. Specifically, the hierarchical structure generation agent 50 derives the distance between content by weighting and integrating the cosine distance derived based on title vector A and title vector B, the cosine distance derived based on summary vector A and summary vector B, and the bidirectional average minimum cosine distance derived based on keyword vector A and keyword vector B. (Step S5-3) The hierarchical structure generation agent 50 creates a distance matrix by converting the distances between content items into a distance matrix.
[0086] Figure 12 is a flowchart showing an example of the operation flow of the information processing device 100 in this embodiment. Referring to Figures 1 and 12, the process by which the hierarchical structure generation agent 50 determines the optimal cluster division number CK will be explained. (Steps S1-4) The hierarchical structure generation agent 50 selects one document from the group of documents to be classified that is not joined to a cluster, and designates it as document A. (Step S2-4) The hierarchical structure generation agent 50 combines document A with the document or cluster that results in the smallest increase in variance when combined, and generates a new cluster.
[0087] (Step S3-4) The hierarchical structure generation agent 50 determines whether there are any documents that are not cluster-joined. If there are any documents that are not cluster-joined, the process proceeds to step S1-4. (Step S4-4) If there are no documents that are not cluster-bound, the hierarchical structure generation agent 50 selects an arbitrary cluster and merges it with the cluster that results in the smallest increase in variance when joined, thereby generating a new cluster.
[0088] (Step S5-4) The hierarchical structure generation agent 50 determines whether the total number of clusters is 1. If the total number of clusters is not 1, the process proceeds to step S4-4. (Step S6-4) The hierarchical structure generation agent 50 creates a tendrogram if the total number of clusters is 1. (Step S7-4) The hierarchical structure generation agent 50 calculates the average silhouette score of all documents at each partitioning stage and identifies the cluster number CK that maximizes the average silhouette score.
[0089] Figure 13 is a flowchart showing an example of the operation flow of the information processing device 100 in this embodiment. The process by which the information processing device 100 performs natural language search will be explained with reference to Figures 1 and 13. Here, as an example, we will explain the case where user U is logged in and the vector space generation agent 84 is generating a (T+R+M+N) dimension vector space.
[0090] (Steps S1-6) The natural language search UI 18 receives a search prompt entered by user U. The natural language search agent 64 obtains the search prompt received by the natural language search UI 18 and the logged-in user's identification information. Based on the obtained user identification information, the natural language search agent 64 obtains user attribute information from the user attribute DB 81. The natural language search agent 64 creates a prompt creation request that includes the user identification information and user attribute information and outputs it to the pseudo-metadata search agent 62.
[0091] (Step S2-6) The pseudo-metadata search agent 62 obtains a prompt creation request from the natural language search agent 64. The pseudo-metadata search agent 62 obtains user identification information, user attribute information, and the content of the prompt entered by the user, which are included in the obtained prompt creation request, and creates a prompt using LLM6 based on the obtained user attribute information and the content of the prompt entered by the user. LLM6 includes a model that has been trained for pseudo-metadata generation.
[0092] (Steps S3-6) The pseudo-metadata search agent 62 uses LLM6 to check whether the user attribute information necessary for searching and the information necessary to generate pseudo-metadata required for user searches are available. LLM6 includes pre-trained functions that check the availability of the user attribute information necessary for searching and the information necessary to generate pseudo-metadata required for user searches, and list the necessary user attribute information and the information necessary to generate pseudo-metadata required for user searches.
[0093] For example, the pseudo-metadata search agent 62 outputs information to the display control unit 66 for querying missing user attribute information and information necessary for generating pseudo-metadata required for user searches. The display control unit 66 obtains the information from the pseudo-metadata search agent 62 for querying missing user attribute information and information necessary for generating pseudo-metadata required for user searches, and displays the information on the UI 10 for querying missing user attribute information and information necessary for generating pseudo-metadata required for user searches. The pseudo-metadata search agent 62 obtains the information entered by user U in response to the information displayed on the UI 10 for querying missing user attribute information and information necessary for generating pseudo-metadata required for user searches.
[0094] (Steps S4-6) The pseudo-metadata search agent 62 determines whether the user attribute information necessary for searching and the information necessary for generating pseudo-metadata required for performing searches from the user are available.
[0095] (Steps S5-6) The pseudo-metadata search agent 62 requests additional information from user U if it determines that the user attribute information necessary for the search and the information necessary to generate the pseudo-metadata required for the search from the user are not available. For example, the pseudo-metadata search agent 62 may request additional information via chat. The process then proceeds to step S1-6.
[0096] (Step S6-6) The pseudo-metadata search agent 62, when it determines that it has the necessary user attribute information for searching and the information necessary to generate pseudo-metadata for performing searches from users, stores the necessary user attribute information for searching and the information necessary to generate pseudo-metadata for performing searches from users, and generates pseudo-metadata based on the stored user attribute information for searching and the information necessary to generate pseudo-metadata for performing searches from users.
[0097] (Step S7-6) The pseudo-metadata search agent 62 places the generated pseudo-metadata into a (T+R+M+N)-dimensional vector space in which multiple contents stored in the content vector space storage unit 86 are embedded, based on the pseudo-metadata it generates.
[0098] (Step S8-6) The pseudo-metadata search agent 62 lists a predetermined number of content items, starting with those closest to the placed pseudo-metadata, and derives their likelihood. When P is the coordinate of a point in the content vector space generated by the pseudo-metadata, the likelihood Li of content X can be expressed, for example, by equation (13). Li(x) = α(1 / D(P,X)) (13) In equation (13), α is an arbitrary constant used to adjust the value. According to equation (13), the closer (smaller) the distance D, the larger the likelihood Li becomes.
[0099] The natural language search agent 64 obtains a predetermined number of content items and information indicating the likelihood of each of those content items from the pseudo-metadata search agent 62. The natural language search agent 64 outputs the obtained predetermined number of content items and information indicating the likelihood of each of those content items to the display control unit 66. The display control unit 66 obtains the predetermined number of content items and information indicating the likelihood of each of those content items from the natural language search agent 64 and displays them on the UI 10.
[0100] Figure 14 is a flowchart showing an example of the operation flow of the information processing device 100 in this embodiment. Referring to Figures 1 and 14, other processes performed by the information processing device 100 for natural language search will be described. Here, as an example, we will describe the case where user U is logged in and the vector space generation agent 84 is generating a (T+R+M+N) dimension vector space.
[0101] (Steps S1-7) The natural language search UI 18 receives a search prompt entered by user U. The natural language search agent 64 obtains the search prompt received by the natural language search UI 18 and the logged-in user's identification information. Based on the obtained user identification information, the natural language search agent 64 obtains user attribute information from the user attribute DB 81. The natural language search agent 64 creates a prompt creation request that includes the user identification information and user attribute information and outputs it to the pseudo-metadata search agent 62.
[0102] (Step S2-7) The pseudo-metadata search agent 62 obtains a prompt creation request from the natural language search agent 64. The pseudo-metadata search agent 62 obtains user identification information and user attribute information included in the obtained prompt creation request, and creates a prompt using LLM6 based on the obtained user attribute information. LLM6 includes a model that has been trained for pseudo-metadata generation.
[0103] (Steps S3-7) The pseudo-metadata search agent 62 uses LLM6 to check whether the user attribute information necessary for searching and the information necessary to generate pseudo-metadata required for user searches are available. LLM6 includes pre-trained functions that check the availability of the user attribute information necessary for searching and the information necessary to generate pseudo-metadata required for user searches, and list the necessary user attribute information and the information necessary to generate pseudo-metadata required for user searches.
[0104] For example, the pseudo-metadata search agent 62 outputs information to the display control unit 66 for querying missing user attribute information and information necessary for generating pseudo-metadata required for user searches. The display control unit 66 obtains the information from the pseudo-metadata search agent 62 for querying missing user attribute information and information necessary for generating pseudo-metadata required for user searches, and displays the information on the UI 10 for querying missing user attribute information and information necessary for generating pseudo-metadata required for user searches. The pseudo-metadata search agent 62 obtains the information entered by user U in response to the information displayed on the UI 10 for querying missing user attribute information and information necessary for generating pseudo-metadata required for user searches.
[0105] (Steps S4-7) The pseudo-metadata search agent 62 excludes dimensions relating to attributes that have not yet been generated from the search space if attribute generation is not complete.
[0106] (Steps S5-7) The pseudo-metadata search agent 62, after excluding dimensions relating to attributes that have not been generated in steps S4-7 from the search space, or after the attribute generation is completed in step S3-7, stores information necessary for generating user attribute information and pseudo-metadata necessary for searching from users, and generates pseudo-metadata based on the stored user attribute information and information necessary for generating pseudo-metadata necessary for searching from users.
[0107] (Steps S6-7) The pseudo-metadata search agent 62 places the generated pseudo-metadata into a (T+R+M+N)-dimensional vector space in which multiple contents stored in the content vector space storage unit 86 are embedded, based on the pseudo-metadata it has generated. The pseudo-metadata search agent 62 generates a subspace containing only the dimensions in which attribute data exists, and searches this space.
[0108] (Step S7-7) The pseudo-metadata search agent 62 lists a predetermined number of content items, starting with those closest to the placed pseudo-metadata, and derives their likelihood. The natural language search agent 64 obtains a predetermined number of content items and information indicating the likelihood of each of those content items from the pseudo-metadata search agent 62. The natural language search agent 64 outputs the obtained predetermined number of content items and information indicating the likelihood of each of those content items to the display control unit 66. The display control unit 66 obtains a predetermined number of content items and information indicating the likelihood of each of those content items from the natural language search agent 64 and displays them on the UI 10.
[0109] Figure 15 is a flowchart showing an example of the operation flow of the information processing device 100 in this embodiment. Referring to Figures 1 and 15, the process by which the information processing device 100 displays the optimal classification for user U will be described. Here, as an example, the case in which the vector space generation agent 84 generates a (T+R+M+N) dimension vector space will be described.
[0110] (Steps S1-8) The information processing device 100 performs the login process. For example, user U enters user identification information and a password on the initial screen (not shown) displayed on the UI 10. The information processing device 100 performs the login process based on the entered user identification information and password.
[0111] (Step S2-8) The display control unit 66 displays the classification creation screen on the UI 10. The classification creation screen includes a display for selecting whether to target content created by the user (user U), content created by others other than user U, or all content.
[0112] First, let's explain what happens when user U chooses to target content they have created. (Step S3-8-1) The vector space generation agent 84 creates a subspace in a (T+R+M+N) dimension vector space in which multiple contents stored in the content vector space memory unit 86 are embedded, by fixing the creator attribute, which is M-dimensional vector data, to itself. (Step S4-8-1) The hierarchical structure generation agent 50 creates a classification on this subspace.
[0113] Next, we will explain what happens when user U chooses to target content created by others. (Step S3-8-2) The vector space generation agent 84 creates a subspace in a (T+R+M+N) dimension vector space in which multiple contents stored in the content vector space memory unit 86 are embedded, by fixing the creator attribute, which is an M-dimensional vector data, to something other than itself (someone other than itself).
[0114] (Step S4-8-2) The hierarchical structure generation agent 50 creates a classification on this subspace. (Step S5-8-2) The hierarchical structure generation agent 50 leaves the lowest-level folder containing previously viewed content as is, and creates a category called "Other" for all other folders, moving them under this category, and then completes the process. By classifying content with a viewing frequency below a pre-set threshold into "Other," the hierarchical structure generation agent 50 can improve visibility.
[0115] Next, we will explain what happens when user U chooses to include all content. (Step S3-8-3) The vector space generation agent 84 creates and completes classification in a (T+R+M+N) dimension vector space in which multiple contents stored in the content vector space memory unit 86 are embedded.
[0116] In the information processing device 100, the vector space generation agent 84 may use LLM to extract creator attributes that are similar to the user attributes of the user themselves. The vector space generation agent 84 may also create a classification based on a subset of the content vector space based on the extracted creator attributes.
[0117] Figure 16 is a flowchart showing another example of the operation flow of the information processing device 100 in this embodiment. Referring to Figures 1 and 16, the process by which the information processing device 100 displays the optimal classification for user U will be described. Here, as an example, the case in which the vector space generation agent 84 generates a (T+R+M+N) dimension vector space will be described.
[0118] (Steps S1-9) The information processing device 100 performs the login process. For example, user U enters user identification information and a password on the initial screen (not shown) displayed on the UI 10. The information processing device 100 performs the login process based on the entered user identification information and password.
[0119] (Step S2-9) The natural language search UI18 accepts the conditions for creating classifications that are entered as prompts on the Chat input screen of LLM (not illustrated). The natural language search agent 64 obtains user identification information and conditions for classification creation entered as prompts from the natural language search UI 18 and outputs them to the pseudo-metadata search agent 62. The pseudo-metadata search agent 62 obtains user identification information and classification creation conditions entered as prompts from the natural language search agent 64. The pseudo-metadata search agent 62 outputs the obtained user identification information and prompts to the LLM (not shown). The LLM (not shown) obtains the user identification information and prompts from the pseudo-metadata search agent 62, refers to the (T+R+M+N) dimension vector space in which multiple contents stored in the content vector space storage unit 86 are embedded, and creates a content subspace that matches the conditions of user U.
[0120] (Step S3-9) The vector space generation agent 84 creates and completes a classification in the content subspace that matches the conditions for user U created by LLM (not shown). Figure 17 shows an example of a content classification screen. As shown in Figure 17, the example content classification screen includes a section for displaying major categories and a section for displaying subcategories. Furthermore, the example content classification screen includes an update button (UPBU).
[0121] The section displaying major categories shows one or more folders. In the example shown in Figure 17, the folders "AAA", "BBB", "CCC", ..., "XXX", ..., and "KKK" are displayed. The section displaying subcategories shows one or more folders that are included in the folder specified by user U, from among the one or more folders displayed in the section displaying major categories. In the example shown in Figure 17, the folders "PPP", "QQQ", "RRR", ..., "SSS", and "TTT" that are included in the "XXX" folder are displayed. Furthermore, pressing the update button UPBU allows the classification structure to be updated.
[0122] According to the information processing device 100 of this embodiment, metadata is generated based on each of multiple contents, a vector space representing the characteristics of the contents is generated based on the generated metadata, and multiple clusters are generated from the multiple contents based on the generated vector space and the semantic similarity of the contents, thereby virtually classifying the contents. Since multiple clusters are generated from multiple contents based on the semantic similarity of the contents, the time required to find the contents can be reduced compared to the case where multiple clusters are generated from multiple contents without being based on the semantic similarity of the contents.
[0123] (Modified examples of the embodiment) Information processing device of a modified embodiment Figure 18 shows an example of an information processing device 150 as a modified embodiment. The information processing device 150 is implemented by a device such as a personal computer, server, smartphone, tablet computer, or industrial computer.
[0124] The information processing device 150 differs from the information processing device 100 in that it includes a hierarchical structure generation agent 55 instead of the hierarchical structure generation agent 50. The hierarchical structure generation agent 55 can be applied to the hierarchical structure generation agent 50. However, the hierarchical structure generation agent 55 differs from the hierarchical structure generation agent 50 in the following respects. The reception unit 25 receives information indicating a hierarchical structure. The hierarchical structure is a classification tree in which content is represented hierarchically in L layers (where L is an integer L > 1). For example, the user can input information indicating the direction of classification, such as by category, by rule, or by user attribute, from the operation unit. The operation unit is, for example, an operating device such as a touch panel, keyboard, or mouse, and is operated by the operator.
[0125] Figure 19 shows an example of a hierarchical structure. Figure 19 shows a two-tiered hierarchical structure when categorized as information indicating the direction of classification. The top-level folder has the "Admission" folder, "Advanced Studies" folder, and "Enrollment Management" folder in the first tier, and the "Admission-Related Documents" folder and "Transfer / Enrollment" folder in the second tier. The attributes of the files that will be contained in each of the multiple folders included in the hierarchical structure are predefined. For example, for each of the multiple folders, the attributes of the files to be contained may be selected, and information indicating the attributes of the selected files may be stored in association with the folder.
[0126] The hierarchical structure generation agent 55 acquires information indicating the hierarchical structure received by the reception unit 25. Based on the acquired information indicating the hierarchical structure, the hierarchical structure generation agent 55 performs classification processing on the content to be classified from among the multiple contents embedded in the N-dimensional vector space stored in the content vector space storage unit 86. Hereinafter, the information indicating the hierarchical structure may be referred to as the "template". Specifically, the hierarchical structure generation agent 55 classifies multiple contents based on information indicating a hierarchical structure and information indicating the attributes of files stored in association with each of the multiple folders. The hierarchical structure generation agent 55 may also classify multiple contents based on metadata generated from each of the multiple contents. For example, the hierarchical structure generation agent 55 stores contents (documents, papers) with metadata of "admission-related documents" as a category in the "admission-related documents" folder.
[0127] Furthermore, the hierarchical structure generation agent 55 may, for example, derive the similarity between each of the multiple contents and the name displayed in the folder in the hierarchical structure, and classify the multiple contents based on the derived similarity. For example, the hierarchical structure generation agent 55 stores content (e.g., documents) that has metadata of "Admission-Related Documents" as a category in the "Admission-Related Documents" folder. For example, the hierarchical structure generation agent 55 stores content (e.g., documents) that has metadata with a high similarity to the content (e.g., documents) template "Admission-Related Documents". The hierarchical structure generation agent 55 stores each of the classified content items in its corresponding folder.
[0128] The hierarchical structure generation agent 55 performs hierarchical clustering on each of the multiple folders at the L-level to create a hierarchical structure from the L-level onward. The hierarchical structure generation agent 55 may, for example, adjust the metadata weighting when performing hierarchical clustering. By adjusting the metadata weighting, the data can be classified to approximate the hierarchical structure specified by the user.
[0129] The hierarchical structure generation agent 55 updates the hierarchical structure in the background when the distribution of content included in the cluster may change, for example, when content is added. The hierarchical structure generation agent 55 generates folder names using LLM5 when generating and updating the hierarchical structure.
[0130] (Processing flow of the information processing device 150) The operation of the information processing device 150 will be explained below with reference to Figures 20 to 23. Figure 20 is a flowchart showing an example of the operation flow of the information processing device 150 in a modified embodiment. Referring to Figures 18 and 20, the processes executed by the information processing device 150 when either or both of the following are performed: adding or deleting content such as documents will be described. Here, the operations after either or both of the following are performed by the vector space generation agent 84 will be described.
[0131] (Steps S1-10) The hierarchical structure generation agent 55 starts crawling and detects differential files. (Step S2-10) The hierarchical structure generation agent 55 determines whether a document has been added or not. (Step S3-10) If the hierarchical structure generation agent 55 determines that a document has been added in step S2-10, it then determines whether or not the document has been deleted.
[0132] (Step S4-10) If the hierarchical structure generation agent 55 determines in step S3-10 that no documents have been deleted, it retrieves information about newly added documents. (Step S5-10) The hierarchical structure generation agent 55 determines whether an existing document exists. (Step S6-10) If the hierarchical structure generation agent 55 determines that an existing document exists, it merges (integrates) the existing distance matrix with the new information.
[0133] (Step S7-10) If the hierarchical structure generation agent 55 determines that there are no existing documents, it fully calculates the distance matrix using only the new information. (Steps S8-10) If the hierarchical structure generation agent 55 determines that a document has been deleted in step S3-10, it removes the deleted document from the existing distance matrix. (Steps S9-10) The hierarchical structure generation agent 55 retrieves information about newly added documents.
[0134] (Step S10-10) The hierarchical structure generation agent 55 calculates a distance matrix between the new document and the existing document. (Step S11-10) The hierarchical structure generation agent 55 calculates a distance matrix between new documents. (Step S12-10) The hierarchical structure generation agent 55 merges the calculation results with existing information to generate a new distance matrix.
[0135] (Step S13-10) If the hierarchical structure generation agent 55 determines in step S2-10 that no documents have been added, it then determines whether or not a document has been deleted. (Step S14-10) If the hierarchical structure generation agent 55 determines in step S13-10 that a document has been deleted, it removes the corresponding location (the location of the deleted document) from the existing distance matrix. (Step S15-10) The hierarchical structure generation agent 55 determines that there are no updates if it determines in step S13-10 that the document has not been deleted.
[0136] Figure 21 is a flowchart showing an example of the operation flow of an information processing device 150, a modified embodiment. Referring to Figures 18 and 21, we will describe a case in which content is clustered hierarchically based on information showing a hierarchical structure, and then a classification structure update process is performed in the background.
[0137] (Step S1-11) The vector space generation agent 84 checks for new documents. The vector space generation agent 84 checks for new documents at predetermined intervals, such as every two hours. (Step S2-11) If the vector space generation agent 84 can identify a new document, it converts the new document into a vector representation and saves it to the content vector space storage unit 86.
[0138] (Step S3-11) The vector space generation agent 84 saves the newly converted document into a vector representation to the content vector space storage unit 86, and then instructs the hierarchical structure generation agent 55 to calculate the overall distance matrix. (Step S4-11) The hierarchical structure generation agent 55 calculates the overall distance matrix based on instructions from the vector space generation agent 84 to calculate the overall distance matrix and stores it in the content vector space storage unit 86.
[0139] (Step S5-11) The user accesses the page. (Step S6-11) The reception unit 25 receives information indicating a template (information indicating a hierarchical structure) entered by the user by operating the control unit.
[0140] (Step S7-11) The hierarchical structure generation agent 55 obtains information indicating a template from the reception unit 25, and performs background processing to update the distance matrix based on the obtained template information. (Step S8-11) The hierarchical structure generation agent 55 notifies the display control unit 66 that the distance matrix is being updated, and notifies the user by displaying on the display unit (not shown) that the distance matrix is being updated.
[0141] (Steps S9-11) After the classification of the template is complete, the hierarchical structure generation agent 55 stores the classification results of the template in the content vector space storage unit 86. (Steps S10-11) After saving the classification results of the template, the hierarchical structure generation agent 55 notifies the display control unit 66 that the update is complete.
[0142] (Step S11-11) The display control unit 66 displays an update button on the display unit based on information indicating that the update from the hierarchical structure generation agent 55 has been completed. (Step S12-11) The user clicks the update button displayed on the display unit. The reception unit 25 receives the latest result request entered by the user when they click the update button.
[0143] (Step S13-11) The hierarchical structure generation agent 55 obtains the latest result request from the reception unit 25. (Step S14-11) The hierarchical structure generation agent 55 outputs information indicating the latest template classification result to the display control unit 66 based on the latest result request it has received. (Step S15-11) The display control unit 66 acquires information indicating the classification result of the latest template output by the hierarchical structure generation agent 55, and causes the display unit to display the classification result of the latest template based on the acquired information indicating the classification result of the latest template.
[0144] Steps S1-11 through S4-11 are executed to create a distance matrix for the entire document, which can then be stored as a cache. If the document is updated, this distance matrix will be recreated. By executing steps S5-11 to S15-11, the information processing device 150 can retain the latest template configuration information used by the user. When the user accesses it again, the information processing device 150 can recreate the hierarchical structure based on the template configuration information. Once the creation of the hierarchical structure is complete, the information processing device 150 displays the classification results. By decomposing and retaining the latest log data, the information processing device 150 can complete processing in a short period of time even for minor changes in the classification structure.
[0145] Figure 22 is a flowchart showing an example of the operation flow of the information processing device 150, a modified embodiment. The case in which the labeling process is automatically performed by LLM5 will be described with reference to Figures 18 and 22.
[0146] (Steps S1-12) The hierarchical structure generation agent 55 determines whether it is a semantic classification. (Step S2-12) The hierarchical structure generation agent 55 generates clusters when it determines that it is a semantic classification.
[0147] (Step S3-12) The hierarchical structure generation agent 55 adds the generated cluster to the labeling queue. (Step S4-12) The hierarchical structure generation agent 55 starts the labeling process. The labeling process may start synchronously when the cluster is added to the labeling queue, or it may start asynchronously.
[0148] (Step S5-12) The hierarchical structure generation agent 55 requests labeling processing from an interactive generation AI service such as LLM5. (Step S6-12) The hierarchical structure generation agent 55 determines whether it has successfully obtained a response to the labeling processing request from the conversational generation AI service.
[0149] (Step S7-12) If the hierarchical structure generation agent 55 fails to obtain a response to the labeling request from the interactive generation AI service, it uses the title of the child document. (Steps S8-12) The hierarchical structure generation agent 55 applies a name to the cluster if it successfully obtains a response to the labeling request from the interactive generation AI service, or if it fails to obtain a response to the labeling request from the interactive generation AI service and uses the title of the child document.
[0150] (Steps S9-12) The hierarchical structure generation agent 55 determines if there is a next task in the labeling queue. If there is a next task in the labeling queue, the process proceeds to step S4-12. (Steps S10-12) The hierarchical structure generation agent 55 waits for a new task.
[0151] (Steps S11-12) The hierarchical structure generation agent 55 determines whether a new task has arrived. If a new task has arrived, the process proceeds to step S4-12. If no new task has arrived, the process proceeds to step S10-12. (Step S12-12) If the hierarchical structure generation agent 55 determines in step S1-12 that it is not a semantic classification, it performs a filtering classification.
[0152] (Step S13-12) The hierarchical structure generation agent 55 determines whether to proceed to the next hierarchical level. If it decides to proceed to the next hierarchical level, it proceeds to step S1-12. (Step S14-12) The hierarchical structure generation agent 55 completes processing if it does not proceed to the next hierarchical level. (Step S15-12) The hierarchical structure generation agent 55 returns the classification result.
[0153] Figure 23 is a flowchart showing an example of the operation flow of the information processing device 150 in a modified embodiment. Referring to Figures 18 and 23, the case in which labeling processing is automatically performed by LLM5 will be described. Here, as an example, the processing of the classification engine, semantic classification function, and labeling queue of the hierarchical structure generation agent 55 will be described separately.
[0154] (Step S1-13) The vector space generation agent 84 instructs the hierarchical structure generation agent 55 to begin classification. (Step S2-13) The classification engine receives the classification start instructed by the vector space generation agent 84 and instructs the semantic classification function to classify hierarchical level 1.
[0155] (Step S3-13) The semantic classification function creates group A based on instructions from the classification engine. (Step S4-13) The semantic classification function requests the labeling queue to name group A.
[0156] (Step S5-13) The semantic classification function creates group B based on instructions from the classification engine. (Step S6-13) The semantic classification function requests the labeling queue to name group B.
[0157] (Step S7-13) The classification engine receives the classification start instructed by the vector space generation agent 84 and instructs the semantic classification function to classify hierarchical level 2. (Step S8-13) The semantic classification function creates subgroup C based on instructions from the classification engine.
[0158] (Steps S9-13) The semantic classification function requests the labeling queue to name subgroup C. (Steps S10-13) The labeling queue requests the LLM service (LLM6) to generate a name for group A.
[0159] (Steps S11-13) The LLM service generates a name for group A based on a request from the labeling queue to generate a name for group A. Here, we continue the explanation assuming that the LLM service has created "Business Document" as the name for group A. The LLM service responds to the labeling queue with "Business Document" as the name for group A. (Steps S12-13) The labeling queue retrieves responses from the LLM service, sets the name of Group A to "Business Documents" based on the retrieved responses, and notifies the classification engine that the name of Group A has been set.
[0160] (Step S13-13) The labeling queue requests the LLM service to generate a name for group B. (Step S14-13) The labeling queue requests the LLM service to generate the name for subgroup C.
[0161] (Step S15-13) The LLM service generates the name of Group B based on a request to generate the name of Group B from the labeling queue. Here, the description continues with the case where the LLM service creates "Technical Document" as the name of Group B. The LLM service responds to the labeling queue with "Technical Document" as the name of Group B. (Step S16-13) The labeling queue obtains the response from the LLM service, sets the name of Group B to "Technical Document" based on the obtained response, and notifies the classification engine that the name of Group B has been set.
[0162] (Step S17-13) The LLM service generates the name of Subgroup C based on a request to generate the name of Subgroup C from the labeling queue. Here, the description continues with the case where the LLM service creates "Manual" as the name of Subgroup C. The LLM service responds to the labeling queue with "Manual" as the name of Subgroup C. (Step S18-13) The labeling queue obtains the response from the LLM service, sets the name of Subgroup C to "Manual" based on the obtained response, and notifies the classification engine that the name of Subgroup C has been set.
[0163] (Step S19-13) The classification engine obtains the notification from the labeling queue that the name of Group A has been set, the name of Group B has been set, and the name of Subgroup C has been set. Based on the obtained notification, the classification engine returns information indicating a named hierarchical structure to the vector space generation agent 84.
[0164] According to a modification of the embodiment, the information processing apparatus 150 classifies a plurality of contents and generates a plurality of clusters based on information indicating a hierarchical structure that represents the contents in L layers hierarchically. If a new document is generated, automatic classification is performed. If an unintended hierarchical structure is generated as a result of the automatic classification, the user will give an instruction for automatic classification again, and the process will be repeated until the structure intended by the user is obtained, which requires time and labor. According to the information processing apparatus 150, since the user can indicate the direction of classification by the information indicating the hierarchical structure, the processing time can be shortened as compared with the case of creating a hierarchical structure without indicating the hierarchical structure. In addition, the information processing apparatus 150 performs the process of updating the classification structure of the contents in the background. The process of automatic document classification can be executed without making the user notice as much as possible. In addition, the information processing apparatus 150 can be automatically labeled by the LLM6.
[0165] <Supplementary Note> [Configuration Example 1] A reception unit that receives contents, [[ID=1^{3}]] Based on each of the information specifying a plurality of physical addresses from a content pool in which information specifying a physical address is given and stored for each of the plurality of contents received by the reception unit, a plurality of contents are acquired, metadata is generated based on each of the plurality of acquired contents, a vector space representing the characteristics of the contents is generated based on the plurality of generated metadata, and a plurality of clusters are generated from the plurality of contents based on the generated vector space and the semantic similarity of the contents, thereby virtually classifying the contents. A content processing unit, A display control unit that causes the display unit to display the result of the content processing unit virtually classifying the contents, An information processing apparatus comprising:
[0166] [Configuration Example 2] The content processing unit, A content pool crawler that receives information identifying the physical address assigned to the content stored in the aforementioned content pool, A metadata generation unit obtains information identifying a physical address from the content pool crawler, obtains content from the content pool based on the obtained information identifying a physical address, and generates metadata based on the obtained content. A vector space generation unit generates an N-dimensional (where N is an integer N>1) vector space as metadata that represents the characteristics of the content based on multiple metadata, An embedding processing unit that embeds content based on metadata into the N-dimensional vector space generated by the vector space generation unit, A hierarchical structure generation unit generates k (where k is an integer k>1) clusters from multiple contents embedded in the N-dimensional vector space, based on the semantic similarity of the contents. Equipped with, The N is the number of attribute items to be vectorized among the plurality of metadata. The information processing device described in Configuration Example 1.
[0167] [Configuration Example 3] The hierarchical structure generation unit acquires a plurality of metadata generated based on each of the plurality of contents, vectorizes each of the acquired plurality of metadata, derives the distance between the vectorized plurality of metadata for each metadata, and generates k clusters based on the derived plurality of distances. The information processing device described in Configuration Example 2.
[0168] [Configuration Example 4] The hierarchical structure generation unit creates a distance matrix based on the distances between the vectorized metadata, and generates k clusters based on the created distance matrices. The information processing device described in Configuration Example 3.
[0169] [Configuration Example 5] The hierarchical structure generation unit weights the distances between the vectorized metadata and generates a distance matrix. The information processing device described in Configuration Example 4.
[0170] [Configuration Example 6] The hierarchical structure generation unit optimizes each of the k clusters based on the average distance between any content and other content included in the cluster containing the content, and the average distance between any content and the cluster closest to it. The information processing device described in Configuration Example 2.
[0171] [Configuration Example 7] The aforementioned vector space generation unit creates M-dimensional (where M is an integer M>1) vector data representing the attributes of the content creator, and generates an M-dimensional vector space. The embedding processing unit embeds the content into an N+M dimensional vector space. The information processing device described in Configuration Example 2.
[0172] [Configuration Example 8] The aforementioned vector space generation unit creates R-dimensional (where R is an integer R > 1) vector data based on management rules and generates an R-dimensional vector space. The aforementioned embedding processing unit embeds the content into an N+M+R dimension vector space. The information processing device described in Configuration Example 7.
[0173] [Configuration Example 9] The vector space generation unit creates T-dimensional vector data (where T is an integer T>1) representing the time axis and attributes, and generates a T-dimensional vector space. The aforementioned embedding processing unit embeds the content into an N+M+R+T dimension vector space. The information processing device described in Configuration Example 8.
[0174] [Configuration Example 10] The aforementioned content processing unit, A pseudo metadata search unit that generates pseudo metadata based on a search prompt and searches for content based on the generated pseudo metadata. comprising The display control unit causes the display unit to display the content searched by the pseudo metadata search unit. The information processing apparatus according to Configuration Example 2.
[0175] [Configuration Example 11] When there is information lacking in the metadata generated based on the content, the metadata generation unit creates information for inquiring about the lacking information, acquires the information input for the created inquiring information, and completes the metadata The information processing apparatus according to Configuration Example 2.
[0176] [Configuration Example 12] The pseudo metadata search unit derives the characteristics of the content based on the search result of the content. The display control unit causes the display unit to display information indicating the characteristics of the content derived by the pseudo metadata search unit. The information processing apparatus according to Configuration Example 10.
[0177] [Configuration Example 13] The pseudo metadata search unit derives a question for narrowing down the content based on the derived characteristics of the content. The display control unit causes the display unit to display information indicating the question for narrowing down the content derived by the pseudo metadata search unit. The information processing apparatus according to Configuration Example 12.
[0178] [Configuration Example 14] When there is lacking information, the pseudo metadata search unit creates information for inquiring about the lacking information. The information processing apparatus according to Configuration Example 10.
[0179] [Configuration Example 15] The aforementioned receiving unit receives information indicating a hierarchical structure that represents the content in a hierarchical manner. The hierarchical structure generation unit generates clusters from multiple contents based on the information indicating the hierarchical structure. The information processing device described in Configuration Example 2.
[0180] [Configuration Example 16] The aforementioned hierarchical structure generation unit generates k clusters from multiple content items in the background. The information processing device described in Configuration Example 2.
[0181] [Configuration Example 17] The aforementioned hierarchical structure generation unit uses a large-scale language model to set the name of each of the k clusters. The information processing device described in Configuration Example 2.
[0182] It is also possible to provide a method for performing processing similar to that performed by an information processing device. [Configuration Example 18] A method of information processing performed by a computer, Accepts multiple content, From a content pool where information identifying the physical address is assigned to each of the received content items and stored, multiple content items are retrieved based on each of the pieces of information identifying the physical address. Metadata is generated based on each of the multiple acquired contents, and a vector space representing the characteristics of the contents is generated based on the multiple generated metadata. Based on the semantic similarity between the generated vector space and the content, multiple clusters are generated from multiple content pieces, thereby virtually classifying the content. The results of virtually classifying the content are displayed on the display unit. Information processing methods.
[0183] It is also possible to provide a program (computer program) that performs the same processing as that performed by the information processing device. [Configuration Example 19] On the computer, Accept multiple types of content, From a content pool in which information identifying the physical address is assigned to each of the multiple pieces of content that have been received and stored, multiple pieces of content are retrieved based on each of the pieces of information identifying the physical address. Based on each of the multiple retrieved content items, metadata is generated, and based on the generated metadata, a vector space representing the characteristics of the content is generated. Based on the semantic similarity between the generated vector space and the content, multiple clusters are generated from multiple content pieces, thereby virtually classifying the content. The results of virtually classifying the content are displayed on the display unit. program.
[0184] Although embodiments of the present invention have been described in detail above with reference to the drawings, the specific configuration is not limited to these embodiments, and design modifications and the like are also included within the scope of the gist of the present invention. Alternatively, a computer program for realizing the functions of the information processing device 100 or information processing device 150 described above may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be loaded into a computer system and executed. The term "computer system" here may include hardware such as an operating system and peripheral devices.
[0185] Furthermore, "computer-readable recording media" refers to writable non-volatile memory such as flexible disks, magneto-optical disks, ROMs, and flash memory, portable media such as DVDs (Digital Versatile Disks), and storage devices such as hard disks built into computer systems.
[0186] Furthermore, "computer-readable recording media" also includes volatile memory (such as DRAM (Dynamic Random Access Memory)) within computer systems that act as servers or clients when programs are transmitted via networks such as the Internet or communication lines such as telephone lines, which retain programs for a certain period of time.
[0187] Furthermore, the above program may be transmitted from a computer system that stores the program in a memory device or the like to another computer system via a transmission medium or by transmission waves within the transmission medium. Here, the "transmission medium" used to transmit the program refers to a medium that has the function of transmitting information, such as a network (communication network) like the Internet or a communication line (communication line) like a telephone line.
[0188] Furthermore, the above program may be intended to implement some of the functions described above. It may also be a so-called differential file (differential program) that can implement the aforementioned functions in combination with programs already recorded in the computer system. [Explanation of symbols]
[0189] 100, 150…Information processing device, 10…UI, 12…Account registration / update UI, 14…Content pool UI, 16…Classification display UI, 18…Natural text search UI, 20…Account registration / update processing unit, 30…Content pool, 40…Virtual content pool construction processing unit, 42…Content pool crawler, 44…Metadata generation agent, 50, 55…Hierarchical structure generation agent, 62…Pseudo-metadata search agent, 64…Natural text search agent, 66…Display control unit, 80…Virtual content pool, 81…User attribute DB, 82…Metadata DB, 83…Usage log DB, 84…Vector space generation agent, 85…Embedding processing unit, 86…Content vector space storage unit
Claims
1. The reception desk that accepts content, A content processing unit virtually classifies content by: obtaining multiple content items from a content pool in which information identifying a physical address is assigned to each of the multiple content items received by the reception unit and stored; obtaining multiple content items based on each of the information identifying a physical address; generating metadata based on each of the obtained multiple content items; generating a vector space representing the characteristics of the content based on the generated multiple metadata; and generating multiple clusters from the multiple content items based on the generated vector space and the semantic similarity of the content. The content processing unit causes the display control unit to display the results of the virtual classification of the content on the display unit, Equipped with, The content processing unit generates k clusters (where k is an integer k > 1) from a plurality of content, and optimizes each of the k clusters based on the average distance between any content and other content included in the cluster containing the arbitrary content, and the average distance between any content and the cluster closest to it.
2. The aforementioned content processing unit, A content pool crawler that receives information identifying the physical addresses assigned to the content stored in the aforementioned content pool, A metadata generation unit obtains information identifying a physical address from the content pool crawler, obtains content from the content pool based on the obtained information identifying a physical address, and generates metadata based on the obtained content. A vector space generation unit generates an N-dimensional (where N is an integer N > 1) vector space as metadata that represents the characteristics of the content based on multiple metadata, An embedding processing unit that embeds content based on metadata into the N-dimensional vector space generated by the vector space generation unit, A hierarchical structure generation unit generates k clusters (where k is an integer k > 1) from multiple contents embedded in the N-dimensional vector space, based on the semantic similarity of the contents. Equipped with, The N is the number of attribute items to be vectorized among the plurality of metadata. The information processing apparatus according to claim 1.
3. The hierarchical structure generation unit acquires a plurality of metadata generated based on each of the plurality of contents, vectorizes each of the acquired plurality of metadata, derives the distance between the vectorized plurality of metadata for each metadata, and generates k clusters based on the derived plurality of distances. The information processing apparatus according to claim 2.
4. The hierarchical structure generation unit creates a distance matrix based on the distances between the vectorized metadata, and generates k clusters based on the created distance matrices. The information processing apparatus according to claim 3.
5. The hierarchical structure generation unit weights the distances between the vectorized metadata and generates a distance matrix. The information processing apparatus according to claim 4.
6. The vector space generation unit creates M-dimensional (where M is an integer M > 1) vector data representing the attributes of the content creator, and generates an M-dimensional vector space. The embedding processing unit embeds the content into an N+M dimensional vector space. The information processing apparatus according to claim 2.
7. The vector space generation unit creates R-dimensional (where R is an integer R > 1) vector data based on management rules, and generates an R-dimensional vector space. The aforementioned embedding processing unit embeds the content into an N+M+R dimension vector space. The information processing apparatus according to claim 6.
8. The vector space generation unit creates T-dimensional vector data (where T is an integer T > 1) that represents time attributes, and generates a T-dimensional vector space. The aforementioned embedding processing unit embeds the content into an N+M+R+T dimension vector space. The information processing apparatus according to claim 7.
9. The aforementioned content processing unit, A pseudo-metadata search unit that generates pseudo-metadata based on a search prompt and searches for content based on the created pseudo-metadata. Equipped with, The display control unit causes the content retrieved by the pseudo-metadata search unit to be displayed on the display unit. The information processing apparatus according to claim 2.
10. The metadata generation unit, if there is any missing information in the metadata generated based on the content, creates information to query for the missing information, retrieves the information entered in response to the created query information, and completes the metadata. The information processing apparatus according to claim 2.
11. The aforementioned pseudo-metadata search unit derives the characteristics of the content based on the search results for the content, The display control unit causes the display unit to display information indicating the characteristics of the content derived by the pseudo-metadata search unit. The information processing apparatus according to claim 9.
12. The aforementioned pseudo-metadata search unit derives questions to narrow down the content based on the characteristics of the derived content, The display control unit causes the display unit to display information indicating a question for narrowing down the content derived by the pseudo-metadata search unit. The information processing apparatus according to claim 11.
13. The aforementioned pseudo-metadata search unit, if there is missing information to search for content, creates information to query for the missing information. The information processing apparatus according to claim 9.
14. The aforementioned receiving unit receives information indicating a hierarchical structure that represents the content in a hierarchical manner. The hierarchical structure generation unit generates clusters from multiple contents based on the information indicating the hierarchical structure. The information processing apparatus according to claim 2.
15. The aforementioned hierarchical structure generation unit generates k clusters from multiple content items in the background. The information processing apparatus according to claim 2.
16. The aforementioned hierarchical structure generation unit uses a large-scale language model to set the name of each of the k clusters. The information processing apparatus according to claim 2.
17. A method of information processing performed by a computer, Accepting multiple content, From a content pool in which information identifying the physical address is assigned to each of the received content items and stored, multiple content items are retrieved based on each of the pieces of information identifying the physical address. Metadata is generated based on each of the multiple acquired contents, and a vector space representing the characteristics of the contents is generated based on the multiple generated metadata. Based on the semantic similarity between the generated vector space and the content, multiple clusters are generated from multiple content pieces, thereby virtually classifying the content. The results of the virtual classification of the content are displayed on the display unit. When virtually classifying the aforementioned content, k clusters (where k is an integer k > 1) are generated from the multiple content items, and for each of the k clusters, the k clusters are optimized based on the average distance between any content item and other content items included in the cluster containing that content item, and the average distance between any content item and the cluster closest to that content item. Information processing methods.
18. On the computer, Accept multiple types of content, From a content pool in which information identifying the physical address is assigned to each of the multiple pieces of content that have been received and stored, multiple pieces of content are retrieved based on each of the pieces of information identifying the physical address. Based on each of the multiple retrieved content items, metadata is generated, and based on the generated metadata, a vector space representing the characteristics of the content is generated. Based on the semantic similarity between the generated vector space and the content, multiple clusters are generated from multiple content pieces, thereby virtually classifying the content. The results of the virtual classification of the content are displayed on the display unit. When virtually classifying the aforementioned content, k clusters (where k is an integer k > 1) are generated from multiple content items, and for each of the k clusters, the k clusters are optimized based on the average distance between any content item and other content items included in the cluster containing that content item, and the average distance between any content item and the cluster closest to that content item. program.
Citation Information
Patent Citations
Information processor and method, and program
JP2008070959A
Document search system and document search program
JP2021056581A
Clustering device, clustering method, and program
JP2022071929A
Document classification filter for search queries
US11036764B1
Content search device using dendrogram
JP2005078245A