Information aggregation apparatus, information aggregation method, and information aggregation program
The information aggregation device addresses the challenge of consolidating diverse supply chain information by vectorizing text, clustering, and generating summary sentences, effectively aggregating overlapping information for improved understanding.
Patent Information
- Application Number
- JP2023187470
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-01
- Publication Date
- 2025-05-15
AI Technical Summary
Existing information aggregation methods struggle to consolidate diverse and unstructured information from various sources in the supply chain, particularly due to variations in granularity and format, as well as the presence of natural language descriptions.
An information aggregation device that extracts relevant information, vectorizes text using a language model, classifies vectors into clusters, generates summary sentences for each cluster, and calculates representative values to aggregate information effectively.
Enables the appropriate aggregation and provision of multiple overlapping pieces of information, improving understanding and reducing communication complexity, especially in complex domains like supply chain vulnerability management.
Smart Images

Figure 2025075939000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to an information aggregating device, an information aggregating method, and an information aggregating program for aggregating and providing a plurality of pieces of information. [Background technology]
[0002] Conventionally, repositories (databases) have been used to allow multiple users to share information obtained from multiple information providers. For example, vulnerability information for a product in the supply chain from procurement of the parts that make up the product to sales consists of information on the parts of the product (Bill of Materials, BOM) and vulnerability information linked to each part, etc. Since this information is sometimes provided by each supplier of the parts that make up the product, the granularity and format of the information are not standardized, making it difficult to consolidate vulnerability information. Non-Patent Document 1 proposes a method of aggregating information using a knowledge graph. By describing information in a specific format and associating multiple pieces of information to form a knowledge graph, it becomes possible to obtain comprehensive information about the supply chain, which is difficult to do by simply checking each piece of information. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] W. Zhang et al., "The construction of a domain knowledge graph and its application in supply chain risk analysis," IEEE International Conference on E-Business Engineering (ICEBE), 2019. [Non-Patent Document 2] A. Radford et al., "Improving language understanding by generative pre-training," 2018,<https: / / cdn.openai.com / research-covers / language-unsupervised / language_understanding_paper.pdf> . Summary of the Invention [Problem to be solved by the invention]
[0004] However, in conventional technology, information had to be formatted into a common format in order to be aggregated. In a supply chain, many businesses register information about a wide variety of products, making it difficult to describe this information in a common format. Even if the information were described in a common format through standardization, individual and specific information may still be described in natural language as supplementary information, making it difficult to aggregate such unstructured information. For example, when explaining that a certain sensor is vulnerable to strong ambient light, it was difficult to aggregate variations in notation such as "strong light" and "bright light" using a mechanical method such as a rule base.
[0005] An object of the present invention is to provide an information aggregating device, an information aggregating method, and an information aggregating program that can appropriately aggregate and provide a plurality of overlapping pieces of information. [Means for solving the problem]
[0006] The information aggregation device of the present invention includes a data extraction unit that extracts a set of information linked to a specified identifier from a repository in which multiple types of information are linked to each other; a vectorization unit that converts strings of specified items contained in each of multiple pieces of information of the same type in the set into vectors using a language model; a classification unit that classifies the strings into clusters using a specified clustering method based on the distance between each of the vectors; a text generation unit that combines the strings classified into each cluster for each cluster and then generates a summary sentence using a generation language model; and an output unit that converts the summary sentence into information that aggregates the specified items, formats the set, and outputs it.
[0007] The information aggregation device may include a representative value calculation unit that calculates, for each cluster, a representative value obtained by statistically processing values of other items, different from the specified item, that are included in each of the multiple pieces of information, and the output unit may treat the representative value as information that aggregates the other items, and may shape and output the set.
[0008] The text generator may aggregate the plurality of pieces of information by weighting them according to the proximity of the pieces of information from the centroid vector of each cluster.
[0009] The text generator may repeatedly combine the character strings a number of times proportional to the weights, and then generate a summary sentence using a generative language model.
[0010] The text generation unit may generate text as the summary sentence by using the language model with the centroid vector of each cluster as an input.
[0011] The multiple types of information may include product information and vulnerability information, and the data extraction unit may extract the vulnerability information for each component linked to a certain product from the product information as the multiple types of information of the same type.
[0012] The information aggregation method of the present invention includes a data extraction step of extracting a set of information linked to a specified identifier from a repository in which multiple types of information are linked to each other; a vectorization step of converting strings of specified items contained in each of multiple pieces of information of the same type among the set into vectors using a language model; a classification step of classifying the strings into clusters using a specified clustering method based on the distance between each of the vectors; a text generation step of combining the strings classified into each cluster for each cluster, and then generating a summary sentence using a generative language model; and an output step of converting the summary sentence into information that aggregates the specified items, formatting the set, and outputting it.
[0013] An information aggregating program according to the present invention causes a computer to function as the information aggregating device. Effect of the Invention
[0014] According to the present invention, a plurality of overlapping pieces of information can be appropriately consolidated and provided. [Brief description of the drawings]
[0015] [Figure 1] FIG. 1 is a diagram illustrating an example of a configuration of a repository system for managing vulnerability information in an embodiment. [Diagram 2] 2 is a block diagram showing a functional configuration of an information aggregating device according to an embodiment. FIG. [Diagram 3] 10 is a flowchart showing a flow of an aggregation process performed by the information aggregating device according to the embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0016] An example of an embodiment of the present invention will now be described. The information aggregation method of this embodiment aggregates overlapping information extracted from a repository using a language model. Note that the focus here is on a collection of hierarchically associated information, that is, information that can be used as a starting point to reference related information as a collection. As an example, we will take vulnerability information in a supply chain.
[0017] FIG. 1 is a diagram illustrating an example of the configuration of a repository system 100 for managing vulnerability information in this embodiment. The information handled in the repository is broadly divided into three types: product information, company information, and vulnerability information.
[0018] The information input unit 110 has an input function corresponding to each information type, acquires information from a corresponding external system, and records it in a corresponding table in the information storage unit 120. The information storage unit 120 configures a table for each type of information and stores each piece of information. Here, the information included in each table is associated with each other via an identifier such as an ID.
[0019] The information reforming unit 130 aggregates information based on the information stored in each table of the information storage unit 120, and outputs the information to an external system. In this embodiment, the information reforming unit 130 includes the information aggregating device 1.
[0020] Product information includes information about the products that make up the hardware, and includes, for example, a value that uniquely identifies the product information (product information ID), the product name, product model number, information about the company that manufactured the product (company information ID), the use or application of the product (product use), and a list of products used in the product (parts list).
[0021] Company information includes information about the company that provides the hardware, and includes, for example, a value that uniquely identifies the company information (company information ID), the company name, the nationality of the company (company nationality), and information about the company's subsidiaries, affiliated companies, business partners, etc. (affiliated company list).
[0022] Vulnerability information is information that includes known vulnerability information related to hardware, and includes, for example, a value that uniquely identifies the vulnerability information (vulnerability information ID), a value that indicates the degree of threat of the vulnerability (vulnerability threat level), the date of vulnerability registration, the date of vulnerability update, the conditions under which the vulnerability occurs (vulnerability conditions), the impact when the vulnerability is attacked (vulnerability impact), the source of the vulnerability information (vulnerability information source), and a list of affected products (product list). Vulnerability information can also include information such as a flag that indicates that doubts arose during testing.
[0023] For example, product information for a "notebook PC" is associated with product information such as a "CPU" and an "SSD" as a parts list, and the product information for the "CPU" and the "SSD" is further associated with corresponding vulnerability information. When aggregating vulnerability information related to a "notebook PC," vulnerability information associated with product information included in a parts list such as "CPU" and "SSD" is aggregated. In this case, if vulnerability information for "CPU" and "SSD" is described in different ways, such as "causing malfunctions when exposed to high temperatures," the information aggregation device 1 will aggregate such information with similar meanings into one and output it.
[0024] FIG. 2 is a block diagram showing a functional configuration of the information aggregation device 1 in this embodiment. The information aggregating device 1 is an information processing device (computer) including a control unit 10, a storage unit 20, and various input / output interfaces.
[0025] The control unit 10 is a part that controls the entire information aggregation device 1, and realizes the functions of each functional block described later by appropriately reading and executing various programs stored in the storage unit 20. The control unit 10 is not particularly limited, but may be a CPU.
[0026] The memory unit 20 is a storage area for various programs and various data for causing the hardware group to function as the information aggregation device 1, and may be, but is not limited to, ROM, RAM, flash memory, a hard disk drive (HDD), a solid state drive (SSD), etc.
[0027] The control unit 10 includes a data extraction unit 11 , a vectorization unit 12 , a classification unit 13 , a text generation unit 14 , a representative value calculation unit 15 , and an output unit 16 .
[0028] The data extraction unit 11 extracts a set of information linked to a specified identifier from a repository in which information of a plurality of information types is linked to one another. Specifically, in this embodiment, the multiple information types include product information and vulnerability information, and the data extraction unit 11 extracts vulnerability information for each of multiple components linked to a certain product (product information ID) in the product information as a set including multiple pieces of information of the same information type.
[0029] The vectorization unit 12 converts character strings of a predetermined item contained in each of a plurality of pieces of information of the same information type from the set extracted by the data extraction unit 11 into vectors using a language model.
[0030] Here, a language model is a model of the occurrence probability of characters or words in a natural language by analyzing a large number of documents written in the natural language in advance, and an example of this is the Generative Pre-trained Transformer (GPT) shown in Non-Patent Document 2. Using a language model, it is possible to vectorize text written in natural language and generate directed text. For example, by vectorizing two texts, it is possible to express the semantic similarity of these texts using a mathematical index such as cosine similarity, and to generate a summary of the input text.
[0031] The classification unit 13 classifies character strings corresponding to each vector generated by the vectorization unit 12 into clusters using a predetermined clustering method based on the distance between each vector. Clustering methods are generally classified into hierarchical clustering algorithms such as Ward's method and non-hierarchical clustering algorithms such as k-means. In hierarchical clustering algorithms, a threshold τ is set for the distance function d, and data that is less than the threshold is grouped into one cluster. In non-hierarchical clustering algorithms, for example, the number of clusters to be generated is specified in the k-means method. In this way, the granularity at which data is aggregated into clusters in the aggregation process can be adjusted as appropriate using parameters of various algorithms.
[0032] The text generating unit 14 combines the character strings classified into each cluster generated by the classifying unit 13, and then generates a summary sentence using a generative language model. At this time, the text generator 14 may aggregate multiple pieces of information after weighting them according to the distance from the centroid vector of each cluster. Specifically, the text generator 14 repeatedly combines character strings a number of times proportional to the weights, and then generates a summary sentence using a generative language model. Furthermore, the text generating unit 14 may generate text using a language model with the centroid vector as an input, and use the generated text as a summary sentence.
[0033] The representative value calculation unit 15 calculates a representative value for each cluster by performing statistical processing on values of items other than the predetermined item targeted by the text generation unit 14, which are included in each of the plurality of pieces of information. The items that are statistically processed by the representative value calculation unit 15 are information related to text information. For example, vulnerability information may be associated with values indicating its importance or impact, and such values can be aggregated together with the text information of the vulnerability information. The representative value obtained by the statistical processing can be set appropriately according to the information, for example, a maximum value, a minimum value, an average value, or the like.
[0034] The output unit 16 replaces the summary generated by the text generation unit 14 with information that summarizes predetermined items, formats the set, and outputs it. At this time, the output unit 16 may further replace the representative value calculated by the representative value calculation unit 15 with information that summarizes other items, and then shape the set.
[0035] FIG. 3 is a flowchart showing the flow of the aggregation process by the information aggregating device 1 in this embodiment. Here, we will show the procedure for consolidating duplicate information when collecting parts information starting from a certain product when sharing vulnerability information in the supply chain. Assume that a product P uses multiple parts, and Ns pieces of text information are associated with the product information for each part. In this case, the information aggregation device 1 aggregates similar information from the Ns pieces of information through aggregation processing, and obtains Nc (≦Ns) pieces of text information.
[0036] In step S1, the data extraction unit 11 extracts and collects multiple pieces of vulnerability information linked by product information IDs or the like from a repository.
[0037] In step S2, the vectorization unit 12 vectorizes each of the character strings of a predetermined item among the collected information. Let S={s 1 ,s 2 ,…,s Ns At this time, the vectorization unit 12 applies a language model to the collected character strings to associate each character string with a vector. If the language model is the function LM(·), the vector corresponding to the i-th string is v i =LM(s i As a result of vectorization, a set of vectors V = {v i :1≦i≦Ns} is obtained.
[0038] In step S3, the classification unit 13 classifies the information vectorized in step S2. Here, two samples v i ,v j The function to obtain the distance between v i ,v j ) As an example of a function d, let us take two samples v i ,v j The Euclidean distance between the samples is used as a distance function, and the operation of classifying each sample into a cluster using this distance function is generally called clustering.
[0039] The classification unit 13 classifies the samples into Nc clusters using a predetermined clustering algorithm. i is the cluster c i (1≦c i ≦Nc). The number of clusters Nc is either explicitly specified as a specific value in advance, or is implicitly determined by setting the above-mentioned threshold τ, as in the hierarchical clustering algorithm.
[0040] In step S4, the text generator 14 generates text as a summary sentence for each cluster classified in step S3. Here, let Tc be the set of strings classified into cluster c, {s i |s i ∈S,c i = c}. Let the function CONCATENATE(·) be a function that combines a set of given texts into one text, and the function SUMMARY(·) be a generative language model that summarizes the given text. Note that the generative language model SUMMARY(·) may be a model different from the above-mentioned LM(·).
[0041] The text generator 14 generates a text t representing each cluster c by the following process: c Generate. t c =SUMMARY(CONCATENATE(Tc)) By processing each cluster c in the same way, Nc pieces of text (summary sentences) are obtained.
[0042] In step S5, the representative value calculation section 15 may further aggregate the related information of the text aggregated for each cluster in step S4. Here, the string to be aggregated corresponding to the i-th vulnerability information is s i , the i-th piece of information in the set of related information X is x i Far away. A set of numerical values (or label information equivalent to numerical values) corresponding to the strings classified into cluster c {x i |x i ∈X,c i =c, 1≦i≦Ns}, the related information r c can be expressed as follows using the aggregation function AGGREGATE(·): r c =AGGREGATE({x i |x i ∈X,c i =c,1≦i≦Ns})
[0043] For example, a value (such as a numerical value or a symbol indicating a level) indicating the importance or impact of vulnerability information may be associated with the vulnerability information. In this case, the representative value calculation unit 15 aggregates the value together with the character string of the vulnerability information. In this case, the aggregation function AGGREGATE(·) is, for example, a function that returns the maximum value of the set.
[0044] In step S6, the output unit 16 replaces the original data with the information (summary sentences and representative values) collected in steps S4 and S5, and outputs the information collected in step S1.
[0045] This aggregation process aggregates and outputs duplicate information from the collected vulnerability information. In step S4, when generating a summary sentence that aggregates character strings, the text generation unit 14 may weight the information to be aggregated using information obtained during clustering.
[0046] For example, the centroid vector of cluster c is g c Then, the vector v of strings belonging to cluster c i By using a distance function for i W i =1 / d(g c ,v i ) can be assigned as the weight w i Let the set of be Wc. At this time, the text generator 14 uses a function CONCATENATE'(Tc,Wc) obtained by modifying the above-mentioned function CONCATENATE(Tc) so that it can receive a set of weights Wc. CONCATENATE'(Tc,Wc) is, for example, i The string is repeatedly concatenated a number of times proportional to . This allows you to specify which information to prioritize when aggregating.
[0047] Also, the center of gravity vector g c is a representative of multiple pieces of information, calculated by weighting each piece of information. Therefore, the text generator 14 uses this center of gravity vector g c Alternatively, a summary may be generated directly using a language model with the above as input.
[0048] According to this embodiment, the information aggregating device 1 refers to and aggregates information on various parts in a supply chain for a certain product. The information aggregating device 1 vectorizes information described in natural language for each part, and performs clustering on the vectorized information using a clustering algorithm. The information aggregating device 1 then aggregates information for each cluster, and generates one summary sentence using a language model. Therefore, the information aggregating device 1 can aggregate a specific type of information as a summary sentence for a collection of information in which multiple pieces of information are linked starting from a certain piece of information (for example, a product information ID). This allows the information aggregating device 1 to appropriately aggregate and provide multiple pieces of overlapping information, which is expected to improve comprehension and reduce the amount of communication when linking with an external system. This improved understandability is particularly useful when a large amount of related information is linked, when advanced knowledge is required to interpret the information, or when the information needs to be interpreted quickly, such as in the supply chain vulnerability management mentioned above.
[0049] In addition to the summary sentence, the information aggregation device 1 can also aggregate values other than text at the same time using the clustering results, thereby making it possible to provide a combination of representative values such as numerical values with the summary sentence, which is expected to further improve understandability.
[0050] Furthermore, the information aggregating device 1 can weight each piece of information based on the relationship between each piece of information within a cluster, and can aggregate information more appropriately.
[0051] In addition, this embodiment can appropriately aggregate a large amount of collected information, for example, by removing duplicates, making it possible to contribute to Goal 9 of the United Nations-led Sustainable Development Goals (SDGs), which is to "build resilient infrastructure, promote sustainable industrialization and foster innovation."
[0052] Although the embodiments of the present invention have been described above, the present invention is not limited to the above-described embodiments. Furthermore, the effects described in the above-described embodiments are merely a list of the most preferable effects resulting from the present invention, and the effects of the present invention are not limited to those described in the embodiments.
[0053] The information aggregation method by the information aggregation device 1 is realized by software. When realized by software, a program constituting this software is installed in an information processing device (computer). These programs may be recorded on a removable medium such as a CD-ROM and distributed to users, or may be distributed by being downloaded to the user's computer via a network. Furthermore, these programs may be provided to the user's computer as a Web service via a network without being downloaded. [Explanation of symbols]
[0054] 1. Information aggregation device 10 Control section 11 Data Extraction Section 12 Vectorization Department 13 Classification section 14 Text Generation 15 Representative value calculation section 16 Output section 20 Memory section
Claims
1. a data extraction unit that extracts a set of information associated with a specified identifier from a repository in which multiple types of information are associated with each other; a vectorization unit that converts character strings of a predetermined item included in each of a plurality of pieces of information of the same type from the set into vectors using a language model; a classification unit that classifies the character strings into clusters using a predetermined clustering technique based on the distances between the vectors; a text generation unit that generates a summary sentence by combining the character strings classified into each cluster using a generative language model; and an output unit that converts the summary sentence into information that summarizes the specified items and formats and outputs the collection.
2. a representative value calculation unit that calculates, for each cluster, a representative value obtained by statistically processing values of items other than the predetermined item that are included in each of the plurality of pieces of information; The information aggregating device according to claim 1 , wherein the output unit treats the representative value as information aggregating the other items, and formats and outputs the set.
3. 3 . The information aggregating device according to claim 1 , wherein the text generating unit aggregates the plurality of pieces of information after weighting them according to a distance from a centroid vector of each of the clusters.
4. The information aggregating device according to claim 3 , wherein the text generating unit repeatedly combines the character strings a number of times proportional to the weights, and then generates a summary sentence using a generative language model.
5. 3 . The information aggregating device according to claim 1 , wherein the text generating unit generates text by using the language model with the centroid vector for each cluster as an input, and sets the text as the summary sentence.
6. The multiple types of information include product information and vulnerability information, The information aggregating device according to claim 1 , wherein the data extracting unit extracts the vulnerability information of each part linked to a certain product from the product information as the plurality of pieces of information of the same type.
7. A data extraction step of extracting a set of information linked to a specified identifier from a repository in which multiple types of information are linked to each other; a vectorization step of converting character strings of a predetermined item included in each of a plurality of pieces of information of the same type from the set into vectors using a language model; a classification step of classifying the character strings into clusters using a predetermined clustering technique based on the distances between the respective vectors; a text generation step of combining the character strings classified into each cluster and then generating a summary sentence using a generative language model; an output step of converting the summary sentence into information that summarizes the specified items, and outputting the set in a formatted form, the information aggregating method being executed by a computer.
8. An information aggregating program for causing a computer to function as the information aggregating device according to claim 1.