Information processing system and information processing method

The system addresses structural and update frequency issues in information processing by automatically collecting, analyzing, and integrating unstructured data, enhancing data reliability and consistency through AI-driven dynamic optimization.

JP7867677B1Active Publication Date: 2026-06-01POLICY INNOVATION JAPAN CO LTD

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
POLICY INNOVATION JAPAN CO LTD
Filing Date
2025-11-28
Publication Date
2026-06-01

AI Technical Summary

Technical Problem

Conventional information processing systems struggle with structural changes in information sources, variations in update frequency, and issues like data duplication, inconsistency, and redundancy, failing to dynamically optimize and integrate unstructured data effectively.

Method used

An information processing system that automatically collects, analyzes, and structures unstructured data using AI and machine learning, optimizing access to external sources, integrating and managing data through data collection, analysis control, structuring, and integrated management means, including dynamic crawling and model selection based on data type and format.

Benefits of technology

This system significantly reduces manual effort, enhances data timeliness, comprehensiveness, and reliability, adapting to structural changes and update frequencies, ensuring consistent and accurate information integration and provision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007867677000001_ABST
    Figure 0007867677000001_ABST
Patent Text Reader

Abstract

This system and method provide an information processing system and method that automatically optimizes access to external information sources, performs semantic analysis of information using AI and machine learning processing, and integrates and structures the information. [Solution] The information processing system 100 uses multiple external information sources MA1 to MA n The system includes a data collection means 10 that periodically or dynamically cycles through to acquire unstructured data, and an analysis control means 20 that automatically determines the type of data based on the format or content of the acquired data and selects a corresponding extraction rule or language model. Furthermore, it includes a structuring means 30 that extracts information units using the selected model and converts them into structured data of a predetermined format, an integrated management means 40 that identifies the same object among multiple structured data sets, integrates duplicates, and generates unified data, and a data provision means 50 that stores the unified data in a database DB and provides it to an external device MC or external system MT via a communication network NT.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention relates to an information processing system and an information processing method, and more particularly to a technology that enables efficient information integration and provision by automatically collecting unstructured data from multiple external information sources and analyzing and structuring said data using AI (artificial intelligence) or machine learning processing. [Background technology]

[0002] In recent years, information in the political and administrative spheres, such as the activities of legislators and ministry officials, policy documents, meeting records, and government announcements, is scattered across numerous external websites and public information sources. This information is often provided as unstructured data in formats such as HTML, PDF, and social media posts, requiring enormous effort to collect and organize manually. Therefore, information processing systems such as RSS readers and web scraping have long been known as means of information gathering.

[0003] These systems automate the collection of news articles and public information by retrieving data from specified URLs or RSS feeds and regularly checking for updates. [Prior art documents] [Patent Documents]

[0004] [Patent Document 1] Japanese Patent Publication No. 2008-257695 [Patent Document 2] Japanese Patent Publication No. 2009-181548 [Overview of the project] [Problems that the invention aims to solve]

[0005] However, conventional information processing systems have problems such as being unable to cope with structural changes in information sources and variations in update frequency, and being prone to duplication, inconsistency, and redundancy in acquired data. Furthermore, dynamic optimization processes that semantically analyze unstructured data and automatically ensure consistency between information sources have not been performed.

[0006] This invention was made to solve the above-mentioned problems, and aims to provide an information processing system and information processing method that automatically optimizes access to external information sources, performs semantic analysis of information using AI and machine learning processing, and integrates and structures the information. [Means for solving the problem]

[0007] The information processing system and information processing method of the present invention include the following means to achieve the above objectives.

[0008] The information processing system of the present invention includes data collection means that periodically or dynamically visits a plurality of external information sources and acquires unstructured data from each external information source; analysis control means that automatically determines the type of unstructured data based on the format or content of the acquired unstructured data and selects an extraction rule or language model corresponding to that type; and using the selected extraction rule or language model, Across multiple external sources A structuring means that extracts information units from unstructured data, converts those information units into a predetermined data format to generate structured data, and takes the generated structured data as input. By combining a normalized proper noun dictionary, similarity calculation, and distance determination in the embedding space, In a structured set of predetermined data formats Across multiple external sources The system includes an integrated management means for identifying identical objects and integrating duplicate data to generate unified data, and a data provision means for storing the generated unified data in a database and transmitting it to an external device or system via a communication network.

[0009] This configuration allows for the automatic collection of unstructured data from multiple external sources, followed by analysis, structuring, and integrated management of that data. This significantly reduces the manual effort required for information gathering and organization, and enhances the timeliness, comprehensiveness, and reliability of information related to unstructured data. Furthermore, it can flexibly adapt to structural changes and differences in update frequencies of external sources, ensuring that information is always up-to-date and consistent.

[0010] Furthermore, the data collection means of the information processing system of the present invention detects the update frequency, communication response time, or structural changes of external information sources, and automatically changes the target of patrolling and the patrolling timing based on these factors.

[0011] This configuration enables efficient crawling based on the update patterns and communication load of each external information source, reducing unnecessary access while collecting the latest information with high accuracy. Furthermore, by automatically detecting structural changes and adjusting crawling accordingly, data acquisition errors due to site configuration changes can be prevented.

[0012] Furthermore, the analysis control means of the information processing system of the present invention extracts syntactic or lexical features of unstructured data, automatically classifies the data type using clustering or a machine learning model, and automatic classification of Based on the results, select an extraction rule or language model.

[0013] This configuration allows for the automatic selection of the optimal extraction model even for unstructured data of different writing styles, formats, and fields, improving analysis accuracy and versatility. Furthermore, by utilizing pre-trained models, adaptation to new data formats can be achieved quickly.

[0014] Furthermore, the structuring means of the information processing system of the present invention calculates a consistency score for information units extracted by selectively combining multiple language models, and generates the structured data based on the consistency score.

[0015] According to this configuration, the variation in extraction results among multiple models can be quantitatively evaluated, and only highly reliable information can be integrated, thus improving the accuracy and consistency of the output data. Also, by leveraging the complementary characteristics among models, a strong structuring process can be realized even for ambiguous expressions peculiar to natural language.

[0016] Moreover, the integration management means of the information processing system of the present invention detects differential information in a plurality of the structured predetermined data formats, and assigns the differential information as update history information to metadata.

[0017] According to this configuration, the history management of data updates is automatically performed, and the past information changes can be traced, making it easier to perform time-series analysis of policy changes and speech changes. Furthermore, by attaching the update history, it can also be applied to the reliability evaluation and forgery detection of information.

[0018] In addition, in the information processing system of the present invention, a series of processes from the data collection means to the data providing means are automatically executed by a scheduler or a workflow engine periodically or by an event trigger.

[0019] According to this configuration, the operation of the entire system is autonomized, and continuous collection, analysis, and provision of the latest data can be realized without manual intervention. Also, since immediate execution by an event trigger is possible, it is also suitable for monitoring information for which quick reporting is required.

[0020] In addition, the external information sources of the information processing system of the present invention are information of , Ministry members of parliament, parliamentary secretaries, bureaucrats , or office staff Regarding administrative personnel

[0021] According to this configuration, statements, policy documents, meeting minutes, etc. of stakeholders with different positions such as members of parliament, bureaucrats, and staff can be managed and analyzed in a unified format in a centralized manner.

[0022] of the present invention The information processing system executesThe information processing method periodically or dynamically cycles through multiple external information sources, acquires unstructured data from each external information source, automatically determines the type of unstructured data based on the format or content of the acquired unstructured data, selects an extraction rule or language model corresponding to that type, and uses the selected extraction rule or language model. Across multiple external sources Information units are extracted from unstructured data, these information units are converted into a predetermined data format to generate structured data, and the generated structured data is used as input. By combining a normalized proper noun dictionary, similarity calculation, and distance determination in the embedding space, In a structured set of predetermined data formats Across multiple external sources The system identifies identical objects, integrates duplicate data to generate unified data, stores the generated unified data in a database, and transmits it to an external device or system via a communication network. [Effects of the Invention]

[0023] The information processing system and information processing method of the present invention can automatically optimize access to external information sources, perform semantic analysis of information using AI and machine learning processing, and integrate and structure the information. [Brief explanation of the drawing]

[0024] [Figure 1] This is a system configuration and functional block diagram showing one embodiment of the information processing system according to the present invention. [Figure 2] This is a functional block diagram of the analysis and control means according to this embodiment. [Figure 3] This flowchart shows the processing flow of the analysis and control means according to this embodiment. [Figure 4] This flowchart shows the process of automatically collecting, analyzing, integrating, and providing information on members of parliament and government officials according to this embodiment. [Figure 5] This is an illustrative diagram showing an example of unstructured data in a different format according to this embodiment. [Figure 6] This figure shows an example of a politician information display screen according to this embodiment. [Modes for carrying out the invention]

[0025] The following describes an information processing system 100 based on an embodiment of the present invention.

[0026] 《Embodiment》 Figure 1 is a system configuration and functional block diagram showing one embodiment of the information processing system according to the present invention. Figure 2 is a functional block diagram of the analysis control means according to this embodiment. Figure 3 is a flowchart showing the processing flow of the analysis control means according to this embodiment. Note that the information processing system 100 according to this embodiment is not limited to the one shown in the figures, and may also include modifications to the illustrations and descriptions within the bounds of common sense.

[0027] (System Overview) The information processing system 100 according to this embodiment is suitable for, for example, automatically collecting unstructured data from external information sources in the political and administrative fields, analyzing and structuring said unstructured data using artificial intelligence (AI), and managing and providing it in an integrated manner.

[0028] The information processing system 100 according to this embodiment uses multiple external information sources MA1 to MA n This is an information processing system that automatically retrieves unstructured data, analyzes and structures the retrieved unstructured data, and integrates the structured data to generate unified data.

[0029] As shown in Figure 1, the information processing system 100 according to this embodiment uses multiple external information sources MA1 to MA n The external information sources MA1 to MA are cycled periodically or dynamically. nThe information processing system 100 comprises: a data collection means 10 that acquires unstructured data; an analysis control means 20 that automatically determines the type of unstructured data based on the format or content of the acquired unstructured data and selects an extraction rule or language model corresponding to that type; a structuring means 30 that uses the selected extraction rule or language model to extract information units from the unstructured data, convert the information units into a predetermined data format, and generate structured data; an integrated management means 40 that takes the generated structured data as input, identifies identical objects in a plurality of predetermined structured data formats, and integrates duplicate data to generate unified data; and a data provision means 50 that stores the generated unified data in a database DB and transmits it to an external device MC or external system MT via a communication network NT. The information processing system 100 also further comprises a scheduler 60A and a workflow engine 60B that control a series of processes of each of the above means.

[0030] The information processing system 100 may be a computer system equipped with a CPU, memory, a network interface for controlling data communication with an external network, and a bus for interconnecting each functional unit. The memory stores programs that cooperate with the CPU to control each functional unit or each unit. Furthermore, the memory also stores master data necessary for controlling each functional unit. The information processing system 100 can also be implemented on a general-purpose computer device, and each functional unit may be implemented by executing a program using a processor.

[0031] The information processing system 100 uses external information sources MA1 to MA n It is connected to the external system MT and the external device MC via the communication network NT.

[0032] External Information Sources MA1~MA n This constitutes the upstream information supply side, providing raw data (unstructured data) to the information processing system 100. External information sources MA1~MA n The information formats are diverse, including HTML, PDF, CSV, images, and text files, and the update frequency and data structure are not uniform.

[0033] In this embodiment, "unstructured data" refers to information data that is not organized according to a tabular format or schema structure in a database or the like. Specifically, this includes HTML, PDF, etc. , Te This includes data in formats that are human-readable but not suitable for machine processing, such as texts, images, audio, or social media posts. In this embodiment, this unstructured data is automatically collected and converted into machine-readable structured data by the analysis control means 20 and the structuring means 30, thereby facilitating searching, integration, and reuse.

[0034] For example, external information sources MA1~MA n This includes various types of publicly available information such as the official websites of the National Diet and local assemblies in the political and administrative fields (meeting minutes, meeting summaries, and bill information), press releases and public notices pages of various ministries, local governments, and independent administrative agencies, activity reports, blogs, and social media posts on the websites of political parties and individual members of parliament, official gazettes, public notices, open data APIs (e.g., data.go.jp), or news article RSS or XML feeds provided by news organizations.

[0035] The external system MT is an external client or information linkage system that receives structured data or integrated data generated by the information processing system 100 of the present invention and uses it secondarily. Specifically, examples include policy analysis tools, administrative information portals, media analysis systems, or corporate information dashboards. The external system MT acquires data from the data provision means 50 through API communication or file linkage. It can also be configured to return feedback information (access logs, evaluation values, etc.) as needed, and this information can be used for learning and optimization of the analysis control means 20.

[0036] The external device MC is a terminal device such as a personal computer or smartphone operated by user U.

[0037] The network NT is, for example, a wired LAN (Local Area Network), a wireless LAN, the Internet, a public switched telephone network, a mobile data communication network, or a combination thereof.

[0038] The data collection means 10 periodically or dynamically circulates through a plurality of external information sources MA1 to MA n and automatically acquires unstructured data such as HTML, PDF, image files, text data, or SNS posts. The data collection means 10 detects the update frequency, communication response time, or structure change of the external information sources MA1 to MA n and automatically changes the target of circulation and the timing of circulation based on these factors.

[0039] In this embodiment, "dynamically circulate" means monitoring the update frequency, communication response state, or change status of the site structure of each information source, and automatically changing the interval and order of circulation according to these conditions.

[0040] The data collection means 10 includes an update detection module and a communication state monitoring module that monitor the behavior of the external information sources MA1 to MA n The update detection module compares the hash value of the most recently acquired data, the update date and time metadata, or the ETag / Last-Modified information of the RSS header for each information source, and automatically determines whether there is an update.

[0041] The communication state monitoring module monitors the HTTP status code, response time, and number of timeouts, etc., and reduces the access frequency when the communication state falls below a certain threshold.

[0042]

[0043] ​Furthermore, the data collection means 10 includes a structure change detection module that automatically detects changes in the site structure. This structure change detection module compares the HTML document structure (tag hierarchy, DOM node configuration, class attributes, etc.) analyzed during the previous acquisition with the structure acquired this time, and detects changes in the web page structure by analyzing the frequency and area of ​​differences. The detection results are reflected in the collection control logic, which is configured to automatically correct scraping rules and XPath patterns.

[0044] The data collection means 10 automatically optimizes the patrol schedule based on these monitoring results. For example, it distributes the server load by shortening the patrol interval for information sources with a high update frequency, extending the patrol interval for information sources with a low update frequency, and shifting the patrol timing for information sources with a long communication response time.

[0045] Furthermore, for information sources where structural changes occur frequently, monitoring is limited to when changes occur, and access frequency is reduced during periods when no changes are detected. This allows for both increased efficiency in the monitoring process and suppression of excessive access to external servers.

[0046] Furthermore, the data collection means 10 records these control results as historical information in the database DB. The accumulated historical data is used for optimizing subsequent cycles using machine learning algorithms (for example, scheduling on workflow engines such as Prefect or Airflow®).

[0047] Next, the analysis control means 20 has the function of automatically determining the type (data type) of the unstructured data acquired by the data acquisition means 10 based on the format or content of the unstructured data, and selecting an extraction rule or language model associated with the determination result.

[0048] As shown in Figure 2, the analysis control means 20 includes a preprocessing unit 21, a feature extraction unit 22, a machine learning classification unit 23, a confidence evaluation unit 24, a rule / model selection unit 25, and an output control unit 26.

[0049] The following describes the processing flow of the analysis control means 20 with reference to the flowchart in Figure 3.

[0050] In this embodiment, "data type" includes at least (i) formal attributes (HTML, PDF, rich text including embedded images, etc.), (ii) content attributes (policy documents, meeting minutes, legislator profiles, news articles, social media posts, etc.), and (iii) update characteristics (static pages / highly updated pages).

[0051] First, the preprocessing unit 21 analyzes the format of the unstructured data and extracts (a) formal meta-features such as MIME type, file extension, HTTP header, PDF metadata (title / author / creation date), HTML DOM structure statistics, text length, paragraph hierarchy, presence or absence of tables, and presence or absence of images / attachments. Simultaneously, the feature extraction unit 22 extracts lexical and syntactic features from the document text, including tokenization, part-of-speech tags, named entities (personal names, job titles, committee names, law names, session numbers, etc.), n-grams, technical term dictionary match rates, keyword frequencies (TF-IDF), and semantic vectors (embeddings) (S201).

[0052] The feature extraction unit 22 generates embedding vectors on a document or paragraph basis based on these features, and represents the semantic distance between sentences in numerical space. This allows us to obtain basic data to be used for calculating similarity between sentences and for clustering (S202).

[0053] The machine learning classification unit 23 uses the features obtained from the extraction and embedding processes to automatically classify the data types using at least one of the following methods: (1) a natural group formation method using clustering (e.g., k-means, hierarchical clustering, DBSCAN, etc.) (S203), or (2) a multi-class classification method using a machine learning model (e.g., logistic regression, SVM, random forest, BERT, etc.) (S204). Clustering is suitable for detecting novel patterns (unknown classes), and machine learning classification is suitable for high-precision identification of known classes; therefore, a hybrid configuration using both methods is also possible (S205).

[0054] Typical classification labels include [Profile / Biography], [Policy Documents (White Papers, Draft Policies, Reports)], [Meeting Records (Minutes, Transcripts)], [Government Announcements / Press Releases], [News Articles], and [Short Social Media Posts (Announcements, Comments)]. These labels can be added, merged, and redefined during operation and are automatically applied in response to training data or cluster updates.

[0055] The confidence evaluation unit 24 assigns a confidence score to the classification result from the machine learning classification unit 23 (S206). This score uses silhouette coefficients or cluster center distance in the case of clustering, or softmax output or calibration probability (Platt scaling, isotonic regression, etc.) in the case of classification models. If the confidence score is below a predetermined threshold, robustness is ensured by (i) provisional extraction using a general-purpose language model, (ii) extraction of additional features and re-evaluation, or (iii) ensemble execution of rule-based extraction and LLM extraction.

[0056] The rule / model selection unit 25 selects the optimal extraction rule or language model for analyzing the unstructured data based on the output results of the confidence evaluation unit 24 and the machine learning classification unit 23 (S207).

[0057] Extraction rules are sets of formal or lexical rules used to mechanically extract specific information units (e.g., names of legislators, positions, committee names, session dates, policy names, statements, social media posts, etc.) from unstructured data. Examples include (a) CSS selectors / XPath for HTML, (b) regular expression templates for semi-structured text, (c) column header mapping for table structures, and (d) dictionary matching and normalization rules (position dictionaries, committee dictionaries, party dictionaries, place name dictionaries).

[0058] Examples of language models include (e) NER (Named Entity Recognition) models, (f) relation extraction models, (g) summary / headline extraction models, and (h) schema creation prompts using large-scale language models (LLMs) (e.g., JSON schema output).

[0059] The rule / model selection unit 25 selects these using a lookup table or a policy model. A policy model is a decision-making model for selecting the optimal extraction rule or language model based on the unstructured data type determination result and features. The policy model takes multiple parameters as input, such as data format (HTML, PDF, image-embedded documents, etc.), content attributes (meeting minutes, policy documents, SNS short messages, etc.), update characteristics, and past extraction accuracy history, and determines the most appropriate processing policy from various extraction pipelines. This makes it possible to dynamically select the optimal analysis procedure even when there are variations in document format or changes in site structure, thereby improving structured accuracy and robustness.

[0060] As an example, label-specific pipelines are defined as follows: (1) "Meeting Records" → Speaker Separation + Speech Extraction Model → JSON Conversion, (2) "Policy Documents" → Paragraph Analysis + Summarization Model → Tagging, (3) "Profiles" → Dictionary Matching Rules → Person / Role Schema Output, (4) "SNS" → Short Sentence Classification + Event Extraction → Activity Schema Output (S208).

[0061] Furthermore, the analysis control means 20 does not determine the model based solely on formal judgment, but also comprehensively evaluates it by using content features in conjunction. For example, even with PDF files, it does not limit itself to simple meeting minutes, but analyzes the occurrence trends of tables, chapter headings, meeting numbers, committee names, etc., using OCR and layout analysis to select the optimal model.

[0062] The output control unit 26 issues appropriate extraction instructions to the structuring means 30 based on the classification results and reliability evaluation, and outputs the results with metadata (model ID, execution time, accuracy index, etc.) added. In addition, the model parameters are automatically updated through online learning / continuous learning, and the duplicate and differential information identified by the integrated management means 40 is used as retraining data (S209).

[0063] Specifically, (i) the identification results of the same person / event by the integrated management means 40, (ii) notification viewing history, and (iii) extraction failure logs are used for retraining to dynamically improve classification and extraction accuracy. This makes it possible to keep up with changes in vocabulary and expressions, as well as changes in site structure, over time.

[0064] In this embodiment, for the Japanese political domain, the feature extraction unit 22 implements (a) dictionary normalization of the name of the House, session period, bill number, committee name, and official title, (b) conversion between Japanese and Western calendars, and (c) unification of variations in the spelling of personal names, thereby improving the classification accuracy and the identity determination accuracy of the integrated management means 40.

[0065] Evaluation metrics include accuracy, recall, F1 score, and unknown detection rate. Extraction is performed by measuring exact match rate, partial match rate, and normalized match rate for each schema item.

[0066] When a threshold is violated, the system automatically switches policies (e.g., LLM → rule-based), saves these operational metrics in a database, and periodically retrains the model and updates the thresholds using the scheduler 60A or workflow engine 60B described later. The scheduler 60A and workflow engine 60B may also be configured to automatically execute and monitor the entire process from collection to delivery using Prefect or an equivalent workflow management mechanism.

[0067] As described above, the analysis control means 20 performs a stepwise process consisting of (i) feature extraction → (ii) clustering or machine learning classification → (iii) automatic selection of extraction rules / language models. This enables the autonomous and highly accurate application of appropriate extraction strategies to documents that have the same format but different content, or documents that have similar content but different format.

[0068] Next, the structuring means 30 analyzes the acquired unstructured data using the extraction rules or language model selected by the analysis control means 20, and extracts information units (for example, "speaker name," "speech content," "speech date and time," "affiliation," "policy theme," etc.).

[0069] (1) Extraction process of information units (structuring process) The structuring means 30 first refers to the extraction rule set or language model group corresponding to the data type (e.g., meeting minutes, policy documents, government announcements, SNS posts, etc.) supplied by the analysis control means 20, and applies it to the target data.

[0070] For structured documents such as HTML, CSS selectors, XPath, or regular expression patterns are applied to extract elements based on tag structure and keywords.

[0071] For PDF or text data, the data is divided based on paragraphs, line breaks, or tabular format, and a named entity recognition model (NER model) or relational recognition model is applied to each unit to detect relationships such as speaker, subject, and action (statement, proposal, decision).

[0072] Furthermore, when using deep learning, a schema creation prompt is input to BERT, RoBERTa, or a Large-Scale Language Model (LLM), and information units are extracted according to a predefined JSON schema. After content validation, the extracted information units are converted into a structured format (e.g., JSO). Nmata The data is organized as a CSV file.

[0073] (2) Selective combination of multiple language models and calculation of consistency scores The structuring means 30 is not limited to extraction using a single language model, but also has the function of selectively combining multiple models for analysis.

[0074] For example, a named entity recognition model (NER model) and an relational recognition model are applied in parallel, and the matching rate of the "person's name," "statement content," "meeting name," and "related theme" detected by each model is evaluated.

[0075] The structuring means 30 calculates a consistency score from the output results of these multiple models. The consistency score is calculated based on the agreement rate of identical items detected by the multiple models, the variance of the score distribution, and the similarity between the model outputs (e.g., Cosine Similarity). In other words, the consistency score is a score calculated based on the agreement rate or similarity of extracted items among the multiple models.

[0076] This consistency score is represented by the following indicators: JPEG0007867677000002.jpg1157

[0077] Here, P (consistency) is the precision of items that matched across models, and R (consistency) is the recall of items that should match. Items with high consistency scores are adopted as reliable information, while those with low scores are re-extracted or imputed. This helps to absorb output variability between different models and reduce false extractions.

[0078] (3) Generation of structured data The structuring means 30 integrates multiple extraction results within the same data based on the consistency score described above and generates structured data for centralized processing by the integrated management means 40.

[0079] In this embodiment, "structured data" refers to data obtained by formatting information units extracted from unstructured data into a machine-readable format according to a predetermined schema or data model.

[0080] For example, items such as the name of the legislator, the content of the statement, the session period, the name of the committee, and the policy classification can be included in the JSO. Nmata This refers to a data structure that is hierarchically organized according to a standard format such as RDF, and can be stored or transmitted while maintaining the relationships between each item. Structured data is used in the integrated management means 40 described later as basic information for identifying identical objects, eliminating duplicates, and managing the history of data obtained from multiple information sources.

[0081] For example, if multiple sources of information regarding the same legislator's statements or policy-related tags exist in different external sources, the structuring method 30 prioritizes and integrates the information with the highest consistency score, eliminating duplication and inconsistencies.

[0082] The integrated data is registered in the database DB along with metadata (extraction date, source URL, score history, etc.). This allows for cross-sectional analysis and visualization at the individual and policy levels using the integrated management means 40 described later.

[0083] Next, the integrated management means 40 takes the generated structured data as input, identifies identical objects in multiple predetermined structured data formats, and integrates duplicate data to generate unified data. In other words, it receives multiple structured data generated by the structuring means 30, identifies and integrates identical objects, and also has the function of detecting and recording differences and change history.

[0084] The integrated management means 40 automatically compares and generates unified data when multiple external information sources exist and different data are obtained regarding the same person or the same event.

[0085] (1) Identification process for identical objects The integrated management means 40 processes each data record (JSO) output by the structuring means 30. N etc.The system analyzes attribute information contained in the text (e.g., name, affiliation, position, date, agenda, meeting name, URL, content of speech) to determine whether they are the same entity or not. This identification is not limited to simple string comparison, but combines normalized proper noun dictionaries (dictionary of member names, political party dictionary, committee dictionary, etc.), similarity calculations (e.g., string distance, Cosine Similarity, Jaccard coefficient), and distance determination in the embedding space (sentence-BERT semantic similarity evaluation).

[0086] For example, even data containing variations in spelling such as "Taro Yamada," "Member of Parliament Taro Yamada," and "Member of the House of Councillors Taro Yamada" can be treated uniformly as belonging to the same person after normalizing the name structure and title. Furthermore, it is possible to group multiple statements belonging to the same meeting using event information (session number, date, committee name, agenda) as a key. Through such determination, a set of candidate mergers for identical subjects is generated.

[0087] Multiple data points determined to be from the same subject are detected as duplicate data and prioritized based on content integrity score, update time, etc.

[0088] (2) Integration of duplicate data and generation of unified data The integrated management means 40 receives multiple structured data generated by the structuring means 30 as input, identifies identical objects, and then executes a process to generate unified data. In other words, for multiple data determined to be the same object, the integrated management means 40 performs a content comparison and executes a process to integrate duplicate data. Specifically, it performs the following integration process: • Duplicate values ​​in the same field are merged based on priority rules (e.g., latest date and time and high confidence score are prioritized). • If there are missing fields, fill them in with values ​​from other data. • Different fields will be stored as metadata with multiple values ​​(e.g., if the "location of the statement" is different).

[0089] The integrated data is recalculated by weighting reliability evaluation scores (consistency score and source reliability) and registered as "unified data" that retains the most consistent content.

[0090] This centralized data is structured hierarchically so that it can be referenced at different levels of granularity, such as individual, meeting, and policy levels.

[0091] (3) Detection of differential information The integrated management means 40 compares the unified data from the previous integration with the data newly received from the structuring means 30 and detects difference information (diff information). The targets of difference detection include the following: • Adding new items (e.g., new positions, committee appointments) • Updating existing items (e.g., changing titles or statements) • Deletion or expiration (e.g., end of term, removal of information source)

[0092] Difference detection uses hash comparisons for each key item and the amount of change in the embedded vector (e.g., cosine distance). After detecting changes, the type of change (addition / update / deletion) is classified, and an update history entry is generated for each data item.

[0093] (4) Addition of update history information (metadata extension) The integrated management means 40 adds the detected difference information to the metadata as "update history information." This history information is maintained for each record. This ensures the accuracy and transparency of the data and can also be used for evaluating the reliability of the information and for time-series analysis.

[0094] Next, the data provision means 50 stores the unified data generated by the integrated management means 40 in the database DB and transmits it to an external device MC or external system MT via the communication network NT.

[0095] (1) Database storage process The data provision means 50 stores the centralized data output from the integrated management means 40 in the database DB. The database DB stores structured data (e.g., JSO). Nmata The schema design is optimized for efficiently storing relational data, with a primary key such as a person ID, meeting ID, or document ID, and includes metadata such as update history information and confidence score.

[0096] The database (DB) can be implemented as an RDBMS (e.g., PostgreSQL, MySQL®) or a NoSQL database (e.g., MongoDB, Elasticsearch). Furthermore, an incremental update method is used for updates, managing both old and new versions while retaining historical data. This enables the reproduction of past state points (time travel queries).

[0097] (2) External transmission process External transmission of data provision means 50 can take the following forms depending on the application.

[0098] 1.API provision method The system responds to requests from external systems (external system MT) via a REST API or GraphQL API, returning the target data in JSON format. This allows external analysis systems, visualization dashboards, or news distribution services to directly access the centralized data of information processing system 100.

[0099] 2. Batch distribution method Regularly centralize the data and save it to files (CSV, JSO). N etc. The data is exported to a file system and sent to an external system via SFTP, HTTPS, or cloud storage (e.g., object storage). This method applies security settings, including transfer schedules, authentication information, and encryption protocols (TLS, SSH).

[0100] 3. Event-driven notification method It is also possible to adopt a configuration that automatically notifies an external system MT of update events using webhooks or Pub / Sub messaging (e.g., MQTT, Kafka, Amazon SNS, etc.) whenever the centralized data is updated.

[0101] In this case, by sending the differential data included in the update history information in the smallest possible units, real-time synchronization on the external system side becomes possible.

[0102] (3) Examples of integration with external devices and systems External System MT is an information analysis platform operated by external users or government agencies / media organizations, etc., and can perform analysis of legislator trends, policy comparisons, visualization of speech trends, etc., using centralized data provided by data provision means 50.

[0103] Furthermore, the external device MC is expected to be a general terminal (PC, smartphone, tablet, etc.), and the user U can view and search data from the information processing system 100 via a web browser or a dedicated application.

[0104] The HTTPS protocol is used for communication, and data is securely transmitted using TLS encryption. Furthermore, it is equipped with a logging module that monitors external usage (number of API calls, response time, error rate, etc.), and load balancing and cache control are also performed to maintain availability and throughput.

[0105] Next, the scheduler 60A manages the start time and priority of each processing job based on periodic execution or event triggers, for example, external information sources MA1~MA n The execution schedule of the data collection means 10 is automatically optimized according to the update frequency and communication response status.

[0106] (1) Periodic execution mode (scheduled trigger) Scheduler 60A periodically launches each processing module according to a pre-configured schedule. For example, it can automatically crawl government and parliamentary websites at 3:00 AM every day, re-fetch PDF white papers from government agencies every Monday, and update social media posts every 5 minutes, allowing users to set the optimal execution cycle for each information source.

[0107] The schedule is defined in cron format or as a policy table in a metadatabase. The Scheduler 60A also logs the execution results (success / failure / warning) of each job and automatically applies a retry policy (e.g., up to 3 retries) in case of abnormalities. This configuration allows the entire system to run unattended and on a regular schedule.

[0108] (2) Event trigger mode (dynamic execution control) The scheduler 60A can also dynamically initiate processing in response to event detection. Examples of event triggers include external information sources MA1~MA n Examples of events include update detection (RSS updates, site structure changes, HTTP status changes, etc.), new file detection (addition of PDF / HTML / CSV files), and external API notifications (Webhooks, Pub / Sub events). When these events occur, the scheduler 60A calls the workflow engine 60B and executes the necessary tasks (collection → analysis → structuring → integration → delivery) sequentially or in parallel. This allows for efficient reprocessing of only the information sources where changes have occurred, optimizing overall computing resources and communication load.

[0109] (3) Job monitoring and recovery control Scheduler 60A monitors the execution status, execution time, output size, and error logs of each task, and performs automatic recovery processing in the event of abnormal termination. Specifically, it includes fail-safe mechanisms such as retrying in case of communication errors, applying alternative models in case of syntax parsing failures, and queuing retransmission in case of database connection failures. This ensures redundancy that allows the system to continue operating without operator intervention.

[0110] The workflow engine 60B is a control module that executes each processing job started by the scheduler 60A sequentially or in parallel based on dependencies, coordinating the processes of data acquisition, analysis, structuring, integration, and delivery.

[0111] For example, flow control is performed such that the analysis control means 20 is activated when the data collection means 10 is completed, and the data provision means 50 is executed after waiting for the results of the structuring means 30 and the integrated management means 40. This allows the entire data processing flow from information collection to distribution to be executed automatically, improving the processing efficiency and reliability of the entire system.

[0112] Workflow definitions are represented in DAG (Directed Acyclic Graph) format and can be implemented using task orchestration platforms such as Prefect, Airflow, and Luigi. This allows individual processes to be reused as independent tasks, and facilitates partial re-execution and rollback in case of errors.

[0113] Figure 4 is a flowchart illustrating the process of automatically collecting, analyzing, integrating, and providing information on members of parliament and government officials according to this embodiment. Figure 5 is an illustrative diagram showing an example of unstructured data in a different format according to this embodiment. Figure 6 is a diagram showing an example of a politician information display screen according to this embodiment.

[0114] Next, following Figure 4, we will explain the process of automatically collecting, analyzing, integrating, and providing information on members of parliament and government officials.

[0115] The scheduler 60A automatically cycles through external information sources MA1 to MA3 (see Figure 5), such as the National Diet, various committees, and administrative agencies, based on a pre-configured schedule or event trigger (S401). This allows it to detect updates to National Diet meeting minutes, meeting summaries, policy announcements, press releases, etc. (S402).

[0116] The data collection means 10 receives an execution command from the scheduler 60A, accesses the detected external information sources MA1 to MA3, and collects new or updated unstructured data (HTML, PDF). 、S The system automatically retrieves NS posts, etc. During retrieval, the system performs structural analysis of the target page, character encoding determination, and removal of unnecessary information (advertisements, banners, etc.) (S402).

[0117] The analysis control means 20 analyzes the acquired unstructured data and automatically determines the data type based on its format and content (S403). Specifically, it performs natural language processing and document feature extraction to classify the data into categories such as "records of parliamentary speeches," "policy documents," "news articles," and "SNS posts." Depending on the classification result, it selects the corresponding extraction rule or language model.

[0118] The structuring means 30 applies selected extraction rules or language models to extract information units such as "member's name," "position," "content of statement," "session period," "committee name," and "related policy" from the document (S404). The extraction results are formatted according to a predetermined schema (JSON format, RDF format, etc.) and output as structured data.

[0119] The integrated management means 40 takes the generated structured data as input, identifies the same member of parliament from among multiple structured data sets, normalizes variations in name spelling (e.g., "Yamada Taro" vs. "Yamada Taro") and differences in job titles, and then integrates the duplicate data (S405). This creates a consistent data record for each member of parliament.

[0120] The integrated management means 40 further detects differential information by comparing it with existing data (S406). For example, changes in position, additions to statements, or updates to policy recommendations are extracted as differentials, and their historical information is attached as metadata. This makes it possible to track changes in politicians' activities over time as a history.

[0121] The data provision means 50 stores the updated centralized data in a database DB and distributes it to an external system MT or external device MC via the communication network NT (S407).

[0122] External systems like MT can be, for example, politician dashboards or policy analysis platforms, capable of performing visualization and statistical processing using received data.

[0123] The external system MT automatically updates the speech history, committee affiliations, and trend analysis by policy area for each politician based on the latest data provided (S408). This allows users to grasp political activity information in near real-time (see Figure 6). In addition, users can centrally manage and analyze speeches, policy documents, meeting records, etc., from stakeholders with different positions, such as legislators, bureaucrats, and staff, in a unified format. This enables cross-platform search and comparative analysis, which previously required individually referring to each ministry's website, parliamentary page, press releases, etc., on a single platform.

[0124] The workflow engine 60B records processing logs and statistical information once it confirms that all processing steps have been completed successfully. This allows for analysis of processing time, success rate, and update frequency. After processing is complete, it sets the next execution schedule in the scheduler 60A and returns the system to standby mode (S409).

[0125] In this embodiment, the scheduler 60A and the workflow engine 60B are configured as independent modules, but the system is not limited to this, and they may be implemented as an integrated control module. In this case, each processing step (collection, analysis, structuring, integration, and provision) can be defined in DAG (directed acyclic graph) format, and the optimal execution order can be determined based on automatic analysis of dependencies. This facilitates the distribution of processing load, dynamic optimization of the execution order, and partial re-execution in the event of a failure.

[0126] Furthermore, the workflow engine 60B may be configured to self-optimize the job schedule for subsequent tasks based on the analysis results of the processing logs. Specifically, it uses processing time, error frequency, or data acquisition amount as features to automatically adjust the next execution interval or degree of parallelism using a machine learning algorithm (e.g., reinforcement learning, Bayesian optimization, etc.). This allows for dynamic optimization of the overall system operation efficiency in response to fluctuations in the update frequency of information sources and network load.

[0127] Furthermore, the functions of Scheduler 60A and Workflow Engine 60B may be configured to work in conjunction with a container orchestration platform on a cloud environment (e.g., Kubernetes®, Docker Swarm, etc.). In this case, each processing task is scaled on a container basis and automatically scales out / scales in according to the system load and number of concurrent executions, making it possible to flexibly handle large amounts of data processing and high-frequency updates.

[0128] In addition, the workflow engine 60B may perform distributed execution control of tasks based on the execution metadata (job ID, dependent tasks, priority) received from the scheduler 60A. This allows tasks to be reassigned by other nodes even in the event of a single node failure, thereby increasing the overall availability and redundancy of the system.

[0129] Furthermore, the above configuration may be simplified by omitting the workflow engine 60B and performing overall control with the scheduler 60A alone, or conversely, by incorporating scheduling functionality into the workflow engine 60B for integrated control. This allows for flexible configuration selection depending on the deployment environment and system scale.

[0130] Finally, the above-described embodiments should be considered in all respects to be illustrative and not restrictive. The scope of the invention is indicated by the claims, not by the embodiments described above. Furthermore, the scope of the invention is intended to include all modifications within the meaning and scope equivalent to the claims. [Explanation of symbols]

[0131] 10…Data collection methods 20…Analysis and control means 21…Pre-processing section 22...Feature extraction unit 23…Machine Learning Classification Department 24…Reliability Evaluation Department 25…Rule / Model Selection Section 26…Output Control Unit 30…Structuring means 40…Integrated management means 50…Method of providing data 60A... Scheduler 60B…Workflow Engine MA1~MA n …external information sources MC…External device MT...External system DB...Database NT... Communications Network

Claims

1. A data collection means that periodically or dynamically visits multiple external information sources and acquires unstructured data from each external information source, An analysis control means that automatically determines the type of unstructured data based on the format or content of the acquired unstructured data and selects an extraction rule or language model corresponding to that type, A structuring means that extracts information units from the unstructured data spanning multiple external information sources using the selected extraction rule or language model, and converts the information units into a predetermined data format to generate structured data; An integrated management means that takes the generated structured data as input, combines a normalized proper noun dictionary, similarity calculation, and distance determination on the embedding space to identify the same object across multiple external information sources in multiple predetermined structured data formats, and integrates duplicate data to generate unified data. A data provision means that stores the generated centralized data in a database and transmits it to an external device or external system via a communication network, An information processing system characterized by comprising the following features.

2. The information processing system according to claim 1, characterized in that the data collection means detects the update frequency, communication response time, or structural changes of the external information source, and automatically changes the target of patrolling and the patrolling timing based on these factors.

3. The information processing system according to claim 1, characterized in that the analysis control means extracts syntactic or lexical features of the unstructured data, automatically classifies the data type using clustering or a machine learning model, and selects the extraction rule or the language model based on the result of the automatic classification.

4. The information processing system according to claim 1, characterized in that the structuring means calculates a consistency score of the information units extracted by selectively combining multiple language models, and generates the structured data based on the consistency score.

5. The information processing system according to claim 1, characterized in that the integrated management means detects difference information in a plurality of structured predetermined data formats and adds the difference information to the metadata as update history information.

6. The information processing system according to claim 1, characterized in that the series of processes from the data collection means to the data provision means are automatically executed periodically or by event triggers by a scheduler or workflow engine.

7. The information processing system according to claim 1, characterized in that the external information source is information relating to members of parliament, parliamentary secretaries, bureaucrats, ministry officials, or administrative personnel.

8. We periodically or dynamically visit multiple external information sources and retrieve unstructured data from each external information source. Based on the format or content of the acquired unstructured data, the type of unstructured data is automatically determined, and an extraction rule or language model corresponding to that type is selected. Using the selected extraction rule or language model, information units are extracted from the unstructured data spanning multiple external information sources, and these information units are converted into a predetermined data format to generate structured data. Using the generated structured data as input, a normalized proper noun dictionary, similarity calculation, and distance determination on the embedding space are combined to identify the same object across multiple external information sources in multiple predetermined structured data formats, and duplicate data is integrated to generate unified data. An information processing method performed by an information processing system, comprising storing the generated centralized data in a database and transmitting it to an external device or external system via a communication network.