Information processing system, information processing method, and program

The information processing system improves AI output reliability by evaluating and filtering input information, addressing noise and error propagation in conventional systems.

JP7863942B1Active Publication Date: 2026-05-22PROPERTY INNOVATION CONSULTING CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
PROPERTY INNOVATION CONSULTING CO LTD
Filing Date
2026-02-13
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Conventional AI-based information processing systems face issues with unreliable output due to insufficient control over the quality and scope of input information, leading to noise, misleading results, and a risk of self-replicating errors in autonomous AI agents.

Method used

An information processing system that evaluates and distinguishes between information to be preferentially used and excluded by AI, using analysis and control mechanisms to improve accuracy and reliability.

Benefits of technology

Enhances the accuracy and reliability of AI processing results by filtering noise and preventing self-replication of errors, ensuring stable and reliable output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007863942000001_ABST
    Figure 0007863942000001_ABST
Patent Text Reader

Abstract

To improve the accuracy and reliability of processing results when using AI. [Solution] An information processing system using artificial intelligence, which evaluates information or a region within a target space that is the target of processing by the artificial intelligence, and, based on the evaluation, performs processing to distinguish between information that should be preferentially used for processing by the artificial intelligence and information that should be excluded or suppressed from processing by the artificial intelligence within the target space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to an information processing system, an information processing method, and a program.

Background Art

[0002] In recent years, with the progress of digital transformation (DX), the importance of information processing systems that utilize artificial intelligence (AI), particularly large language models (LLMs), has been increasing in corporate activities and research and development. For example, in the fields of intellectual property (IP) landscapes and market research, attempts have been made to analyze vast amounts of text data such as patent gazettes and market reports using AI to extract technological trends and market trends.

[0003] Also, in the fields of business automation and decision-making support, the introduction of AI agents that autonomously infer and act to achieve given goals has been progressing. These technologies are expected to efficiently process large amounts of information that humans cannot handle, contributing to the derivation of new findings and the improvement of business efficiency.

[0004] Conventionally, in these systems, an approach of performing comprehensive analysis and learning by using as much of the collected information as possible as input data and learning data has been common.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Patent Document 2

Patent Document 3

Patent Document 4

Summary of the Invention

[0006] Conventional AI-based information processing systems have faced the challenge of unreliable output due to insufficient control over the quality and scope of input information.

[0007] For example, the information being analyzed may contain not only useful information for understanding trends, but also noise and misleading information. If these are processed indiscriminately, the AI ​​may produce erroneous analysis results influenced by noise, or generate information that is not based on facts (hallucination).

[0008] In particular, in autonomous AI agents, there is a risk that the internal state will deteriorate due to retraining on incorrect inference results, leading to a vicious cycle where errors self-replicate.

[0009] Therefore, this disclosure aims to provide an information processing system, an information processing method, and a program that can improve the accuracy and reliability of processing results when using AI. [Means for solving the problem]

[0010] An information processing system according to a first aspect of this disclosure is an information processing system using artificial intelligence, which evaluates information or areas within a target space that is the subject of processing by the artificial intelligence, and, based on the evaluation, performs processing to distinguish between information that should be preferentially used for processing by the artificial intelligence and information that should be excluded or suppressed from processing by the artificial intelligence within the target space.

[0011] An information processing method relating to a second aspect of this disclosure is an information processing method performed by an information processing system using artificial intelligence, which evaluates information or a region within a target space that is the subject of processing by the artificial intelligence, and, based on the evaluation, performs processing to distinguish between information that should be preferentially used for processing by the artificial intelligence and information that should be excluded or suppressed from processing by the artificial intelligence within the target space.

[0012] A program according to a third aspect of this disclosure causes an information processing system using artificial intelligence to perform the following actions: evaluate information or regions within a target space that is the subject of processing by the artificial intelligence, and, based on the evaluation, perform processing to distinguish between information that should be preferentially used for processing by the artificial intelligence and information that should be excluded or suppressed from processing by the artificial intelligence within the target space. [Effects of the Invention]

[0013] According to one aspect of this disclosure, it is possible to provide an information processing system, an information processing method, and a program that can improve the accuracy and reliability of processing results when using AI. [Brief explanation of the drawing]

[0014] [Figure 1] Figure 1 shows an example of the configuration of an information processing system according to the first embodiment. [Figure 2] Figure 2 shows an example of the configuration of a server device according to the first embodiment. [Figure 3] Figure 3 shows an example of the basic operation of a server device according to the first embodiment. [Figure 4] Figure 4 shows a first embodiment of the operation of the analysis unit according to the first embodiment. [Figure 5] Figure 5 shows a second embodiment of the operation of the analysis unit according to the first embodiment. [Figure 6] Figure 6 shows a third embodiment of the operation of the analysis unit according to the first embodiment. [Figure 7] Figure 7 shows a fourth embodiment of the operation of the analysis unit according to the first embodiment. [Figure 8] Figure 8 shows an example of the basic operation of the server device related to the modification example. [Figure 9] Figure 9 shows the mechanism of the vicious cycle of information in an AI agent according to the second embodiment. [Figure 10]FIG. 10 is a conceptual diagram showing the relationship between inputs, outputs, and mediating variables connecting them in an AI agent according to the second embodiment. [Figure 11] FIG. 11 is a functional block diagram showing a configuration example of a server device according to the second embodiment. [Figure 12] FIG. 12 is a diagram schematically showing the distribution of information in a target space as sand piles according to the second embodiment. [Figure 13] FIG. 13 is a diagram showing a state in which sand pile B is deleted and sand piles A and C remain according to the second embodiment. [Figure 14] FIG. 14 is a diagram showing a state in which sand pile C is deleted and only sand pile A, which is a candidate for the correct answer, remains according to the second embodiment. [Figure 15] FIG. 15 shows a state in which sharpening processing is performed on sand pile A according to the second embodiment. [Figure 16] FIG. 16 is a diagram showing an image in which boundary conditions are further set around sharpened sand pile A according to the second embodiment. [Figure 17] FIG. 17 is a flowchart showing an operation example of a server device according to the second embodiment.

Embodiments for Carrying Out the Invention

[0015] An information processing system according to an embodiment of the present disclosure will be described while referring to the drawings.

[0016] The information processing system according to the present disclosure is an information processing system using artificial intelligence (AI), which evaluates information or regions in a target space that is the target of AI processing, and based on the evaluation, distinguishes, in the target space, information to be preferentially used for AI processing and information to be excluded or suppressed from AI processing.

[0017] Here, "artificial intelligence (AI)" may mean, for example, a machine learning model, a deep learning model, a large-scale language model (LLM), or an autonomous agent using these. "Target space" may mean, for example, a dataset to be analyzed (patent information, market information, etc.), or a logical information space or behavioral space in which the AI ​​agent explores or learns. "Evaluation" may mean, for example, an analysis of the correlation between information and derived trends, or a confirmation of the consistency between information and known correct answer structures. "Distinguishing process" may mean, for example, a process of classifying information into significant and unsignificant items, or a process of deleting or sharpening specific areas from the search space.

[0018] According to the above configuration, when the AI ​​processes information or domains, it can evaluate the quality of the information or domain beforehand or dynamically, clearly distinguishing between useful information and harmful or useless information. This prevents noise information and inappropriate training data from negatively impacting the AI's processing, improving the accuracy and reliability of the AI's output (analysis results, inferences, actions, etc.).

[0019] The information processing system disclosed herein provides foundational technology to solve problems in AI-based processing such as "Garbage In, Garbage Out" (only poor quality output can be obtained from poor quality input) and, in its more severe form, "Garbage Everywhere" (self-propagation of misinformation). This system solves the above problems through two main approaches (first embodiment and second embodiment).

[0020] The first approach (corresponding to the first embodiment) involves "noise reduction" and "purification" when analyzing trends from large amounts of data. In IP landscape (registered trademark; hereinafter the same) and market research, AI collects all information (target space) that is considered relevant. Here, IP landscape can mean market research, customer trend research, technology and patent trend research, etc. However, this includes noise such as defensive applications unrelated to actual technology trends and stealth marketing articles that do not reflect the reality of the market. This system evaluates (analyzes) the correlation (weights of mediating variables, etc.) between the provisional trend information that has been derived and the individual input information. Then, it retrospectively distinguishes (classifies) "first information (information that should be used preferentially)" that truly contributed to the formation of trends and "second information (information that should be excluded or suppressed)" that has a low contribution. This makes it possible to perform in-depth analysis using only first information and to make highly accurate future predictions.

[0021] The second approach (corresponding to the second embodiment) involves "controlling the search space" when operating an autonomous AI agent. Here, the target space corresponds to the search range in which the AI ​​selects thoughts and actions. In an initial state without any control, the AI ​​will search a space that includes erroneous judgments and uncertain areas. This system evaluates the areas within the space based on known business rules (correct answer structure) and excludes (distinguishes) erroneous areas from the learning and search targets. Furthermore, areas that are found to be inappropriate as learning progresses are sequentially excluded, and the space is reshaped so that only correct answer candidates remain. By prioritizing inference within this reshaped space, hallucination and runaway behavior are prevented, and stable behavior that can withstand practical applications is achieved.

[0022] Thus, the technology disclosed herein manages and / or controls the quality of information handled by AI from both analysis and control perspectives, dramatically improving the practicality of AI systems.

[0023] (1) First Embodiment The first embodiment will be described with reference to Figures 1 to 8.

[0024] (1.1) Background technology and challenges In recent years, the importance of data-driven decision-making has increased in corporate management and R&D strategy formulation. In particular, in the field of technology development, an activity called IP landscaping (IPL), which accurately grasps the technological trends of one's own company and others and predicts future technological trends, is attracting attention. IP landscaping visualizes competitive trends, technological gaps, or promising technological areas in specific technological fields by collecting and analyzing vast amounts of technical information such as patent information and research paper information.

[0025] Furthermore, accurately understanding market trends is essential in the fields of product development and marketing. Market trend research involves collecting a wide variety of market information, such as product sales performance, customer reviews, and word-of-mouth on social networking services (SNS), to analyze consumer needs and changing trends.

[0026] These trend surveys are increasingly utilizing artificial intelligence (AI), particularly generative AI technologies such as large-scale language models (LLMs). Generative AI is expected to significantly improve the efficiency of analytical work by extracting highly relevant information from large amounts of text data and performing summarization and classification. For example, it is becoming possible to have AI summarize trends without having to manually read thousands or tens of thousands of patent documents or market reports.

[0027] However, conventional AI-based trend analysis methods had their drawbacks. Specifically, conventional analysis methods use all the collected information (target information) as input to derive trends, which makes the analysis results prone to noise.

[0028] For example, in IP landscapes, there is a large amount of information that is not necessarily important in indicating future technological trends, such as idea-level patents that have been filed but are not actually used in business, or defensive patents that are only intended to deter competitors from entering the market. If analysis is performed with this noisy information included, AI may draw incorrect conclusions (misleading results) that differ from the actual technological trends.

[0029] Similarly, in market trend surveys, some biased opinions or intentionally written posts aimed at sales promotion can act as noise, potentially distorting the true overall market trends. Since generative AI tends to be strongly influenced by information it has learned and frequently occurring information, it is difficult to ensure sufficient accuracy and reliability of analysis unless noise is properly separated from the input information.

[0030] Furthermore, there is a risk of hallucination, a phenomenon where artificial intelligence generates false or misleading information and presents it as fact.

[0031] This embodiment describes an information processing system capable of improving the accuracy and reliability of trend surveys and analyses.

[0032] (1.2) System Configuration First, the configuration of the information processing system 1 according to this embodiment will be described. Figure 1 is a diagram showing an example of the configuration of the information processing system 1 according to this embodiment.

[0033] The information processing system 1 according to this embodiment includes a server device 100 and a terminal device 200. The server device 100 and the terminal device 200 are connected to each other so as to be able to communicate with each other via a network 5 such as the Internet.

[0034] The server device 100 is an information processing device that performs information processing according to this embodiment. The server device 100 may be composed of a general-purpose computer such as a workstation or a personal computer, or it may be logically implemented by cloud computing.

[0035] The terminal device 200 is a computer used by the user, such as a personal computer, tablet, or smartphone. The terminal device 200 accesses the server device 100 via a web browser or dedicated application and provides a user interface for displaying analysis results and performing various operations. Users can access the server device 100 from the terminal device 200 via the network 5 and utilize the trend survey service described later.

[0036] Figure 2 shows an example of the configuration of the server device 100 according to this embodiment. The server device 100 according to this embodiment includes a communication unit 110, a storage unit 120, and a processing unit 130.

[0037] The communication unit 110, under the control of the processing unit 130, communicates with external devices such as terminal devices 200 via the network 5. The storage unit 120 includes storage media such as ROM, RAM, HDD, and SSD, and stores programs for executing the processing according to this embodiment, as well as various types of information used in the processing (target information, trend information, classification results, etc.).

[0038] In this embodiment, the storage unit 120 has storage areas such as a target information storage unit 121, a trend information storage unit 122, and a classification result storage unit 123. The target information storage unit 121 stores target information, which is the information subject to trend surveys. The trend information storage unit 122 stores trend information derived by the first derivation unit 132. The classification result storage unit 123 stores first information and second information classified by the first classification unit 134. Details of this information will be described later.

[0039] The processing unit 130 includes a processor such as a CPU and controls the operation of the entire server device 100 by executing a program stored in the storage unit 120. The processing unit 130 operates as various functional units such as the first acquisition unit 131, the first derivation unit 132, the first analysis unit 133, the first classification unit 134, the output unit 135, and the analysis unit 136. Some or all of the functions of the processing unit 130 may be implemented as an AI agent. Details of the AI ​​agent will be described in the second embodiment.

[0040] The first acquisition unit 131 acquires target information relating to the subject of the investigation. The first acquisition unit 131 may acquire target information from the target information storage unit 121. The first acquisition unit 131 may also acquire target information from an external database, website, etc., via the communication unit 110.

[0041] Here, "subject of investigation" refers to the subject of the investigation and analysis, and includes anything that the user wishes to understand the trends of, such as a specific technology field, a specific product, a specific market, or a specific company or organization. In this embodiment, we assume a technology trend investigation as a trend investigation, and will mainly explain an example where the subject of investigation is a specific technology field. For example, the trend investigation may be an activity called IP landscape (IPL) that accurately grasps the technology trends in a specific technology field and predicts future technology trends. However, the subject of investigation may be a specific company or organization. The subject of investigation may be a specific technology field within a specific company or organization.

[0042] "Target information" refers to a collection of information related to the subject of the investigation. For example, if the subject of the investigation is a specific technological field, then patent publications, papers, and technical news articles related to that field would be considered target information. For instance, IP Landscape collects and analyzes vast amounts of technical information, such as patent information and paper information, as target information to visualize competitive trends, technological gaps, or promising technological areas in a specific technological field. On the other hand, if the subject of the investigation is a product (or service), then the target information would include the specifications of that product (or service), sales data, pricing information, customer reviews, and mentions on social media.

[0043] The first derivation unit 132 derives trend information regarding the trends of the subject of the survey based on the target information acquired by the first acquisition unit 131. "Trend information" refers to information derived from the target information that shows the temporal changes, trends, and characteristic patterns of the subject of the survey. For example, future technological trends (such as research and development trends in the technological field), changes in market trends, and changes in the business strategies of specific companies fall under the category of trend information.

[0044] The first analysis unit 133 analyzes the correlation between the target information acquired by the first acquisition unit 131 and the trend information derived by the first derivation unit 132. "Correlation" refers to the relationship that shows how much each element (input variable) constituting the target information influences the derived trend information (target variable). This correlation includes not only statistical correlation coefficients but also broader concepts. For example, in machine learning models, it can be expressed as a weight (parameter) that indicates the degree to which certain input data (feature words, etc.) influences the output result, or as a transfer function or sensitivity that defines the relationship between input and output. Furthermore, in addition to regression analysis and sensitivity analysis, covariance structure analysis, which is a structural equation modeling method, can be used to analyze correlations.

[0045] The first classification unit 134, based on the results of the analysis by the first analysis unit 133, classifies the target information acquired by the first acquisition unit 131 into first information, which is significant for deriving trend information, and second information, which is less significant than the first information. "First information" refers to a group of information that has a large contribution to deriving trend information, i.e., "significant" information. For example, this includes a group of patents that are core in characterizing a certain technological trend, or customer reviews that greatly influenced product sales. "Second information" refers to a group of information that has a small contribution to deriving trend information, or is "less significant" information that becomes noise in the analysis. For example, this includes patents for defensive purposes that are not very relevant to actual business, or special claims due to individual circumstances.

[0046] Thus, in this embodiment, first, the first acquisition unit 131 acquires comprehensive target information regarding the subject of the investigation from an external database, website, etc., and stores it in the target information storage unit 121.

[0047] Next, the first derivation unit 132 derives trend information of the subject of the investigation based on this entire set of target information. In this derivation process, for example, a learning model such as a generative AI may be used to generate a summary (report) that summarizes trends and key points from a large amount of text data.

[0048] Next, the first analysis unit 133 analyzes the correlation between the input target information and the output trend information. This analysis identifies which target information strongly influenced the derivation of the trend information. For example, if a neural network model is used, it analyzes which paths from input nodes (corresponding to specific words or data) to output nodes (corresponding to specific trends) were activated, and what the weights and sensitivities were at that time. It is also possible to analyze the quantitative relationship between specific information elements and trends using statistical methods such as the Pearson correlation coefficient.

[0049] Furthermore, if we consider input nodes as explanatory variables, output nodes as the dependent variable, and the connections between these nodes as mediating variables, then once we learn a mediating variable with a high degree of influence or contribution, such as frequently occurring words or characteristic words (in other words, a high weight or sensitivity), even when we try to update it with new input, the influence of the previous mediating variable may remain, and the impact on the new population may not be reflected.

[0050] Furthermore, if maliciously generated output nodes or a large amount of false information are fed into the input nodes, the AI ​​has no choice but to learn from them, and once a result is output, the recipient tends to believe it.

[0051] Therefore, to prevent hallucination, it is effective to verify the mediating variables using a population that is thought to be free of noise. Furthermore, when comparing with a new population, the presence or absence of changes in the aforementioned mediating variables allows for accurate understanding and verification not only of the AI ​​results but also of the processing and how users perceive it.

[0052] Finally, the first classification unit 134 classifies information with a high contribution to the derivation of trend information as "first information" and information with a low contribution as "second information," based on the analysis results. This classification serves to separate "signals" and "noise" in the analysis. This classification clarifies the basis for why the trends were derived and prepares a high-purity dataset for subsequent more detailed analysis (in-depth analysis).

[0053] Therefore, according to this embodiment, it is possible to objectively classify the information used in trend surveys into "significant information" that truly contributed to the derivation of trends and "information of low significance (noise)" that does not. This enhances the reliability of the analysis results and helps users (analysts) accurately grasp essential trends without being misled by noise information.

[0054] In this embodiment, the output unit 135 may output at least one of the first information and second information classified by the first classification unit 134. For example, the output unit 135 has a function to present the results of analysis and classification in a format that the user can recognize via the communication unit 110. Specifically, the output unit 135 may display the results on the display of the terminal device 200, generate an analysis report (e.g., a PDF file or a CSV file), or transmit the data to another system.

[0055] The output unit 135 may present the results, classified into first and second information, in an easily understandable manner to the user. For example, it may display a list of target information and indicate whether each piece of information belongs to the first or second information using labels or color coding. It is also possible to extract and display only the first information or only the second information. This allows the user to grasp at a glance what kind of information was important in shaping the trend, and conversely, what kind of information was judged to be noise.

[0056] Therefore, according to this embodiment, users can clearly understand the basis of the analysis results. In particular, knowing what the important information (first information) was makes it possible to make judgments based on highly accurate information in future strategy formulation and decision-making.

[0057] Here, we will explain the difference between this method and analysis using causal inference (so-called causal AI), which has seen significant research in recent years. Unlike statistical analysis, causal inference seeks causal relationships (cause and effect), so caution is sometimes required when using machine learning. This is because the data used is premised on the link between cause and effect, so data that is not linked may become noise. Furthermore, particular care is required in handling confounding factors, covariates, and the aforementioned mediating variables (also called intermediate variables) that affect both causal and experimental aspects. While causal analysis, which takes such influences into account, is gaining prominence, it may require preparation such as experimental design, and there are also labor-intensive, physical, and technical difficulties in handling existing or acquired data. Moreover, in actual market analysis, causal inference that is thought to have little change is also necessary, but the mediating variables such as the aforementioned confounding factors and covariates often change, especially depending on the market environment and the measurement and analysis period, and have a significant impact on the results. For example, behavioral changes brought about by changes in how social media is used and the spread of AI technology can occur discontinuously and rapidly as trends, exceeding the predictions of causal inference. Furthermore, when analyzing causal inference over time, it is difficult to incorporate the aforementioned mediating variables into a model that excludes them over time, and there is a risk that the question of whether it is acceptable to ignore the influence of the mediating variables may become a point of contention. Even in such cases, it is possible to easily consider constructing a suitable model by comparing a method that also focuses on mediating variables, such as the present invention, with causal inference. For example, if the present invention is used to analyze that the influence of the mediating variables is high, or that the mediating variables change over time, it may be possible to reconsider the causal inference model.

[0058] In addition, the concept of counterfactuals in causal estimation is sometimes used as part of the estimation or verification. However, since the present application extracts populations with different learning results from the first and second information, it is also possible to perform a similar verification by comparing the two populations with different learning results.

[0059] In this embodiment, the target information is technical information that includes at least one of patent information and non-patent information. "Technical information" refers to information related to technology. Specifically, it includes "patent information" such as published patent gazettes and patent publications issued by the Japan Patent Office, and "non-patent information" such as academic papers, articles on technology news sites, and conference presentation materials. The first derivation unit 132 derives technology trend information as trend information, and the first analysis unit 133 analyzes the correlation between the technical information and the technology trend information. The first information is information that is significant for deriving technology trend information from the technical information. "Technology trend information" is, for example, information that shows the direction of development, trends, activities of major players, and the lineage of technological evolution in a particular technology field.

[0060] This embodiment is particularly intended for investigating technological trends in IP landscapes. The first acquisition unit 131 comprehensively acquires patent and non-patent information from a database related to a technological field specified by the user (e.g., "all-solid-state batteries"). The first derivation unit 132 derives technological trends in the field from this large amount of technological information (e.g., "The main development company is Company X, and recently it has shifted from electrolyte material development to control circuit development"). The information derived here may correspond to the "correct answer structure" (correct answer information) in the second embodiment described later. The first analysis unit 133 and the first classification unit 134 classify the patent groups and papers (first information) that were "significant" in deriving these technological trends, and the information that was not very relevant or was noise (second information). For example, "patent groups that are likely to be used in the future" would correspond to first information, and "mere idea patents, so-called applications that are unlikely to be used for protection purposes" would correspond to second information.

[0061] Therefore, by applying this embodiment to technology trend surveys, it is possible to efficiently identify truly important and noteworthy literature (first information) from a large volume of technical literature. This dramatically improves the accuracy of IP landscape analysis, enabling the interpretation of other companies' true strategies and the more accurate formulation of one's own research and development strategy.

[0062] In this embodiment, the first information may include at least one of the following: characteristic word information, engineer information, and technology classification information. "Characteristic word information" refers to information on keywords and technical terms that accurately express the technology content. "Engineer information" refers to information on persons involved in the technology development, such as inventors, researchers, and authors of papers. "Technology classification information" refers to code information or category information for systematically classifying technology, and includes, for example, patent classifications such as IPC (International Patent Classification), FI (File Index), and F-term (File Forming Term).

[0063] Thus, information significant to technological trends (primary information) can be understood not only as a single document, but also as "characteristics" commonly found in those documents. For example, if a technological trend is strongly associated with a specific "characteristic term" (e.g., "sulfide-based solid electrolytes"), a specific "engineer" (e.g., "Mr. B of Research Institute A"), or a specific "technological classification" (e.g., "H01M10 / 0562"), then this information itself constitutes part of the primary information. By clarifying these characteristics, we can gain a deeper understanding of the essence of the trend.

[0064] Therefore, according to this embodiment, it is possible not only to list significant literature, but also to extract common characteristics (keywords, key figures, core technological fields) from it. This makes it possible to grasp the specific content of technological trends from multiple perspectives and lead to more detailed analysis.

[0065] In this embodiment, the first derivation unit 132 may perform derivation using a learning model that takes target information as input and outputs trend information. A "learning model" refers to a machine learning or deep learning model that learns patterns and relationships from data and is constructed to make predictions, judgments, or generations for unknown data. This also includes large-scale language models (LLMs) such as ChatGPT® and Gemini®, which have attracted attention in recent years.

[0066] Thus, the first derivation unit 132 may utilize a learning model to efficiently derive trend information from a large amount of target information. For example, text data from tens of thousands of patent documents can be input into the learning model to generate a summary of trends in that technical field. This learning model may be a base model that has been pre-trained with general-purpose data, or it may be a model that has been further trained (fine-tuned) with data from a specific domain (e.g., patents or medicine).

[0067] By using learning models, it is possible to quickly and objectively derive trend information from large amounts of unstructured data (such as text) that would require enormous time and effort to process manually, thereby significantly improving the efficiency of the entire analysis process.

[0068] The analysis unit 136 performs an analysis of the trends of the subject of the investigation based on the first information classified by the first classification unit 134. The analysis unit 136 includes, for example, a second acquisition unit 136a, a second derivation unit 136b, a second analysis unit 136c, a second classification unit 136d, an extraction unit 136e, and a comparison unit 136f.

[0069] For example, the analysis unit 136 may have a function to perform a more advanced, detailed analysis (in-depth analysis) by prioritizing the use of the first information, which has been judged as "significant" by the first classification unit 134, over the second information. After the second information, which is noise, is removed and the first information, which is of high purity, is obtained, the analysis unit 136 may prioritize this selected first information as the target of analysis to obtain more essential and reliable analysis results. This is the "in-depth" phase, and it is expected that detailed causal relationships and patterns that were not visible in the initial analysis will be discovered. For example, by performing a detailed analysis after completely eliminating the influence of noise information, the quality of the analysis can be greatly improved, such as identifying the true factors behind trends or making more accurate future predictions.

[0070] As a first embodiment of the operation of the analysis unit 136, the second derivation unit 136b may prioritize the first information over the second information to derive trend information regarding the trends of the subject under investigation, and output the derived trend information (for example, notify the user via the terminal device 200). For example, the second derivation unit 136b may derive trend information regarding the trends of the subject under investigation based on the first information, without relying on the second information. This is an example of "in-depth analysis" by the analysis unit 136. Here, the second derivation unit 136b derives trend information again, using only the first information as input. While the first derivation unit 132 derived trends from the entire target information (first information + second information), the second derivation unit 136b derives trends from only the selected first information. This resets the learning bias from the noise information that the first analysis model had, and re-evaluates it using only pure information.

[0071] This is expected to improve the accuracy of trend information. Because it is derived from a clean dataset with noise removed, it is possible to obtain trend information that is sharper and captures the essence better than the trend information generated by the first derivation unit 132.

[0072] As a second embodiment of the operation of the analysis unit 136, the second acquisition unit 136a may acquire new target information relating to the subject of investigation, and the extraction unit 136e may extract information corresponding to the first information from the new target information and output the extracted information (for example, notifying the user via the terminal device 200). This is another function of the analysis unit 136 and provides a "fixed-point observation" function of information. This corresponds to the SDI (Selective Dissemination of Information) function in patent searches. First, features (for example, specific keywords, technology classifications, company names, etc.) included in the first information (significant information) identified in the first step are set as a search profile and a prompt to indicate. Then, when the second acquisition unit 136a acquires new target information added daily (e.g., newly published patent gazettes), the extraction unit 136e filters only the information that matches the search profile in real time and notifies the user.

[0073] This allows for efficient and continuous monitoring of the latest information related to key trends once they have been identified. This enables early detection of signs of significant change and prompt response.

[0074] As a third embodiment of the operation of the analysis unit 136, the second acquisition unit 136a acquires new target information relating to the subject of investigation, the second derivation unit 136b derives new trend information relating to the trends of the subject of investigation based on the new target information, the second analysis unit 136c analyzes the new correlation between the new target information and the new trend information, and the comparison unit 136f performs at least one of the following: a comparison between the trend information and the new trend information, and a comparison between the correlation and the new correlation, and outputs information based on the comparison results (for example, notifying the user via the terminal device 200).

[0075] In this embodiment, the analysis unit 136 provides a function to perform "time-series comparison". For example, the comparison unit 136f compares an analysis performed one year ago (generating past trend information and past correlations) with an analysis performed with the latest data (generating new trend information and new correlations). For example, it may analyze not only how last year's technological trends and this year's technological trends have changed (comparison of trend information), but also what caused those changes, that is, how the parameters that have become more important between last year and this year (weights of characteristic words, etc.) have changed (comparison of correlations).

[0076] This allows us to not only capture the "change" in trends themselves, but also trace back and analyze the "factors" behind those changes. This enables us to understand the underlying structural changes beyond superficial shifts, leading to more accurate future predictions and concrete strategic adjustments.

[0077] As a fourth embodiment of the operation of the analysis unit 136, the second acquisition unit 136a acquires new target information relating to the subject of investigation, the second derivation unit 136b derives new trend information relating to the trends of the subject of investigation based on the new target information, the second analysis unit 136c analyzes the new correlation between the new target information and the new trend information, the second classification unit 136d classifies the new target information into new first information that is significant for deriving new trend information and new second information that is less significant than the new first information based on the results of the analysis, the comparison unit 136f compares the first and second information with the new first information and new second information, and outputs information based on the comparison results (for example, notifying the user via the terminal device 200).

[0078] This embodiment also performs a "time-series comparison," but the objects of comparison differ from those in the third embodiment. Here, the comparison unit 136f compares past "significant / non-significant" classification results (first information, second information) with current "significant / non-significant" classification results (new first information, new second information). For example, if a technology that was classified as "low significance (second information)" last year is classified as "significant (new first information)" this year, it can be said that this is an extremely important signal indicating that the importance of that technology has increased.

[0079] This allows us to directly capture not only changes in trends themselves, but also changes in the very value system of "what is important." Therefore, it becomes possible to detect more fundamental and significant changes, such as paradigm shifts in industries or strategic shifts in companies, at an early stage and take proactive measures.

[0080] (1.3) Example of server device operation Next, referring to the flowcharts in Figures 3 to 8, an example of the operation of the server device 100 according to this embodiment will be described. Here, the operation will be described when the trend survey is a technology trend survey (IP landscape).

[0081] (1.3.1) Basic Operation Examples Figure 3 shows an example of the basic operation of the server device 100 according to this embodiment. This example of operation is for investigating and analyzing technological trends in a specific technical field.

[0082] In step S1, the first acquisition unit 131 acquires target information related to the subject of investigation. In this example, for example, when a user operates the terminal device 200 and specifies a specific technical field (e.g., "LiDAR technology in autonomous driving") as the subject of investigation, the first acquisition unit 131 of the server device 100 comprehensively acquires technical information (patent publications, papers, etc.) related to the specified technical field from external patent databases and academic literature databases.

[0083] In step S2, the first derivation unit 132 derives trend information regarding the trends of the subject of investigation based on the target information acquired by the first acquisition unit 131. For example, the first derivation unit 132 takes all the acquired technical information as input and uses a learning model (e.g., a generative AI) to derive technical trend information. For example, trend information is generated that states, "In recent years, there has been an increase in applications related to MEMS-type LiDAR, with companies A and B being the main players."

[0084] In step S3, the first analysis unit 133 analyzes the correlation between the target information acquired by the first acquisition unit 131 and the trend information derived by the first derivation unit 132. For example, the first analysis unit 133 analyzes the extent to which each piece of input technical information contributed to the derivation of the trend information and clarifies the correlation.

[0085] In step S4, the first classification unit 134 classifies the target information acquired by the first acquisition unit 131 into first information, which is significant for deriving trend information, and second information, which is less significant than the first information, based on the results of the analysis by the first analysis unit 133. In other words, the first classification unit 134 classifies the technical information based on the analysis results. For example, a group of patents with a high contribution that mention "MEMS method" or "Company A" is classified as "first information," while patents of other methods with low relevance or patents considered to be noise are classified as "second information."

[0086] In step S5, the output unit 135 may output at least one of the first information and second information classified by the first classification unit 134. The output unit 135 may also generate a report containing a list of the classified first and second information, as well as common features of the first information (technical classification, keywords, etc.), and display it on the user's terminal device 200.

[0087] (1.3.2) Example of operation of the analysis unit First to fourth embodiments of the operation of the analysis unit 136 will be described. The first to fourth embodiments may be carried out separately and independently, or two or more embodiments may be carried out by combining at least partially.

[0088] (1.3.2.1) First embodiment of the operation of the analysis unit Figure 4 shows a first embodiment of the operation of the analysis unit 136.

[0089] In step S11, the second derivation unit 136b prioritizes using the first information over the second information classified by the first classification unit 134 to derive trend information regarding trends in a specific technological field. For example, the second derivation unit 136b may acquire only the first information (significant patent groups) classified by the first classification unit 134 and derive the technological trends again based on this clean dataset from which noise has been removed.

[0090] In step S12, the second derivation unit 136b outputs the derived trend information (for example, notifying the user via the terminal device 200).

[0091] This allows us to obtain more accurate and essential trend information than the initial analysis.

[0092] (1.3.2.2) Second embodiment of the operation of the analysis unit Figure 5 shows a second embodiment of the operation of the analysis unit 136.

[0093] In step S21, the second acquisition unit 136a acquires new target information relating to a specific technical field. For example, the second acquisition unit 136a acquires new technical information that is published on a daily basis as new target information relating to a specific technical field.

[0094] In step S22, the extraction unit 136e extracts information corresponding to the first information from the new target information. For example, the extraction unit 136e sets the characteristics of the first information (e.g., technical classification "G01S 7 / 481", keyword "solid state") as a search profile or prompt, and extracts information that matches this profile from the new technical information.

[0095] In step S23, the extraction unit 136e outputs the extracted information (for example, notifying the user via the terminal device 200).

[0096] This allows for efficient and continuous monitoring of the latest information related to key trends once they have been identified. This enables early detection of signs of significant change and prompt response.

[0097] (1.3.2.3) Third embodiment of the operation of the analysis unit Figure 6 shows a third embodiment of the operation of the analysis unit 136.

[0098] In step S31, the second acquisition unit 136a acquires new target information relating to a specific technical field.

[0099] In step S32, the second derivation unit 136b derives new trend information regarding trends in a specific technological field based on the new target information. Here, the second derivation unit 136b may derive the technology trend information using a learning model (e.g., generative AI).

[0100] In step S33, the second analysis unit 136c analyzes the new correlation between the new target information and the new trend information.

[0101] In step S34, the comparison unit 136f performs at least one of the following: a comparison between the past trend information derived in step S2 and the new trend information derived in step S32, and a comparison between the past correlation derived in step S3 and the new correlation derived in step S33. For example, the memory unit 120 stores the results of analysis using past data (e.g., one year ago) (past trend information, past correlations, past first / second information), and uses the stored information for comparison.

[0102] In step S35, the comparison unit 136f outputs information based on the comparison results (for example, notifying the user via the terminal device 200). This allows for the identification of information such as "how the trends have changed and what the causes are," and the output of the results.

[0103] (1.3.2.4) Fourth embodiment of the operation of the analysis unit Figure 7 shows a fourth embodiment of the operation of the analysis unit 136.

[0104] In step S41, the second acquisition unit 136a acquires new target information relating to a specific technical field.

[0105] In step S42, the second derivation unit 136b derives new trend information regarding trends in a specific technological field based on the new target information. Here, the second derivation unit 136b may derive the technology trend information using a learning model (e.g., generative AI).

[0106] In step S43, the second analysis unit 136c analyzes the new correlation between the new target information and the new trend information.

[0107] In step S44, the second classification unit 136d classifies the new target information into new first information that is significant in order to derive new trend information, and new second information that is less significant than the new first information, based on the results of the analysis in step S43.

[0108] In step S46, the comparison unit 136f outputs information based on the comparison result (for example, notifying the user via the terminal device 200).

[0109] Thus, in this embodiment, the classification results of "first information / second information" from the past and present are compared (S45). This allows for a direct understanding of, for example, "how the content of information considered important has changed," and the results are output (S46). This enables users to keenly perceive turning points in their strategies, etc.

[0110] (1.4) Example of changes The above-described embodiment mainly explained the case where the trend survey is a technology trend survey (IP landscape). In contrast, this modified example mainly explains the case where the trend survey is a market trend survey. In the configuration and operation of the above-described embodiment, "technical information" and "technology trend information" can be read as "market information" and "market trend information," respectively.

[0111] Specifically, in this modified example, the target information regarding the subject of the survey is market information that includes at least one of product information and service information. The first derivation unit 132 derives market trend information as trend information, and the first analysis unit 133 analyzes the correlation between market information and market trend information. In this modified example, the first information is information that is significant for deriving market trend information from the market information.

[0112] Here, "market information" refers to information about goods and services in the market. "Product information" includes product specifications, prices, sales volume, etc. "Service information" includes service details, fees, number of contracts, etc. In addition to these, posts on social media, reviews on e-commerce sites, news articles, blog posts, etc., may also be included in market information. "Market trend information" refers to information that shows trends in a particular market, changes in consumer needs, trends of competing products, price fluctuation trends, etc.

[0113] Thus, this modified example applies the above-described embodiment to market trend research. For example, assuming the apparel industry, the first acquisition unit 131 acquires market information (e.g., sales data from e-commerce sites, posts by influencers on social media, reviews from general users, etc.) about a specific product (e.g., sneakers of a specific brand). The first derivation unit 132 derives market trends (e.g., "retro designs that are also comfortable to walk in are trending") from this information. Then, the first classification unit 134 classifies the information that was significant in deriving this trend (e.g., posts by specific influencers, highly-rated reviews that directly led to sales, etc.) as first information, and the information that was not significant (e.g., posts that did not affect sales, outlier low-rated reviews, etc.) as second information.

[0114] This allows us to identify the factors (primary information) that truly influence consumer purchasing behavior and trends. This enables us to develop data-driven, effective product development and marketing strategies, rather than relying on intuition or experience.

[0115] In this example of modification, the first piece of information may include at least one of the following: characteristic word information, information provider information, information medium information, product classification information, and service classification information. Here, "information provider" refers to the source of market information, and includes, for example, influencers and general users who post on social media, reviewers who write on review sites, and journalists who write news articles. "Information medium" refers to the media on which market information is published, and includes social media such as X (formerly Twitter) (registered trademark) and Instagram (registered trademark), specific e-commerce sites, blogs, and news sites. "Product classification" and "service classification" refer to information used to categorize products and services.

[0116] This allows us to identify information that is significant to market trends (primary information) as common characteristics. For example, if market trends are strongly associated with a specific "keyword" (e.g., "Y2K fashion"), a specific "information provider" (e.g., "influencer C"), a specific "information medium" (e.g., "short videos on video-sharing SNS"), or a specific "product category" (e.g., "platform sneakers"), then this information constitutes part of the primary information.

[0117] Furthermore, even when a large amount of fake or false information is disseminated in "short videos on video-sharing SNS," comparing the first and second pieces of information classified as described above, and examining the mediating variables, makes it possible not only to grasp meaningful information without being misled by such noise, but also to avoid hallucination.

[0118] Therefore, this example of modification allows for the multifaceted identification of specific factors shaping market trends. This provides insights that lead to concrete and actionable actions, such as which keywords to focus on, which influencers and media outlets are important, and which product categories to concentrate on.

[0119] For example, if the aforementioned "Y2K fashion," "influencer C," and "platform sneakers" are strongly related, it becomes possible to examine in detail which characteristic term contributes most, and in what order the influence spread, by analyzing the correlation between the characteristic terms (mediating variables). Furthermore, by understanding the interactions between the aforementioned mediating variables and their relationships over time, advanced insights can be obtained. Specifically, it becomes possible to analyze even specific market movements, such as whether "influencer C" has the greatest influence, or whether it is "Y2K fashion" or "platform sneakers," or whether the conditions for the interaction between the three are "and" or "or," or whether there are changes in the interactions over time or changes in the mediating variables themselves.

[0120] For example, the first analysis unit 133 may query the learning model (e.g., generative AI) used to derive market trends for the degree of influence, interactions, time series analysis, etc. Alternatively, the first analysis unit 133 may use statistical methods to perform partial correlation analysis or time series analysis between components.

[0121] Figure 8 shows an example of the basic operation of the server device 100 related to this modification example. Here, we assume a case where market trends for a specific product category are investigated and analyzed.

[0122] In step S1a, the first acquisition unit 131 acquires target information related to the subject of the survey. In this modified example, for example, if the user specifies a specific product category (e.g., "women's sneakers") as the subject of the survey, the first acquisition unit 131 acquires relevant market information (product reviews, influencer posts, sales rankings, etc.) from e-commerce sites, social media, fashion blogs, etc.

[0123] In step S2a, the first derivation unit 132 derives trend information regarding the trends of the subject of the investigation based on the target information acquired by the first acquisition unit 131. For example, the first derivation unit 132 derives market trend information (e.g., "Platform shoes with light colors are gaining popularity, and sales have surged following a post by influencer C") from the acquired market information.

[0124] In step S3a, the first analysis unit 133 analyzes the correlation between the target information acquired by the first acquisition unit 131 and the trend information derived by the first derivation unit 132. For example, the first analysis unit 133 analyzes which reviews or posts strongly influenced the derivation of this market trend and clarifies the correlation.

[0125] In step S4a, the first classification unit 134, based on the results of the analysis by the first analysis unit 133, classifies the target information acquired by the first acquisition unit 131 into first information, which is significant for deriving trend information, and second information, which is less significant than the first information. For example, the first classification unit 134 classifies reviews containing characteristic words such as "platform sole" and "pale color," and posts by "influencer C," etc., as "first information," and classifies other information that did not have much impact on sales as "second information."

[0126] In step S5a, the output unit 135 may output at least one of the first information and second information classified by the first classification unit 134. By presenting the classification results to the user, the output unit 135 can support the planning of the next product and the formulation of promotion strategies.

[0127] (2) Second Embodiment The second embodiment will be described with reference to Figures 9 to 17.

[0128] (2.1) Background technology and challenges In recent years, with the technological advancement of large-scale language models (LLMs), the practical application of AI agents capable of autonomously performing multiple tasks (Agentic AI) has been rapidly progressing. An AI agent is defined as a system that perceives its environment, performs reasoning, and autonomously uses tools to act in order to achieve its goals. While conventional chatbots and standalone large-scale language models are "knowledge retrieval and generation systems" that passively generate answers and summaries based on knowledge within training data in response to user prompts, AI agents are clearly distinguished in that they are "autonomous execution systems" that actively engage in trial and error with the aim of completing tasks and solving problems.

[0129] Specifically, AI agents generally operate by repeating a cycle of Perception, Planning, Action, and Reflection. In "Perception," they understand the user's instructions and the current situation, and in "Planning," they infer the processes necessary to achieve the goal. This inference function is the core of the AI ​​agent and not only outputs knowledge, but also includes a "task decomposition" function that breaks down complex instructions (e.g., "create and book a travel plan") into actionable smaller tasks such as "check the schedule," "search for flights," "compare hotels," and "execute the reservation," as well as "logical thinking" that determines the next move based on the results of a previous action, and even a "self-correction" function that changes search terms or rewrites the code and retries when an error occurs.

[0130] Furthermore, "learning" in AI agents does not necessarily mean real-time updating of the model's weight parameters (learning in the narrow sense). In general operation, even if the model itself remains fixed, it is possible to accumulate experience and behave intelligently over time by using "in-context learning," which temporarily maintains rules and user preferences through the history of dialogue within the context, or by accumulating past behavior logs and success patterns in external memory such as a vector database and referencing them during subsequent task executions using RAG (Retrieval-Augmented Generation) technology. In other words, an AI agent is a system that obtains a substantial learning effect as its external memory increases.

[0131] From an implementation technology perspective, "ReAct prompting," which involves running a Thought, Action, and Observation loop for the LLM, is widely used as a fundamental principle. Furthermore, in terms of architecture, in addition to a single-agent configuration where a single LLM handles all processes, a multi-agent system has also been put into practical use in which multiple specialized agents with different roles, such as planning (PM), coding (development), and test execution (tester), cooperate to perform complex tasks. This is expected to lead to increased efficiency in advanced business automation such as web research, coding, and PC operation assistance, as well as in decision support.

[0132] However, in the implementation of AI agents, while certain results are achieved during the proof-of-concept (PoC) phase, there are many cases where the transition to full-scale operation stalls or fails. Specifically, AI agents that function appropriately under limited conditions may become overly sensitive to the diversity of input data and changes in the environment as the scale of operation expands and they are applied to actual business operations, leading to a decrease in the stability and reliability of their behavior. Once an incorrect judgment or action occurs, the cost of manual monitoring and correction increases, ultimately undermining the labor-saving benefits of introducing AI agents.

[0133] Traditionally, the causes of such problems have been attributed to individual technical factors such as insufficient training data, performance limitations due to the number of model parameters, or inadequate prompt engineering. However, even in environments with equivalent technical infrastructure and data quality, stable operation can be achieved in some cases but not in others. Therefore, the fundamental cause is thought to lie not simply in model performance, but in the structural design of the system.

[0134] In this embodiment, this problem is viewed as a "lack of control" issue stemming from the fact that the AI ​​agent is not designed as a "controlled object." Unlike conventional generative AI (stateless systems) that generate one-off responses, AI agents often have a cyclical structure in which they use their own output as input for subsequent reasoning or actions, and retain the results as memory, as described above. In other words, AI agents have aspects of a dynamic system (stateful system) in which their internal state transitions over time.

[0135] In such dynamic systems, if operation begins without a clear control design—including what should be learned, what should not be learned, and the scope of exploration allowed—the system's behavior is likely to become unstable. Because AI agents inherently lack the social norms and contextual understanding abilities that humans possess, they may learn and reinforce incorrect judgments that succeed by chance, or answers that are logically consistent but inappropriate for the business context, as positive rewards.

[0136] Figure 9 illustrates the mechanism of the vicious cycle of information in AI agents. Referring to Figure 9, if an AI agent is allowed to learn without limiting its learning space when the "correct structure" such as correct decision criteria, unacceptable exceptions, and conditions for stopping a decision in a task is not defined (START), the AI ​​agent will perform a disorderly search.

[0137] In this process, biases may occur in the internal parameters, or "parametric variables," that determine the behavior of the AI ​​agent. Figure 10 is a conceptual diagram showing the relationship between the input (INPUT), output (OUTPUT), and the parametric variables connecting them in an AI agent. As shown in Figure 10, a multi-layered network of parametric variables, composed of weights and activation states, exists between the input (explanatory variables) and the output (dependent variable). The AI ​​agent's inference is performed via paths on this network, but if proper control is not in place, a path that accidentally generates an appropriate output for specific data in the PoC stage (for example, a path that differs from the original correct solution structure, shown as a thin line in Figure 10, or a path that overfits a specific keyword) may be reinforced as an effective path. This means convergence to a local optimum.

[0138] Referring again to Figure 9, the bias in the parameters described above can lead to bias and misleading results in the inference process. Furthermore, due to the cyclical structure unique to AI agents, these inappropriate inference results are fed back to the input side as training data or preconditions for the next inference.

[0139] The conventional "Garbage In, Garbage Out (GIGO)" principle dictates that the quality of input data determines the quality of output. However, in AI agents, the re-input of low-quality output can lead to self-propagation of errors, resulting in a decline in the overall quality of data and decision criteria within the system. In this specification, this state is referred to as "Garbage Everywhere." Once this state is reached, even if attempts are made to improve the quality of input data retrospectively, it becomes difficult to obtain appropriate output because the parameter variables that serve as the AI ​​agent's decision criteria are already fixed in an inappropriate state.

[0140] As a result, the AI ​​agent begins to make incorrect judgments with high confidence, making external corrections (such as additional data training or prompt adjustments) less effective. As shown in Figure 9, simply adding data at this stage may worsen the situation, as additional training is performed on top of a biased structure, potentially reinforcing the tendency to make incorrect judgments.

[0141] Furthermore, the anthropomorphic concept of "training" is sometimes used when introducing AI agents, but this can also contribute to a lack of control design. Humans can learn norms and tacit knowledge through failure and make autonomous corrections, but AI agents are mathematical models that optimize based on designed evaluation functions, etc., and there is no guarantee that they will autonomously correct themselves in a desirable direction. Uncontrolled learning carries the risk of causing the system to change in an unintended direction.

[0142] Furthermore, this challenge becomes even more pronounced in areas where the correct answer is unknown or fluid, such as exploring new business opportunities. While AI agents can generate highly probable options based on existing data, they struggle to consider practical constraints such as cost, legal regulations, and feasibility, as well as non-data-driven factors like human emotional responses. As a result, there is a tendency for AI agents to lean towards either options that are extensions of existing businesses (overfitting) or options that are unlikely to be feasible (hallucination) (the inverse smile curve phenomenon), which hinders effective exploration.

[0143] As described above, simply improving the performance of the model is insufficient for applying AI agents to continuous business operations and advanced search tasks. It is necessary to reconsider the AI ​​agent as a "controlled object" that undergoes state transitions, clearly separate the learning process and the inference process, and develop new control methods to appropriately design and shape the search space.

[0144] (2.2) Operation overview The following describes the operation overview of the information processing system 1 according to this embodiment.

[0145] Figure 11 is a functional block diagram showing an example configuration of the server device 100a according to the second embodiment. The server device 100a has an architecture optimized for strictly managing the autonomous behavior of AI agents as an object of engineering "control," rather than using vague concepts such as the conventional "training."

[0146] As shown in Figure 11, the server device 100a comprises a communication unit 210, a storage unit 220, and a processing unit 230. The processing unit 230 functions as an AI agent processing unit 231, a learning control unit 232, and an inference control unit 233 by executing a program. While each of these units operates autonomously, they work closely together to ensure the quality of the AI ​​agent.

[0147] The AI ​​agent processing unit 231 is an execution entity that uses a large-scale language model (LLM) or the like as its core engine to perform inference, judgment, and action generation for input tasks. An important point in this embodiment is that the AI ​​agent processing unit 231 does not operate without constraints, but is under strict management (control) by the learning control unit 232 and the inference control unit 233. The AI ​​agent processing unit 231 is not merely a text generator, but has an agent function that autonomously calls external tools (search engines, databases, computers, etc.) to perform tasks, but its "use of tools" and "criteria for judgment" are limited to the range defined by the control unit.

[0148] The memory unit 220 has an internal state memory unit 221. The internal state memory unit 221 is a database that stores the "internal state" that determines the behavior of the AI ​​agent processing unit 231. This includes not only the weight parameters of the neural network, but also "mediating variables" specific to this embodiment. Mediating variables are a group of parameters that intervene between input and output and define the AI ​​agent's judgment tendencies, value criteria (what is considered important), search sensitivity (how deeply to investigate), and memory retention (what is retained in long-term memory). The internal state memory unit 221 also holds information on the "correct answer structure (anchors as known correct answers)" described later, and the "big paths (main roads of thought)" that should be preferentially traversed in the search space, forming the identity and judgment axis of the AI ​​agent.

[0149] The learning control unit 232 is the command center that oversees the AI ​​agent's "learning phase." Its role is not simply to load data. The learning control unit 232 is responsible for defining and shaping the "target space (search space)" that the AI ​​agent should explore. Specifically, it removes inappropriate regions from the search space and sharpens the correct answer candidate regions based on on-site work rules and constraints. It also functions as a gatekeeper to prevent data generated by the AI ​​agent from self-replicating and mixing with the learning data, and updates the parameter variables in the internal state memory unit 221 using only verified, highly reliable data. In other words, the learning control unit 232 controls "what the AI ​​agent learns and what it doesn't learn" at the design level, and manages the direction of the AI's growth.

[0150] The inference control unit 233 is a supervisor that oversees the "inference phase" of the AI ​​agent. Based on the search space shaped and sharpened by the learning control unit 232, it causes the AI ​​agent processing unit 231 to perform practical inference. The inference control unit 233 performs fixed control (freeze) to prevent the internal state (parameters) from being unintentionally changed during inference, and monitors in real time to ensure that the output of the AI ​​agent does not deviate from the pre-set boundary conditions (guardrails). If a deviation is detected, it either cuts off the output or executes a fail-safe function to correct it to a safe response, prioritizing safety in actual operation. Alternatively, it controls the target space as defined by the boundary conditions.

[0151] Thus, the server device 100a is configured to ensure stability, predictability, and reliability in the long-term operation of the AI ​​agent by clearly separating the learning and inference processes functionally and temporally, and by performing appropriate control interventions in each phase.

[0152] (2.2.1) First aspect In the first aspect of this embodiment, the learning control unit 232 defines and shapes the target space by performing evaluation (evaluating information or regions within the target space that are the target of processing by the AI ​​agent) and distinction processing (distinguishing between information that should be preferentially used for processing by the AI ​​agent and information that should be excluded or suppressed from processing by the AI ​​agent) in the learning phase, and the inference control unit 233 causes the AI ​​agent to perform inference processing based on the shaped target space in the inference phase.

[0153] Here, "AI agent" may mean, for example, a program or module based on a large-scale language model (LLM) that autonomously performs task decomposition, information gathering, judgment, and action generation in response to a given goal. "Learning phase" may mean, for example, a period or process in which the internal state (parametric variables, knowledge base, etc.) of the AI ​​agent is updated, built, or adjusted. "Inference phase" may mean, for example, a period or process in which the learned internal state is used to generate output (answer, action, etc.) for a specific input. "Target space" may mean, for example, a logical space composed of a collection of information, knowledge, action candidates, or parameters that the AI ​​agent is targeting for exploration, reference, or learning, and may be rephrased as "search space" or "search target space," etc.

[0154] According to the first embodiment, the operation of an AI agent clearly separates the "learning phase" and the "inference phase," allowing for control tailored to the specific purpose of each phase (designing the search space and executing within that space). This fundamentally prevents behavioral instability caused by the AI ​​agent haphazardly exploring a sea of ​​unorganized information, as well as performance degradation due to the accumulation of incorrect learning. In other words, by managing the AI ​​agent with an engineering-based and reproducible approach of "controlling" it, rather than the uncertain approach of "nurturing" it, it becomes possible to realize a system that can operate safely and continuously in practical settings, going beyond the PoC (proof of concept) level. Furthermore, separating learning and inference makes it easier to ensure traceability (trackability) regarding the learning state at which inference was performed.

[0155] Furthermore, in the introduction of agents using conventional generative AI, the boundary between learning and inference was often blurred at the start of operation, which was a major cause of failure. The expectation that the AI ​​would naturally become smarter with use is often disappointed in uncontrolled environments. In reality, there is a high risk of falling into a "Garbage Everywhere" state where incorrect output becomes the next input, and noise information self-propagates as it is learned. In particular, in the case of agents that use tools autonomously, if incorrect tool usage or inappropriate data referencing is learned, the entire system may malfunction.

[0156] In this embodiment, the information processing system 1 addresses this structural problem by first defining and shaping the target space, which is the "arena" in which the AI ​​agent should operate, through the learning control unit 232. This is equivalent to providing the AI ​​agent with an accurate "map" and "compass," and preemptively eliminating dangerous areas that should not be entered (for example, areas where there is a risk of leakage of personal information or areas where unfounded hallucinations are likely to occur) and unproductive exploration areas. Then, the inference control unit 233 allows the AI ​​agent to behave freely only within this prepared, safe, and high-density space. This configuration eliminates instability and drift phenomena (deterioration of accuracy over time) during long-term operation that were not visible under the limited conditions of the PoC stage, and elevates the AI ​​agent from a mere experimental tool to a "controllable social implementation system" that can be entrusted with responsible tasks.

[0157] (2.2.2) Second aspect In a second aspect of this embodiment, the learning control unit 232 may set a known correct answer structure that indicates a valid judgment criterion or constraint as an initial condition, and reconfigure the target space to exclude information or regions in the target space that do not satisfy the initial condition as regions that the AI ​​agent should not explore.

[0158] Here, "known correct answer structure" may mean, for example, rules already established in business operations, mandatory procedures, prohibitions that must be strictly followed, or reliable decision logic derived from past successful cases. "Initial conditions" may mean, for example, fixed constraints or reference points (anchors) that are initially given when the AI ​​agent begins learning or exploring.

[0159] According to the second aspect, before the AI ​​agent begins full-scale learning and exploration, it is possible to structurally eliminate exploration in clearly wrong directions and judgments that are unacceptable in business contexts. This prevents the "initial value sensitivity trap," where the AI ​​agent prematurely converges to a local minimum (described later as "Sandpile B") that appears plausible but is actually incorrect in practice, and allows it to maintain the correct direction from the early stages of learning. As a result, learning efficiency is dramatically improved, and a decrease in reliability in the initial stages of operation can be avoided. In addition, since the exploration can start from a position close to the "correct answer" from the beginning, the time to convergence is shortened, and it also contributes to saving computational resources.

[0160] Figure 12 schematically represents the distribution of information in the target space as a "sand pile." In Figure 12, "A" in the center represents the region of the correct candidate that should ultimately be reached, "B" on the left represents the region that appears to be the correct answer but is actually misleading, and "C" on the right represents the region that is found to be unsuitable after the search.

[0161] The processing performed by the learning control unit 232 in this embodiment corresponds to the process of deleting "pile B" in Figure 12 before the start of the search (first-order pruning). Pile B is close to the initial position (current knowledge) and is likely to receive a high score in the short-term evaluation function, making it appear as an attractive "shortcut" to the AI ​​agent. For example, this includes errors such as presenting a discontinued bus route as the "shortest route" or making inappropriate decisions based on outdated manuals. Because these appear logically consistent, it is difficult for the AI ​​to notice the errors on its own, and once learned, they are difficult to correct.

[0162] The learning control unit 232 physically separates this sandpile B from the search space by setting the correct structure of the field (e.g., the latest train schedule, legal regulations, safety standards, company regulations) as an anchor. Technically, possible methods include setting the learning weight to zero for data that contradicts the correct structure, or providing a strong negative reward. Figure 13 shows the state after sandpile B has been removed, leaving sandpile A and sandpile C. In this way, by clearly indicating the "mountains that should not be climbed" in the initial stages and reconstructing the search space, it is possible to prevent the AI ​​agent from getting lost and concentrate resources on promising areas.

[0163] (2.2.3) Third aspect In a third aspect of this embodiment, the learning control unit 232 may set at least one of the evaluation function, reward, sensitivity, and update intensity in the learning process as a parameter that defines the judgment tendency or memory retention degree of the AI ​​agent with respect to the reconstructed target space (for example, the symmetric space from which the sand pile B has been removed), and cause the AI ​​agent to execute the learning process based on said setting.

[0164] Here, "mediating variables" may refer to, for example, a group of internal parameters that intervene between the input and output of an AI agent and determine its behavior and judgment "quirks," and may include weight coefficients, biases, activation function parameters, etc. "Evaluation function" may refer to, for example, a mathematical formula or algorithm that quantifies the quality of the AI ​​agent's output. "Sensitivity" may refer to, for example, an indicator that shows how much the output changes in response to a change in input, or how much the internal state is updated in response to new information.

[0165] According to the third aspect, the "learning method" of the AI ​​agent itself can be precisely controlled and designed. Rather than simply using symptomatic control that prohibits output (black-box control), by performing "gray-box" control that directly adjusts the internal value judgment criteria (what is considered good, what to react strongly to, and what to ignore), it becomes possible to guide and stabilize the behavior of the AI ​​agent from within to match the "correct answer structure" required in the field. As a result, even when training with the same data, it becomes possible to acquire "smarter" behavior that is more in line with the intentions of the field.

[0166] Furthermore, the behavior of the AI ​​agent is determined not only by the superficial output results, but also by the state of the underlying parameter variables. As shown in Figure 10, there are countless paths from input to output, but by setting the parameter variables, a specific path (an "essential path" that aligns with the correct answer structure) will be preferentially selected.

[0167] The learning control unit 232 incorporates, for example, indicators that are considered important in the field (quality, delivery time, safety, etc.) into the evaluation function with weights, and sets the evaluation axis in a multidimensional manner. For example, by giving higher rewards to "accuracy of answers" and "clear explanation of rationale" than to "response speed," the AI's personality can be adjusted to be more cautious. Furthermore, mediating variables are set to reduce sensitivity to accidental successes and uncertain information (preventing overreactions), thereby increasing tolerance to noise. In addition, the system adjusts the intensity of memory updates, strongly retaining important learning experiences and making experiences that are close to noise easier to forget. As a result, the AI ​​agent can learn with qualitative judgment criteria rather than simply depending on the amount of data. This can be described as a process that instills the "values" and "instincts of judgment" of the field into the AI ​​agent, forming an axis for autonomous judgment.

[0168] (2.2.4) Fourth aspect In a fourth aspect of this embodiment, the learning control unit 232 may evaluate the suitability of the information generated or searched in connection with the execution of the learning process by the AI ​​agent to the correct answer structure or other constraints, and exclude from the target space any information that is determined not to suit the AI ​​agent's internal state and should not be used in subsequent learning or inference.

[0169] Here, exclusion may mean, for example, not only pre-emptive static filtering, but also the adaptive elimination of inappropriate information based on the results of the AI ​​agent's actual exploration. "Internal state" may mean, for example, the state of knowledge, memories, context, or model parameters held by the AI ​​agent.

[0170] According to the fourth aspect, even in areas where the quality cannot be judged without actually performing the search (so-called "sandpit C"), the learning process can be dynamically evaluated and inappropriate elements can be eliminated. This prevents inappropriate information that becomes apparent as the search progresses, noise newly generated during operation, or information that becomes inappropriate due to changes in the external environment from being mixed into the learning data. As a result, the target space is continuously purified over time, narrowing down to a higher purity of candidate regions for the correct answer, thus improving performance and stabilizing it in long-term operation.

[0171] The remaining "pile C" in Figure 13 represents an area that cannot be uniformly eliminated in the initial stages. For example, this includes manufacturing methods that appear effective in theory but do not fit the company's equipment constraints, or new business ideas whose effectiveness cannot be determined without experimentation. These are areas where "you won't know until you try," and require trial and error.

[0172] The learning control unit 232 causes the AI ​​agent to perform additional learning (exploration) under its control and monitors the results. For example, in the process of generating new business ideas, the AI ​​agent generates proposals, and the learning control unit 232 simulates them. If it is found that the output from region C is unstable, inconsistent, or does not meet on-site constraints (cost, legal regulations, etc.), the learning control unit 232 will remove region C as a target for "second-stage pruning".

[0173] Figure 14 shows the state after sandpile C has been removed, leaving only sandpile A, which is a candidate for the correct answer. Through this process, the AI ​​agent can concentrate on the promising region A without wasting resources on unnecessary searches, thereby improving both search efficiency and the quality of the answer.

[0174] (2.2.5) Fifth aspect In a fifth aspect of this embodiment, the learning control unit 232 may perform a sharpening process that adjusts the set parameters to increase the probability of the AI ​​agent selecting the region that remains as information to be preferentially used in the target space, thereby reducing the variability of the judgment.

[0175] Here, "sharpening" may refer to a process that adjusts the AI ​​agent to converge more precisely to the correct answer by, for example, trimming the area around the correct answer region (pile A) in the search space and making the gradient of the evaluation function steeper. "Selection probability" may refer to, for example, the probabilistic tendency for a specific one to be selected from among multiple action candidates or answer candidates.

[0176] According to the fifth aspect, the "sharpness" and "reproducibility" of the AI ​​agent's judgment can be dramatically improved. Not only is the range of correct answers narrowed, but fluctuations in judgment within that range are suppressed, and the agent is guided to always select the optimal solution. As a result, the same (or very close) high-quality output is always returned for the same input, eliminating the "whims" and "variability" inherent in AI, and significantly increasing the reliability of the AI ​​agent in actual work. This reduces rework and verification costs caused by ambiguous answers.

[0177] Figure 15 shows the state after the sand pile A has been sharpened. Compared to the state in Figure 14, the base of the pile has been further eroded, resulting in a steeper shape towards the summit (optimal solution). If the sand pile is gentle, where one lands near the summit tends to be left to chance, and answers that are "generally correct, but not optimal" are likely to be generated. However, with a sharpened sand pile, one is inevitably led to the summit. In Figure 15, the base of the sand pile A contains data sets that appear frequently but have a low impact, such as information that is well known and can be performed as a matter of course, or general information. In other words, this base area is a region that has little influence on the exploration within sand pile A or the reasoning towards the optimal solution. Therefore, by eliminating this base area, reasoning similar to climbing to the optimal solution is effectively supported. Furthermore, this elimination makes it possible to sharpen the sand pile A without distorting or destroying the space, while preserving the shape of sand pile A, that is, the important core part of the exploration space.

[0178] This sharpening effect is achieved by increasing the sensitivity of the mediating variables and clarifying the difference (gradient) in rewards. For example, by giving low rewards to vague answers such as "I think that..." and extremely high rewards to specific and well-founded answers such as "Based on rule X, Y is true," the AI ​​agent is strongly motivated to make more accurate inferences.

[0179] Furthermore, in a sharpened space, subtle differences in conditions (input fluctuations) are prevented from manifesting as large differences in output, while still allowing for a sensitive response to important differences. This is the final stage of control processing that evolves AI agents from "machines that give answers somewhat intuitively" to "systems that make precise judgments like professionals" and enhances their reasoning logic.

[0180] (2.2.6) Sixth aspect In a sixth aspect of this embodiment, the inference control unit 233 may, during the inference phase, set boundary conditions to prevent deviation from the target space after sharpening, and perform control to block or modify the output of the AI ​​agent when the output exceeds the boundary conditions. Alternatively, the control may be performed using the target space in which the boundary conditions have been set.

[0181] Here, “boundary conditions” may mean, for example, rules, thresholds, or guardrails that define the range of behaviors or outputs that an AI agent is allowed to perform. “Blocking or correcting” may mean, for example, filtering out inappropriate remarks, replacing them with safe, canned responses, or escalating them to a human.

[0182] According to the sixth aspect, it functions as a "last line of defense" that guarantees the stable behavior shaped during the learning phase during the inference phase. Even if the AI ​​agent attempts to deviate from the shaped space due to unexpected input, adversarial attacks, or probabilistic fluctuations, the system can forcibly switch to the safe side (fail-safe), minimizing operational risks (compliance violations, spread of misinformation, system downtime, etc.). This allows the concept of defense in depth to be implemented in the AI ​​system.

[0183] Figure 16 shows an image where boundary conditions (guardrails) are further set around the sharpened sand pile A. In conventional AI implementations, there were many cases where control relied solely on these guardrails, but this was like trying to control an unstable engine with only brakes, and had its limitations. In this embodiment, it is important that the guardrails are used only as a supplement after the shaping of sand pile A (control during the learning phase) is complete.

[0184] In a shaped sandcastle A, the AI ​​agent should ideally follow the correct path, but in the real world, perfect prediction is impossible, and LLMs inherently possess probabilistic behavior. The inference control unit 233 applies rules such as "do not respond if confidence level is below a certain level," "stop if certain prohibited words are included," "correct if the numerical response exceeds physical constraints," and "mask patterns that appear to be personal information," thereby ensuring robust operational safety. This is a safety mechanism to prevent fatal failures while respecting autonomy.

[0185] (2.2.7) Seventh aspect In a seventh aspect of this embodiment, the learning control unit 232 may, during the inference phase, fix the internal state of the AI ​​agent so that the set parameter variables are not changed, thereby prohibiting the immediate updating of the parameter variables based on the results of the inference process.

[0186] Here, "fixed control" may mean, for example, temporarily freezing the learning function (weight update, etc.) and operating in an inference-only mode. "Immediately" may mean, for example, learning is performed immediately after inference is performed, or in real time.

[0187] According to the seventh aspect, it is possible to prevent drift (performance degradation) due to unexpected learning during operation and data contamination due to malicious input. By separating inference and learning temporally and functionally, a robust operational system can be established that allows for the continuous provision of a stable version of the service while systematically advancing learning using only data that has been validated in the background. This eliminates the risk of the AI ​​agent unexpectedly changing and ensures the reproducibility and auditability of the system.

[0188] Furthermore, if an AI agent learns the content of user interactions in real time (online learning), there is a risk of malicious user attacks (such as prompt injection) or the agent behaving erratically if it takes user feedback based on misunderstandings at face value. This is a known vulnerability, as seen in cases where corporate chatbots have been trained to make discriminatory remarks.

[0189] In this embodiment, during the inference phase, the parameter is treated as "read-only," and no changes in its state are permitted. This ensures that the AI ​​agent always maintains the tested behavior as designed. Learning should always be performed "offline" or "separately" (e.g., batch learning), and the user interaction environment is defined not as a learning environment, but as a place to demonstrate results. This ensures robustness suitable for enterprise use.

[0190] (2.2.8) Eighth aspect In the eighth aspect of this embodiment, the learning control unit 232 may, during the learning phase, prohibit the use of inference results generated by the AI ​​agent as learning data without verification, and may only use data whose conformance to the correct answer structure has been verified as the target for learning.

[0191] Here, "without verification" may mean automatically adopting data without, for example, human verification, comparison with reliable external data sources, or going through predetermined quality check logic. "Verified data" may mean, for example, highly reliable data that has been determined to conform to the ground truth.

[0192] According to the eighth aspect, it is possible to break the vicious cycle of "Garbage Everywhere" (self-propagation of errors) in AI agents. By breaking the cycle in which the AI ​​itself relearns the erroneous information it outputs as fact, it becomes possible to maintain a high level of knowledge base quality (information hygiene) even in long-term operation. This prevents a situation where the AI ​​becomes "confidently lying" over time (model collapse).

[0193] The root cause of the vicious cycle shown in Figure 9 lies in the fact that inference results are directly fed back into the learning process (uncritical intake of self-generated data). For example, if an incorrect summary generated by AI is incorporated as "correct data" in the next learning stage, the error is reinforced and the truth is buried. This phenomenon is also known as "model collapse" in learning using synthetic data.

[0194] In this embodiment, a strong gate (filter) is installed in this return path. Specifically, data generated by the AI ​​agent is initially placed in a "holding" state (quarantine state), and is only approved as training data after being compared with the correct structure and reviewed by a human (Human-in-the-loop). Unapproved data is either discarded or stored only as an analysis log. This strict information quarantine process prevents the introduction of noise and maintains the integrity of the AI ​​agent.

[0195] (2.2.9) The ninth aspect In the ninth aspect of this embodiment, the parameter includes the update frequency or update intensity of the internal memory held by the AI ​​agent, and the learning control unit 232 may adjust the parameter in the learning process so that the closer the information is to the correct answer structure, the higher the degree to which it is retained in the internal memory.

[0196] Here, "internal memory" may mean, for example, a vector database, a long-term memory module, or historical information within a context window. "Update intensity" may mean, for example, the degree to which existing memory is overwritten by new information, or the priority or duration for which new information is retained as memory.

[0197] According to the ninth aspect, it becomes possible to manage memory in a balanced way according to the importance of the information. Important information (information close to the correct structure, immutable rules, etc.) is less likely to be forgotten and is strongly retained, while unimportant information (temporary exceptions, noise, etc.) is controlled to be less likely to be retained or to be forgotten early. This allows for effective use of the AI ​​agent's memory capacity while ensuring that the basis of its judgment does not waver. As a result, it is possible to create an agent that does not forget "important things" even after long-term operation, and to prevent catastrophic forgetting.

[0198] Just as with human memory, "what an AI agent remembers and what it forgets" is a crucial factor in determining the quality of its intelligence. If everything is remembered equally, important knowledge will be buried in noise and cannot be retrieved.

[0199] In this embodiment, memory weighting is performed based on distance from the correct structure (anchor). For example, information regarding official company rules and success stories (core principles) is given a high update intensity and is retained as long-term memory. On the other hand, records of temporary exception handling and uncertain information (everyday context) are given a low update intensity and are kept in short-term memory or forgotten early. By performing this hierarchical control of memory, the AI ​​agent can always maintain the "ideal state" while flexibly adapting to new information.

[0200] (2.2.10) Tenth aspect In a tenth aspect of this embodiment, the learning control unit 232 may execute a learning process involving a state update for the AI ​​agent only when at least one of the following conditions is met: user approval, evaluation results from an external system, or fulfillment of predetermined verification conditions.

[0201] Here, "state update" may mean, for example, rewriting parameters, adding data to the knowledge base, or updating the version. "Validation conditions" may mean, for example, that the accuracy test score exceeds a certain threshold, that a specific regression test is passed, or that safety in the simulation environment is confirmed.

[0202] According to the tenth aspect, the timing of learning execution can be strictly controlled (gate control). Instead of uncontrolled continuous learning, by allowing gradual growth only at times when quality is guaranteed, humans can reliably control the direction of the AI ​​agent's evolution. This prevents unintended deterioration and changes where the responsibility is unclear, enabling well-governed AI operations within the organization.

[0203] This introduces concepts similar to "release decisions" and "CI / CD (Continuous Integration / Continuous Delivery)" in software development into the AI ​​learning process. Instead of the AI ​​agent changing on its own, its state is only updated after a judgment (approval) is made on whether "this change should be applied."

[0204] For example, after inputting new training data, automated tests (regression tests) are run to confirm that the accuracy of answers against existing correct answer structures has not decreased. Alternatively, important changes are not applied until a human manager presses the approval button (promotion from staging to production). By implementing such gates, the quality of the AI ​​agent is kept above a certain level, and unexpected quality degradation is prevented.

[0205] (2.2.11) Eleventh aspect In the eleventh aspect of this embodiment, the learning control unit 232 may record the internal state of the AI ​​agent or the values ​​of the parameter variables before and after the execution of the learning process, and restore the state of the AI ​​agent to any point in time based on the record.

[0206] Here, "restore (rollback)" may also mean completely restoring, for example, the model parameters, memory state, or configuration files of an AI agent to a specific point in the past (a snapshot).

[0207] According to the eleventh aspect, even if performance degrades due to learning or a fatal error is learned, it can be immediately restored to the previous normal state. This enhances the availability and resilience of the AI ​​system's operation, minimizing downtime and negative impact on business operations. It also functions as a risk hedge when attempting exploratory learning, shortening the mean time to recovery (MTTR).

[0208] Furthermore, because AI agents have complex internal states, once their balance is disrupted (for example, if they start using a biased dataset as a result of learning from it, they may become abusive on certain topics), it can be difficult to identify the cause or partially repair the problem.

[0209] In this embodiment, a "save point (snapshot)" is created after each learning session, providing an environment where you can "start over" at any time. This is a crucial lifeline when performing experimental learning or highly uncertain tasks such as exploring new business ventures. The reassurance of being able to quickly return to "the best state from last week" if something goes wrong makes it possible to actively engage in learning (such as exploring sandpit C) and A / B testing.

[0210] (2.2.12) The twelfth aspect In a twelfth aspect of this embodiment, the learning control unit 232 may, during the learning phase, gradually reduce or reshape the range of information included in the target space and control the AI ​​agent so that its inference results converge along a specific evaluation gradient.

[0211] Here, "evaluation gradient" may mean, for example, the direction in which the value of the evaluation function improves (the slope of the mountain to climb). "Convergence" may mean, for example, that the search range is narrowed to a specific area (potential correct answer) and the output stabilizes. "Stepwise" may mean, for example, going through a sequence of processes as shown in Figures 12 to 15 (deletion of B → evaluation and deletion of C → sharpening of A).

[0212] According to the twelfth embodiment, the AI ​​agent's learning process can be advanced in a stepwise and planned manner, from broad exploration to detailed refinement. This realizes the overall picture of the so-called "sandcastle model," where the optimal solution is reliably and efficiently guided to the solution without destroying the internal correct answer structure by gradually removing material from the outside. This prevents the learning process from going astray or diverging, and allows the system to reach a practical level of performance in the shortest possible time. Furthermore, learning efficiency is maximized by progressing from easy tasks (elimination of inappropriate areas) to difficult tasks (refinement of the correct answer area), similar to curriculum learning.

[0213] This comprehensively defines the control method using the "sandcastle collapse model" of this embodiment. Instead of immediately seeking the correct answer (the peak), the process involves first eliminating "errors (sandcastle B)," then verifying and eliminating "uncertain elements (sandcastle C)," and finally refining the accuracy (sharpening) within the remaining possibilities (sandcastle A).

[0214] By adhering to this sequence, the AI ​​agent can avoid initial erroneous assumptions (convergence to local optima) and reach the designer's intended goal (correct solution structure) without getting lost. This transforms the "training" of the AI ​​agent from subjective trial and error to a reproducible "engineering process," providing a foundation for ensuring the reliability and predictability of the entire system.

[0215] (2.4) Example of operation Referring to the flowchart in Figure 17, an example of the operation of the information processing system 1 (particularly the server device 100a) according to this embodiment will be described. This example of operation is a specific processing flow based on the "sand pile collapse model" for controlling the search space and learning process of the AI ​​agent.

[0216] In the following explanation, each step is primarily performed by the learning control unit 232, but the inference control unit 233, the AI ​​agent processing unit 231, and even human judgment (administrators, field staff, etc.) may intervene as appropriate.

[0217] In step S101, the learning control unit 232 extracts and defines the correct structure of the work environment. This is the most important step as "preparation" before running the AI ​​agent, and it is not simply data collection, but the process of verbalizing the "framework" of the work. Specifically, the learning control unit 232 receives input from the user (for example, the person in charge of the work or an experienced worker in the field) and defines what constitutes the "correct answer" in the work and what constitutes "unacceptable behavior".

[0218] The correct answer structure defined here includes the following elements:

[0219] • Mandatory / Non-Negotiable Rules: Laws, safety standards, company regulations, etc., that must be strictly followed in order to perform the job. For example, this includes clear thresholds and prohibitions such as "the amount of a specific chemical substance must be less than X%" or "personal information will not be transmitted externally without approval." These are areas where AI's "creativity" should not be exercised, but rather where strict adherence is required.

[0220] • Frequently occurring and repeatedly used decisions: These are patterned decision logics that are frequently used in daily operations. For example, rules such as "apply a special discount if the customer rank is A and has a purchase history of less than one year" or "place an automatic order if the inventory level falls below a certain value." These are core elements of improving operational efficiency and are standard practices that should be reliably taught to AI agents.

[0221] • Cases where failure is unacceptable: Past trouble cases and serious patterns that could lead to brand damage. These are defined as "negative correct structures (structures to be avoided)." For example, this includes inappropriate response flows that have led to complaints in the past, or sales pitches that violate compliance. By clearly identifying these, the risk of AI agents repeating the same mistakes is reduced.

[0222] Step S101 clarifies the "core" of the space the AI ​​agent should explore and the "boundaries" that it must absolutely not enter. If this definition remains ambiguous, it can cause the AI ​​agent to wander aimlessly in later stages, so it is recommended to take the time to reach a consensus.

[0223] In step S102, the learning control unit 232 performs initial learning (fixed learning) of the correct answer structure. The learning control unit 232 instructs the AI ​​agent to learn the correct answer structure extracted in step S101 not as a "search target," but as a "prerequisite (anchor)."

[0224] Here, the AI ​​agent's parameters are adjusted to stabilize using this correct answer structure as a baseline. In other words, judgments that conform to the correct answer structure are given high rewards, and judgments that deviate are given strong penalties, thereby constructing an "irretrievable path" (a main road of thought) within the AI ​​agent's internal state. This learning process is called "fixed learning" because it aims to firmly establish the correct answer as "knowledge" rather than allowing the AI ​​agent to freely experiment. Through this process, even when encountering unfamiliar situations, the AI ​​agent will make judgments based on this "anchor," making it less likely to make fundamental judgment errors. This can be described as a phase in which the AI ​​agent is thoroughly instilled with "professional ethics" and "basic operations."

[0225] In step S103, the learning control unit 232 performs an initial classification of the search space. The learning control unit 232 conceptually classifies the space of all information or action candidates accessible to the AI ​​agent (target space) into the following three regions (sand piles).

[0226] • Region A (Group of candidate solutions; equivalent to "Pile A"): A region that is consistent with the correct solution structure defined in step S101 and where further optimization and discovery are expected. This is the main region that we ultimately want the AI ​​agent to explore. For example, it may include derivatives of existing successful patterns or unexplored areas that are extensions of the correct solution structure.

[0227] • Domain B (Unsuitable as a learning target; equivalent to "Sandpile B"): A domain that contradicts the correct answer structure but superficially resembles the correct answer, or where short-term rewards are easily obtained. Examples include procedures based on outdated manuals, textbook answers that do not suit the actual situation on site, or shortcuts that seem efficient but carry long-term risks. This is a domain that can mislead the AI ​​agent, and if left unchecked, the AI ​​agent will easily converge on this domain.

[0228] • Area C (Potential for error / misdirection; equivalent to "Sandpile C"): This area cannot be immediately determined as correct or incorrect at this point and cannot be evaluated without actually conducting exploration (simulation, etc.). Examples include new business ideas, unverified technologies or methods, or strategies whose effectiveness may change due to changes in the external environment. This area is treated as a "pending" area.

[0229] In step S104, the learning control unit 232 performs the deletion of pile B (first pruning). This is an extremely important step that is performed *before* the AI ​​agent begins a full-scale search.

[0230] The learning control unit 232 compares the region B, which is clearly inconsistent with the correct structure (anchor) fixed in step S102, with the correct structure and completely excludes the region B from the learning and exploration targets. For example, it may delete the relevant data from the learning dataset, set a filter to prohibit access to region B, or provide a strong negative reward to discourage the user from selecting an action belonging to region B.

[0231] If the search is started without this process, the AI ​​agent will easily converge on the "closest and easiest to climb" sand dune B, distorting the entire learning process (a failure due to initial sensitivity). For example, this prevents a situation where an AI agent that has learned from outdated data continues to recommend discontinued products to customers. This step proactively blocks that risk and ensures that the AI ​​agent starts learning in the right direction from the beginning.

[0232] In step S105, the learning control unit 232 reconstructs the search space. As a result of removing the sandpile B, the search space consists only of the promising region A and the unknown region C. The learning control unit 232 defines this purified space as the new search range for the AI ​​agent. This ensures that the AI ​​agent's resources are not wasted on the useless region B, but are instead concentrated on the potentially valuable regions A and C.

[0233] In step S106, the learning control unit 232 determines whether or not retraining of the correct answer structure is necessary. For example, if the definition of the correct answer structure itself needs to be modified due to the deletion of region B, or if new preconditions are added, the system returns to step S101 and restarts the process (Yes). For example, this may occur if the scope of application of the basic rule (region A) needs to be redefined as a result of deleting a specific exception handling as region B. If retraining is not necessary, the system proceeds to the next step (No).

[0234] In step S107, the learning control unit 232 designs the evaluation function and parameters. From here, the control design for the exploration including the unknown region C begins. The learning control unit 232 sets parameters that will serve as guidelines for the exploration in regions A and C.

[0235] • Evaluation Criteria: What constitutes "good"? For example, in the case of exploring new business ventures, instead of just "profitability," set multifaceted evaluation criteria such as "feasibility," "customer response (positive emotional response)," "social impact," and "consistency with company resources." This prevents tunnel vision in AI agents.

[0236] • Rewards: Which behaviors should be reinforced? High rewards are set for new discoveries that align with the correct structure, while low rewards are set for high-risk bets or weakly supported inferences.

[0237] • Sensitivity: How much the AI ​​reacts to changes in input. Adjust the sensitivity appropriately to avoid overreacting to noise (unnecessary information or temporary trends in region C). This will help stabilize the AI ​​agent's behavior.

[0238] In step S108, the learning control unit 232 performs additional learning (under control). Based on the evaluation function designed in step S107, the learning control unit 232 causes the AI ​​agent to continue learning within the shaped space (A+C).

[0239] This learning process is not laissez-faire autonomous learning, but rather conducted under strict supervision. The learning control unit 232 monitors changes in the AI ​​agent's internal state (parameters) and allows updates only within a set range (gate control). For example, if the parameter changes rapidly due to learning, or if a specific parameter shows an abnormal value, the learning process is immediately stopped and measures such as a rollback are taken. This prevents "runaway" or "unintended changes" during the learning process.

[0240] In step S109, the learning control unit 232 performs evaluation and deletion (second pruning) of region C. As a result of the additional learning in step S108, the "misleading and destabilizing factors" that were included in region C become apparent.

[0241] For example, if a simulation of a new idea (part of domain C) reveals that the costs are not justifiable or that the risks are too high under certain market conditions, the learning control unit 232 will determine that part as "unsuitable." Similarly, if the AI ​​agent's responses become inconsistent or it starts making inexplicable inferences, the learning data that caused this is identified and removed. Then, before proceeding to the next exploration phase, this unsuitable part is removed from the exploration space.

[0242] This process is not a one-time decision, but may be repeated as the search progresses. This further narrows the search space, reduces uncertainty, and converges it towards a more reliable region centered around sandpile A.

[0243] In step S110, the inference control unit 233 performs a search (inference). Here, after first pruning (deletion of B) and second pruning (deletion of C), the refined search space (sharpened sandpile A) is used to allow the AI ​​agent to derive a solution to the actual business problem.

[0244] At this stage, the search space is sufficiently shaped, and the gradient of the evaluation function is steep (sharp), allowing the AI ​​agent to select the optimal solution with high probability without hesitation. In other words, the AI ​​agent is like driving on a "well-maintained highway," exhibiting high performance and stability. Furthermore, the inference control unit 233 applies boundary conditions (guardrails) as needed, activating a final safety mechanism to prevent the inference results from deviating from the shaped space. For example, it checks for prohibited words in the output content and verifies the validity of numerical values ​​in real time, and blocks the output if there are problems. Alternatively, it may recognize the target space as an even narrower range or an even sharper target space with boundary conditions set, increasing the likelihood of obtaining an appropriate optimal solution in a short time.

[0245] In step S111, the learning control unit 232 or the inference control unit 233 determines whether to terminate the search. If the desired accuracy or result is obtained, or if the user instructs termination, the process is terminated (Yes).

[0246] If sufficient results are not obtained, or if the correct answer structure needs to be revised due to environmental changes, the process returns to step S107 (redesigning the evaluation function) or, in some cases, step S101 (redefining the correct answer structure) to continue the control loop (No). For example, if the previous "correct answer" becomes invalid due to changes in market trends, the process returns to redefining the correct answer structure and constructing a new search space.

[0247] As demonstrated by the above operational examples, the process of introducing and operating an AI agent can be approached by first designing and shaping the search space, rather than first training it with data. This allows for the construction and maintenance of a robust AI agent capable of withstanding production use, without being misled by the apparent success of the proof-of-concept (PoC) phase. In particular, the process of early elimination of pile B (initial misdirection) and the gradual evaluation and elimination of pile C (potential risks) is an essential procedure for avoiding an "uncontrollable" state of the AI ​​agent. This provides a systematic methodology for elevating the AI ​​agent from a mere "convenient tool" to a "reliable partner" embodying the wisdom and standards of the organization.

[0248] (3) Other embodiments In the embodiment described above, at least a portion of the processing performed by the server device 100 may be modified to be performed on the terminal device 200. In this case, at least a portion of each part included in the processing unit 130 may be moved to the terminal device 200, or at least a portion of each part included in the storage unit 120 may be moved to the terminal device 200.

[0249] The operation flow and operation examples in the above-described embodiments do not necessarily have to be executed chronologically in the order shown in the flowchart. For example, the steps in the operation may be executed in a different order than that shown in the flowchart, or they may be executed in parallel. Also, some of the steps in the operation may be deleted, or further steps may be added to the process.

[0250] A program may be provided that causes a computer (information processing device) to perform the operations according to the above embodiment. The program may be recorded on a computer-readable medium. Using a computer-readable medium, it is possible to install the program on a computer. Here, the computer-readable medium on which the program is recorded may be a non-transient storage medium. The non-transient storage medium is not particularly limited, but may be a storage medium such as a CD-ROM or DVD-ROM.

[0251] The functions realized by the information processing system 1 may be implemented in a circuit or processing circuitry, including a general-purpose processor, an application-specific processor, an integrated circuit, an ASIC (Application Specific Integrated Circuit), a CPU (a Central Processing Unit), conventional circuits, and / or a combination thereof, which is programmed to realize the described functions. A processor, including transistors and other circuits, is considered a circuit or processing circuitry. A processor may be a programmed processor that executes a program stored in memory. In this specification, a circuit, a unit, or a means is hardware programmed to realize or execute the described functions. Such hardware may be any hardware disclosed herein, or any hardware known to be programmed to realize or execute the described functions. If such hardware is a processor which is considered a type of circuit, then the circuit, a means, or a unit is a combination of hardware and software used to constitute such hardware and / or processor.

[0252] As used herein, the terms “based on” and “according to” do not mean “based solely on” or “according solely to” unless otherwise specified. “Based on” means both “based solely on” and “based at least partially on.” Similarly, “according to” means both “based solely on” and “according at least partially to.” Furthermore, the terms “include,” “comprise,” and their variations do not mean to include only the listed items, but may include only the listed items, or may include additional items in addition to the listed items. Also, the term “or” as used herein is not intended to mean exclusive OR. Where articles are added by translation, such as a, an, and the in English, these articles are considered plural unless the context clearly indicates otherwise.

[0253] Although the embodiments have been described in detail above with reference to the drawings, the specific configuration is not limited to those described above, and various design changes can be made without departing from the gist of the invention. [Explanation of Symbols]

[0254] 1: Information Processing System 5: Network 100, 100a: Server equipment 110: Communications Department 120: Storage section 121: Target Information Storage Unit 122:Trend information storage unit 123:Classification result storage unit 130: Processing Unit 131: 1st acquisition part 132: 1st derivation part 133: 1st Analysis Department 134: 1st classification part 135: Output section 136:Analysis Department 136a:Second acquisition part 136b: Second derived part 136c:Second Analysis Department 136d: 2nd classification section 136e:Extraction part 136f: Comparison section 200: Terminal device 210: Communications Department 220: Storage section 221: Internal state memory unit 230: Processing Unit 231: AI Agent Processing Unit 232: Learning Control Unit 233: Inference Control Unit

Claims

1. An information processing system that evaluates information or regions within a target space that is the target of processing for an artificial intelligence capable of performing learning processing and inference processing, and performs a process based on the evaluation to distinguish between information that should be preferentially used for processing by the artificial intelligence and information that should be excluded or suppressed from processing by the artificial intelligence, The aforementioned target space includes a first acquisition unit that acquires target information relating to the subject of the investigation, Based on the aforementioned target information, a first derivation unit derives trend information regarding the trends of the subject of the investigation, As part of the aforementioned evaluation, the first analysis unit analyzes the correlation between the aforementioned target information and the aforementioned trend information, Based on the results of the analysis, the system includes a first classification unit that classifies the target information into first information that is significant for deriving the trend information and second information that is less significant than the first information. Information processing system.

2. The system further includes an output unit that outputs at least one of the first information and the second information classified by the first classification unit. The information processing system according to claim 1.

3. The aforementioned subject information is technical information that includes at least one of patent information and non-patent information, The first derivation unit derives technology trend information as the trend information, The first analysis unit analyzes the correlation between the technical information and the technical trend information, The first information is significant information for deriving the technological trend information from the technical information. The information processing system according to claim 1.

4. The first information includes at least one of the following: characteristic word information, engineer information, and technology classification information. The information processing system according to claim 3.

5. The aforementioned information is market information that includes at least one of product information and service information. The first derivation unit derives market trend information as the trend information, The first analysis unit analyzes the correlation between the market information and the market trend information, The first information is significant information for deriving the market trend information from the market information. The information processing system according to claim 1.

6. The aforementioned first information includes at least one of the following: characteristic word information, information provider information, information medium information, product classification information, and service classification information. The information processing system according to claim 5.

7. The first derivation unit performs the derivation using a learning model that takes the target information as input and outputs the trend information. The information processing system according to claim 1.

8. The system further includes an analysis unit that performs an analysis on the trends of the subject of the investigation based on the first information classified by the first classification unit. The information processing system according to claim 1.

9. The analysis unit has a second derivation unit that uses the first information in priority over the second information to derive trend information regarding the trends of the subject of the investigation, and outputs the derived trend information. The information processing system according to claim 8.

10. The aforementioned analysis unit, A second acquisition unit that acquires new target information concerning the aforementioned survey target, The system includes an extraction unit that extracts information corresponding to the first information from the new target information and outputs the extracted information. The information processing system according to claim 8.

11. The aforementioned analysis unit, A second acquisition unit that acquires new target information concerning the aforementioned survey target, A second derivation unit derives new trend information regarding the trends of the subject of the survey based on the new target information, The second analysis unit analyzes the new correlation between the aforementioned new target information and the aforementioned new trend information, A comparison unit that performs at least one of the following: a comparison between the aforementioned trend information and the aforementioned new trend information, and a comparison between the aforementioned correlation and the aforementioned new correlation, and outputs information based on the comparison results. The information processing system according to claim 8.

12. The aforementioned analysis unit, A second acquisition unit that acquires new target information concerning the aforementioned survey target, A second derivation unit derives new trend information regarding the trends of the subject of the survey based on the new target information, The second analysis unit analyzes the new correlation between the aforementioned new target information and the aforementioned new trend information, Based on the results of the analysis, a second classification unit classifies the new target information into new first information that is significant for deriving the new trend information, and new second information that is less significant than the new first information. The system includes a comparison unit that compares the first information and the second information with the new first information and the new second information, and outputs information based on the comparison result. The information processing system according to claim 8.

13. An information processing system that evaluates information or regions within a target space that is the subject of processing by artificial intelligence, and performs a process based on the evaluation to distinguish between information that should be preferentially used for processing by the artificial intelligence and information that should be excluded or suppressed from processing by the artificial intelligence, The aforementioned artificial intelligence is an AI agent capable of autonomously performing learning and inference processes. The information processing system in question is In the learning phase, a learning control unit defines and shapes the target space which becomes the search range for the AI ​​agent to perform the learning process and the inference process by performing the evaluation and the distinction process, The inference phase includes an inference control unit that causes the AI ​​agent to perform the inference process based on the reshaped target space. Information processing system.

14. The learning control unit sets a known correct answer structure that indicates a valid judgment criterion or constraint as an initial condition, and reconfigures the target space to exclude information or regions that do not satisfy the initial condition from the target space as regions that the AI ​​agent should not explore. The information processing system according to claim 13.

15. The learning control unit sets at least one of the evaluation function, reward, sensitivity, and update intensity in the learning process as a parameter for the reconstructed target space, and causes the AI ​​agent to execute the learning process based on this setting. The information processing system according to claim 14.

16. The learning control unit evaluates the suitability of the information generated or searched in connection with the execution of the learning process by the AI ​​agent to the correct answer structure or other constraints, and excludes information that is determined not to suit the conditions from the target space as information that should not be reflected in the internal state of the AI ​​agent and should not be used in subsequent learning or inference. The information processing system according to claim 15.

17. The learning control unit performs a sharpening process on the regions remaining as information to be preferentially used in the target space, thereby increasing the selection probability of the AI ​​agent's inference and reducing the variability of its judgments. The information processing system according to claim 16.

18. The inference control unit, in the inference phase, sets boundary conditions to prevent deviation from the target space after the sharpening process, and if the output of the AI ​​agent exceeds the boundary conditions, it blocks or modifies the output, or controls the target space as defined by the boundary conditions. The information processing system according to claim 17.

19. The learning control unit, in the inference phase, maintains a fixed internal state of the AI ​​agent so that the set parameter is not changed, and prohibits the immediate updating of the parameter based on the results of the inference process. The information processing system according to claim 15.

20. The learning control unit prohibits the use of inference results generated by the AI ​​agent as training data without verification during the learning phase, and only uses data whose conformance to the correct answer structure has been verified as training data. The information processing system according to claim 14.

21. The aforementioned parameter includes the update frequency or update intensity of the internal memory held by the AI ​​agent. The learning control unit adjusts the parameter in the learning process so that information closer to the correct answer structure is retained more effectively in the internal memory. The information processing system according to claim 15.

22. The learning control unit executes a learning process involving a state update for the AI ​​agent only when at least one of the following conditions is met: user approval, evaluation results from an external system, or fulfillment of predetermined verification conditions. The information processing system according to claim 13.

23. The learning control unit records the internal state or parameter values ​​of the AI ​​agent before and after the execution of the learning process, and restores the state of the AI ​​agent to any point in time based on the record. The information processing system according to claim 13.

24. The learning control unit, in the learning phase, gradually reduces or reshapes the range of information included in the target space and controls the AI ​​agent's inference results to converge along a specific evaluation gradient. The information processing system according to claim 13.

25. An information processing method performed by an information processing system, comprising: evaluating information or regions within a target space that is the target of processing for an artificial intelligence capable of performing learning processing and inference processing; and, based on the evaluation, distinguishing between information that should be preferentially used for processing by the artificial intelligence and information that should be excluded or suppressed from processing by the artificial intelligence within the target space, The steps include: acquiring target information related to the subject of the investigation as the target space; Based on the aforementioned target information, the step of deriving trend information regarding the trends of the subject of the survey, The aforementioned evaluation includes a step of analyzing the correlation between the target information and the trend information, The method includes the step of classifying the target information into first information that is significant for deriving the trend information, and second information that is less significant than the first information, based on the results of the analysis. Information processing methods.

26. An information processing method performed by an information processing system, comprising: evaluating information or regions within a target space that is subject to processing by artificial intelligence; and, based on the evaluation, distinguishing between information that should be preferentially used for processing by the artificial intelligence and information that should be excluded or suppressed from processing by the artificial intelligence within the target space, The aforementioned artificial intelligence is an AI agent capable of autonomously performing learning and inference processes. The information processing method is In the learning phase, the process includes defining and shaping the target space, which becomes the search range for the AI ​​agent to perform the learning process and the inference process, by performing the evaluation and the distinction process. The inference phase includes the step of causing the AI ​​agent to perform the inference process based on the reshaped target space. Information processing methods.

27. ​​Cause an information processing system to execute the information processing method described in Claim 25 or 26. program.